Multimedia processing method, related device, storage medium and computer program product

By acquiring media description information of multimedia, determining thematic semantics, and generating features by matching reference content from the content library, the problem of inaccurate feature extraction in existing technologies is solved, and multimedia features can be accurately acquired without relying on optical character recognition and automatic speech recognition.

CN116881392BActive Publication Date: 2026-02-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210308761.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2026-02-13
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

Existing technologies rely on optical character recognition and automatic speech recognition when acquiring multimedia features, which leads to inaccurate feature extraction.

Method used

By acquiring media description information of multimedia, determining its thematic semantics, and matching similar reference content from the content library, media features are generated, avoiding reliance on optical character recognition and automatic speech recognition.

Benefits of technology

It enables accurate acquisition of multimedia media features without relying on optical character recognition and automatic speech recognition, enriching the semantics of feature representation and improving feature accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116881392B_ABST
    Figure CN116881392B_ABST
Patent Text Reader

Abstract

The application discloses a multimedia processing method, related equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring media description information of a target multimedia, and determining feature information of the target multimedia according to the media description information, wherein the feature information of the target multimedia is used for describing theme semantics of the target multimedia; acquiring one or more reference contents matched with the theme semantics from a content library; and finally generating media features of the target multimedia according to the feature information of the target multimedia and feature information of each reference content. According to the embodiment of the application, the media features of the multimedia can be accurately acquired.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a multimedia processing method, related equipment, a storage medium and a computer program product. BACKGROUND

[0002] In the scenarios of content recommendation, content classification and the like, media features of the multimedia need to be acquired first, so as to facilitate subsequent recommendation, classification, screening and the like of the multimedia. Generally, there are two ways to acquire the media features, the first way is to directly take features of descriptive text such as a title and a content summary of the extracted multimedia as the media features, and the second way is to first identify media content text in the multimedia through optical character recognition and automatic speech recognition, and then take features of the extracted descriptive text and the media content text as the media features. The media features extracted by the former way represent less content about the target multimedia, so that the information represented by the media features is not accurate, and the media features extracted by the latter way are too dependent on the recognition effect of the optical character recognition and the automatic speech recognition, so if the recognition effect is not good, the media features extracted finally will also be wrong, so the media features extracted by the existing way are not accurate enough. Therefore, how to accurately acquire the media features of the multimedia is a problem to be solved at present. SUMMARY

[0003] Embodiments of the present application provide a multimedia processing method, related equipment, a storage medium and a computer program product, which can accurately acquire media features of the multimedia.

[0004] In one aspect, an embodiment of the present application provides a multimedia processing method, comprising:

[0005] acquiring media description information of a target multimedia, and determining feature information of the target multimedia according to the media description information, the feature information of the target multimedia being used to describe theme semantics of the target multimedia;

[0006] acquiring the theme semantics described by the feature information of the target multimedia, and acquiring one or more reference contents matching the theme semantics from a content library; any reference content in the one or more reference contents corresponds to a corresponding feature information;

[0007] generating media features of the target multimedia according to the feature information of the target multimedia and feature information of each reference content.

[0008] In one aspect, an embodiment of the present application provides a multimedia processing device, comprising:

[0009] The acquisition unit is configured to acquire media description information of a target multimedia, and determine feature information of the target multimedia according to the media description information, wherein the feature information of the target multimedia is used to describe a theme semantic of the target multimedia.

[0010] The acquisition unit is further configured to acquire one or more reference contents matching the theme semantic from a content library according to the theme semantic described by the feature information of the target multimedia, wherein any reference content of the one or more reference contents corresponds to a corresponding feature information.

[0011] The processing unit is configured to generate a media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content.

[0012] In an aspect, the embodiments of the present application further provide a computer device, comprising:

[0013] The processor is adapted to implement one or more computer programs.

[0014] The computer storage medium stores one or more computer programs, and the one or more computer programs are adapted to be loaded by the processor and execute the multimedia processing method.

[0015] In an aspect, the embodiments of the present application further provide a computer storage medium, which stores one or more computer programs, and the one or more computer programs are adapted to be loaded by the processor and execute the multimedia processing method.

[0016] In an aspect, the embodiments of the present application further provide a computer program product or a computer program, which comprises a computer program, and the computer program is adapted to be loaded by the processor and execute the multimedia processing method.

[0017] In the embodiment of the present application, in the process of acquiring the media feature of the target multimedia, the media description information of the target multimedia can be acquired first, and the feature information of the target multimedia is determined according to the media description information, wherein the feature information of the target multimedia is used to describe the theme semantics of the target multimedia; then, one or more reference contents matching the theme semantics are acquired from the content library, so as to finally generate the media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content. It can be seen that in the embodiment of the present application, without relying on optical character recognition and automatic speech recognition technologies, one or more reference contents matching the theme semantics of the target multimedia are acquired from the content library, so that more text information of the media content of the target multimedia can be acquired, so that the media feature of the target multimedia generated according to the feature information of the target multimedia and the feature information of each reference content can accurately and sufficiently represent the media content of the target multimedia. Therefore, the embodiment of the present application can accurately acquire the media feature of the multimedia. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0019] Figure 1 is a schematic diagram of a multimedia processing system architecture provided by an embodiment of the present application;

[0020] Figure 2 is a flowchart of a multimedia processing method provided by an embodiment of the present application;

[0021] Figure 3 is a flowchart of another multimedia processing method provided by an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of a twin network model architecture provided by an embodiment of the present application;

[0023] Figure 5 is a schematic diagram of determining feature similarity provided by an embodiment of the present application;

[0024] Figure 6 is a flowchart of another multimedia processing method provided by an embodiment of the present application;

[0025] Figure 7 is a structural schematic diagram of a multimedia processing device provided by an embodiment of the present application;

[0026] Figure 8 is a structural schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0027] The present application provides a multimedia processing method. Before generating media features of a target multimedia, the method can obtain a theme semantics (i.e. the meaning of the central idea or main content) of the target multimedia based on media description information of the target multimedia. Then, the method can further obtain reference content from a content library, the reference content expressing a meaning similar or identical to the theme semantics of the target multimedia according to the theme semantics of the target multimedia. Thus, the feature information of the target multimedia determined by the media description information and the feature information of the reference content are both used to generate media features of the target multimedia. Therefore, the generated media features not only represent the theme semantics of the target multimedia, but also represent the semantics of the reference content similar to the theme semantics of the target multimedia. The media features can accurately and sufficiently represent the media content of the target multimedia. It can be seen that the semantics represented by the media features is more abundant, and the semantics represented by the media features is more accurate. In addition, the reference content is obtained from the content library by the theme semantics of the target multimedia, and no optical character recognition and automatic speech recognition are used. Thus, the generated media features are not incorrect due to the poor recognition effect of the optical character recognition and automatic speech recognition.

[0028] In the present application, the media refers to a carrier for transmitting and storing information, and the information includes text, image, video, audio, etc. The multimedia refers to a kind of transmission media combining two or more media. In the present application, the target multimedia can be composed of only one kind of media, such as only video, only audio, etc. or can be composed of multiple media, such as video with music, image with text, etc. The media description information of the target multimedia refers to text information for describing the media content of the target multimedia, which can be the title, introduction, tag (i.e. a keyword tag for searching and finding) of the target multimedia, etc.

[0029] In addition, a feature refers to a characteristic of a thing that is different from other things, and feature information refers to the most basic information that can reflect the characteristics of a thing. Therefore, the feature information of the target multimedia can represent the media content with the characteristics in the target multimedia, so as to distinguish the media description information from other information; that is, the feature information of the target multimedia can be used to describe the theme semantics of the target multimedia. The theme semantics refers to the meaning of the central idea or main content of the target multimedia. Specifically, the feature information can be composed of one or more words that can represent the characteristics of the target multimedia, or composed of one or more symbols that can represent the characteristics of the target multimedia.

[0030] In the embodiments of the present application, since the media description information can only contain the title of the target multimedia, or only contain the introduction of the target multimedia, the feature information determined by the media description information contains only the title feature, or the introduction feature. When the media description information contains both the title of the target multimedia and the introduction of the target multimedia, the feature information determined by the media description information contains both the title feature and the introduction feature. Therefore, the feature information can be composed of one or more feature vectors, which can be obtained by inputting the media description information of the target multimedia into a pre-trained language model for feature encoding. The pre-trained language model refers to a model capable of encoding natural language, including BERT (Bidirectional Encoder Representations from Transformer), XLNet (Auto-Regressive and Auto-Encoding Model), LayoutLM (Pre-training of Text and Layout for Document Image Understanding), etc. The process of encoding text by the pre-trained language model is a common technical means for those skilled in the art, and will not be described here. For ease of illustration, in the absence of special instructions, all embodiments below will be described in detail with respect to the steps involved in the multimedia processing method provided by the embodiments of the present application, taking the feature information including one or more feature vectors as an example.

[0031] In addition, the content library refers to a module with a storage function, such as a local or remote database, a local or remote memory, and the like. The content library stores N texts, where N is a positive integer greater than a preset threshold, and the preset threshold can be a specific value, such as 0, 9999, or the like, or a value range, such as 0-999, 8000-10000, or the like. Preferably, since the texts in the content library need to be matched with various multimedia theme semantics, the number of texts stored in the content library can be relatively large, and the value of N can also be relatively large. The reference content refers to a text in the content library that matches the theme semantics. The texts stored in the content library can be articles such as news and information content, sentences such as titles and content summaries, or keywords, and the like, which are not limited herein. Optionally, the texts stored in the content library can also be media content texts of images, videos, and audio. Specifically, before storing the image, video, or audio into the content library, the content text of the image, video, or audio can be obtained through text recognition (such as optical character recognition), speech recognition (such as an automatic speech recognition model such as a hidden Markov model), or manual recognition, and the like, and then the content text is stored into the content library. Optionally, since the feature information of the target text in the content library can be used as the feature information of the reference content for expanding the media features of the target multimedia, and is closely related to the accuracy of the media features of the target multimedia, the content text of the image, video, or audio needs to be repeatedly checked through manual checking or system checking, or the like, before being stored into the content library as the target text, so that the content text of the image, video, or audio is error-free and matches the image, video, or audio.

[0032] Based on the above multimedia processing method, an embodiment of the present application provides a multimedia processing system, which can be seen from Figure 1 , Figure 1The multimedia processing system shown may include a terminal device 101 and a server 102. The terminal device 101 may include any one or more of smartphones, tablets, laptops, desktop computers, smart in-vehicle devices, and smart wearable devices. Various clients (apps) can run on the terminal device, such as multimedia playback clients, social clients, browser clients, news feed clients, educational clients, etc. The server 102 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal device 101 and the server 102 may communicate directly or indirectly via wired or wireless communication, which is not limited herein.

[0033] In one embodiment, the above multimedia processing method can be performed solely by... Figure 1 The multimedia processing system described herein is executed by terminal device 101. The specific execution process is as follows: the terminal device first obtains the media description information of the target multimedia, and based on this information, obtains the subject semantics of the target multimedia; then, based on the subject semantics of the target multimedia, the terminal device retrieves one or more reference contents from the content library; finally, based on the feature information of the target multimedia and the feature information of each reference content, the terminal device generates the media features of the target multimedia. Optionally, the above multimedia processing method can also be executed solely by... Figure 1 The execution process of the server 102 in the multimedia processing system shown can be found in the specific execution process of the terminal device, and will not be repeated here.

[0034] In another embodiment, the above-described multimedia processing method can run in a multimedia processing system, which may include terminal devices and servers, wherein the multimedia processing method can be... Figure 1The multimedia processing system shown utilizes terminal devices and server devices to jointly complete the process. Specifically, the terminal device acquires the target multimedia and its media description information, then uploads the target multimedia and its media description information to the server. Upon receiving the media description information, the server determines the target multimedia's feature information based on the media description information. Then, the server obtains the thematic semantics described by the target multimedia's feature information to retrieve one or more reference contents matching the thematic semantics from the server's content library. Finally, the server generates media features for the target multimedia based on the feature information of the target multimedia and the feature information of each reference content. Optionally, the server can also transmit the generated media features to the terminal device so that the terminal device can classify, filter, and recommend the target multimedia based on its media features.

[0035] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating a multimedia processing method provided in an embodiment of this application. This multimedia processing method can be executed by the aforementioned terminal device or server, such as... Figure 2 As shown, the multimedia processing method includes steps S201-S203:

[0036] S201, Obtain media description information of the target multimedia, and determine the feature information of the target multimedia based on the media description information.

[0037] In this embodiment, the media description information can be stored in association with the target multimedia, thus allowing direct retrieval of the media description information from the corresponding database or storage. For example, when a user uploads a short video, they edit and save the video name and description; therefore, the video name and description can serve as the media description information stored in association with the short video.

[0038] Optionally, media description information can also include text information describing important media content in the target multimedia. For example, when the target multimedia is a video with sound, key image frames in the video (generally the first few frames or a frame serving as the video cover) and corresponding key audio segments can be identified first. Then, text recognition is performed on the key image frames, and speech recognition is performed on the key audio segments. Finally, the text information of the identified key image frames and key audio segments is used as the media description information of the target multimedia. Since the media description information is only text information describing important media content in the target multimedia, and not text information corresponding to all media content in the target multimedia—that is, the media description information here is just a short sentence, not a long article describing all media content in the target multimedia—the probability of errors during recognition is relatively low because the amount of content requiring speech and text recognition is small. Furthermore, when errors occur in speech and text recognition, the impact of the error on the semantics of the short sentence is far less than the impact of the error on the semantics of the entire long article.

[0039] S202, Obtain the topic semantics described by the feature information of the target multimedia, and retrieve one or more reference contents that match the topic semantics from the content library.

[0040] In this embodiment, since the feature information of the target multimedia can characterize the core and distinctive media content of the target multimedia, and the thematic semantics of the target multimedia refers to the meaning of the central idea or main content of the media content presented by the target multimedia, the feature information of the target multimedia can characterize the thematic semantics of the target multimedia. Therefore, the thematic semantics of the target multimedia can be obtained directly by analyzing the feature information of the target multimedia. The content library mentioned in this embodiment can be a database or storage device in the terminal device or server executing step S202. For example, if the target multimedia is an image depicting tulips, crowds, and the sun, then the distinctive media content in the image can be determined to be flowers, tourists, and sunshine. Therefore, the thematic semantics of the image can be obtained as "flower viewing on a sunny day" based on the characteristics of the image.

[0041] In addition, matching the theme semantics refers to that the text semantics of the reference content is similar or approximately the same as the theme semantics. For example, assuming that video A is a target multimedia, and through analyzing the feature information of video A, it is determined that the theme semantics B described by video A is "China team wins one and loses three in the World Cup preliminary match and ranks fifth", and the text semantics of text 1 in the content library is "correct feeding method of 3-month-old cat", the text semantics of text 2 is "analysis of the game situation of the China team in 4 matches of the World Cup preliminary match", and the text semantics of text 3 is "problem solving method of high school reading comprehension"; through comparing the theme semantics B with the text semantics of the three texts, it is determined that the theme semantics B is most similar to the text semantics of text 2, and thus it is determined that text 2 is the reference content matching the theme semantics B.

[0042] In addition, any one of the one or more reference contents corresponds to a corresponding feature information. The specific explanation of the feature information about the reference content can be referred to the aforementioned feature information about the target multimedia, which is not described herein.

[0043] S203, generating the media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content.

[0044] In an embodiment, the manner of generating the media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content can be: performing feature fusion on the feature information of the target multimedia and the feature information of each reference content to obtain fused features; and taking the fused features as the media feature of the target multimedia. The feature fusion refers to a method and process of fusing two or more different features into one feature. Specifically, performing feature fusion on the feature information of the target multimedia and the feature information of each reference content can be to perform normalization processing on the feature information of the target multimedia and the feature information of each reference content, that is, to unify the representation of the feature information of the target multimedia and the feature information of each reference content. For example, the feature information of the target multimedia is a 4x6-dimensional feature vector, and the feature information of each reference content is an 8x9-dimensional feature vector, and then the feature information of the target multimedia and the feature information of each reference content are converted into a 2x3-dimensional feature vector, and then the outer product of the two 2x3-dimensional feature vectors is performed to obtain a feature vector, thereby realizing feature fusion.

[0045] Optionally, according to the feature information of the target multimedia and the feature information of each reference content, the manner of generating the media feature of the target multimedia can also be: performing feature splicing on the feature information of the target multimedia and the feature information of each reference content to obtain spliced features; and then taking the spliced features as the media feature of the target multimedia. The feature splicing can be directly adding multiple features to obtain a new feature, or can be pre-setting the weight of each feature, and then adding multiple features according to the weight to obtain a new feature, which is not limited herein.

[0046] Based on the above description, the media feature of the target multimedia obtained by the multimedia processing method proposed in the embodiments of the present application can accurately represent the media content of the target multimedia without relying on recognition technologies such as speech recognition or text recognition. Therefore, the multimedia processing method can be widely applied in multiple application scenarios, including but not limited to: scenarios requiring classification of target multimedia, scenarios requiring screening of target multimedia, and scenarios requiring recommendation of target multimedia.

[0047] In the embodiments of the present application, before generating the media feature of the target multimedia, the feature information of the target multimedia is determined based on the media description information of the target multimedia, and then the theme semantics of the target multimedia is obtained, so that the reference content whose meaning expressed by the central idea or main content is similar or identical to the theme semantics of the target multimedia can be conveniently obtained from the content library according to the theme semantics of the target multimedia subsequently; so that the media feature of the target multimedia can finally be generated by fusing or splicing the feature information of the target multimedia and the feature information of the reference content. That is, the media content of the target multimedia described by the media description information is limited, so that the theme semantics of the target multimedia represented by the feature information of the target multimedia obtained finally is not rich. The embodiments of the present application achieve that the text capable of accurately representing the media content of the target multimedia can be obtained without relying on technologies such as optical character recognition and automatic speech recognition, by obtaining the reference content matching the theme semantics of the target multimedia from the content library, and taking the feature information of the reference content as one of the feature information used to generate the media feature of the target multimedia. Moreover, the semantics represented by the media feature generated based on the feature information of the target multimedia and the feature information of the reference content not only contains the theme semantics of the target multimedia, but also contains the semantics of the reference content similar or identical to the theme semantics of the target multimedia, so that the media feature can accurately and sufficiently represent the media content of the target multimedia. Therefore, by using the above multimedia processing method, the feature information used to generate the media feature of the target multimedia can be expanded without relying on technologies such as optical character recognition and automatic speech recognition, so that the media feature of the multimedia can be accurately obtained.

[0048] Please refer to Figure 3 , Figure 3 is a schematic flowchart of another multimedia processing method provided by the embodiment of the present application, which can also be executed by the terminal device or the server mentioned above. As shown in Figure 3 , the multimedia processing method comprises steps S301-S306:

[0049] S301, obtaining media description information of a target multimedia, and determining feature information of the target multimedia according to the media description information.

[0050] S302, obtaining a theme semantic described by the feature information of the target multimedia.

[0051] In steps S301-S302, since the aforementioned media description information refers to text such as title and synopsis, the way of determining the feature information of the target multimedia according to the media description information can be inputting the media description information into a first feature extraction module of a Siamese network model (the Siamese network model refers to a network model composed of two model branches with the same parameters, which are respectively referred to as a first feature extraction module and a second feature extraction module), and performing feature encoding processing on one or more sentences in the media description information by the first feature extraction module to obtain one or more feature vectors of the target multimedia, which constitute the feature information of the target multimedia. Exemplarily, the Siamese network model can be a Sentence-Bert model, which is composed of two Bert model branches sharing the same parameters. By the Bert model branches in the Sentence-Bert model, feature encoding of the sentences can be realized, so that each sentence can obtain a feature vector with semantics. In the embodiment, the first feature extraction module refers to one of the Bert model branches of the Sentence-Bert model, and the second feature extraction module refers to the other Bert model branch of the Sentence-Bert model except the one indicated by the first feature extraction module.

[0052] It should be noted that the closer the distance between the feature vectors of the sentences with similar semantics is. As can be seen, each feature vector has semantics, that is, the feature information of the target multimedia is used to describe the theme semantic of the target multimedia, so the theme semantic can be directly obtained by analyzing the feature information of the target multimedia, and the analysis process can be decoding processing, semantic extraction, etc., which will not be described here.

[0053] S303, performing text segmentation processing on the target text according to the content library to obtain a text segmentation set corresponding to the target text.

[0054] In the embodiments of the present application, the target text refers to any one of the N texts stored in the content library. The text segmentation processing refers to segmenting the target text into one or more sentences or phrases. The one or more texts obtained after the text segmentation processing of the target text are one or more segmented texts, and the segmented text set contains the one or more segmented texts. Specifically, the text segmentation processing can be performed by detecting punctuation marks such as “,” “.” and “!” to determine whether a sentence has ended. If it is determined that a sentence has ended, the sentence is segmented from the text to form a separate sentence (i.e., a segmented text).

[0055] S304, generating feature information of the target text according to the text segmentation set of the target text, and calculating the feature similarity between the feature information of the target text and the feature information of the target multimedia.

[0056] In the embodiments of the present application, the way of generating the feature information of the target text according to the text segmentation set of the target text can be: first, performing feature extraction processing on each segmented text in the text segmentation set respectively to obtain the text features corresponding to each segmented text respectively; then, the feature set composed of the text features corresponding to each segmented text is the feature set corresponding to the text segmentation set; finally, the feature set corresponding to the text segmentation set is taken as the feature information of the target text.

[0057] In specific implementation, the way of performing feature extraction processing on each segmented text in the text segmentation set respectively can be inputting each segmented text in the text segmentation set into the second feature extraction module in the Siamese network model mentioned in steps S301 to S302, and performing feature encoding processing on each segmented text by the second feature extraction module to obtain the feature vector corresponding to each segmented text respectively, and the feature vector corresponding to each segmented text respectively is the text feature corresponding to each segmented text respectively. That is, the feature set corresponding to the text segmentation set is composed of the feature vectors corresponding to each segmented text of the target text respectively.

[0058] In addition, the feature similarity refers to a similarity between the feature information of the target text and the feature information of the target multimedia. The feature information of the target text is composed of text features of each segmented text in the target text. Therefore, the feature similarity between the feature information of the target text and the feature information of the target multimedia can be calculated in the following manner: the similarity between the text features of each segmented text in the target text and the feature information of the target multimedia is calculated, and then the feature similarity is determined by counting the similarity between the text features of each segmented text and the feature information of the target multimedia. Specifically, the feature correlation between the feature information of the target multimedia and any text feature in the feature set (i.e., the similarity between the features) can be calculated to obtain a first feature correlation between the feature information of the target multimedia and any text feature, wherein the first feature correlation refers to the similarity between the feature information of the target multimedia and the any text feature. Then, the first feature correlation between the feature information of the target multimedia and each text feature is taken as the feature similarity between the feature information of the target text and the feature information of the target multimedia.

[0059] In actual applications, the specific process of the above feature correlation calculation can be that the feature vector of the target multimedia (i.e., the feature information of the target multimedia) and the feature vector corresponding to any segmented text (i.e., any text feature) are first subjected to vector similarity calculation, and then the calculated vector similarity between the two feature vectors is taken as the first feature correlation between the feature information of the target multimedia and any text feature. Specifically, the similarity calculation method between the two feature vectors can be the calculation of the Euclidean distance, Pearson correlation coefficient or cosine distance of the two feature vectors, and is not limited herein. In addition, since the first feature correlation is used to indicate the similarity between the two feature vectors, the first feature correlation can be a specific numerical value, such as 80%, 0.763, etc.

[0060] Optionally, since the feature information of the target multimedia mentioned in step S302 may include multiple feature vectors of the target multimedia, when there are multiple feature vectors of the target multimedia, the similarity between each feature vector of the target multimedia and the feature vector corresponding to any segmented text can be calculated to obtain multiple feature correlation degrees. Then, the minimum or maximum feature correlation degree among the multiple feature correlation degrees can be used as the first feature correlation degree between the feature information of the target multimedia and any text feature; alternatively, the average feature correlation degree of the multiple feature correlation degrees can be calculated, and then the average feature correlation degree can be used as the first feature correlation degree between the feature information of the target multimedia and any text feature, which is not limited here.

[0061] In a practical implementation, the similarity calculation module in the Siamese network model can also be used to calculate the similarity between the feature information of the target multimedia and the feature information of all texts in the content library. For example, please refer to the appendix. Figure 4 The diagram illustrates the architecture of a Siamese network model, which includes a first feature extraction module 401, a second feature extraction module 402, and a similarity calculation module 403. Both the first and second feature extraction modules 401 and 402 consist of a BERT model and a pooling layer. The BERT model is used for feature encoding of the sentence, while the pooling layer ensures that each feature vector has the same feature dimension, i.e., a fixed feature dimension.

[0062] Firstly, the target multimedia is determined as a short video Q, then the title Sentence A (i.e., media description information) of the short video Q is obtained, and the Sentence A is input into the first feature extraction module 401, and the feature vector u (i.e., the feature information of the short video Q) is obtained after the Sentence A is encoded by the first feature extraction module 401; then, the Sentence B (i.e., the segmented text) in the target text Y of the content library is input into the second feature extraction module 402, and the feature vector v (i.e., the text feature) is obtained after the Sentence B is encoded by the second feature extraction module 402; then, the vector similarity of the feature vector u and the feature vector v is calculated by the similarity calculation module 403 through the cosine function Cosin-sim, and the vector similarity of the feature vector u and the feature vector v is taken as the first feature correlation degree of the feature information of the short video Q and the text feature of the Sentence B. Since the target text Y has other segmented texts in addition to the Sentence B after the text segmentation processing, after the first feature correlation degree of the feature information of the target multimedia and the text feature of the Sentence B is calculated, the other segmented texts in the target text Y need to be input into the second feature extraction module 402 one by one, so as to obtain the feature vector corresponding to each segmented text, and then the first feature correlation degree of the feature vector corresponding to each segmented text and the feature information of the short video Q is calculated through the similarity calculation module 403. Finally, the first feature correlation degree of the feature vector corresponding to all the segmented texts of the target text Y and the feature information of the short video Q is taken as the feature similarity between the feature information of the target text Y and the feature information of the short video Q.

[0063] In addition, the twin network model is a trained network model, and the architecture of the network model before training can be referred to the architecture schematic diagram of the twin network model in the Figure 4 training process of the twin network model specifically includes: first, obtaining a training sample pair and a sample label of the training sample pair. The training sample pair contains two training samples, one of the training samples in the training sample pair can be a title, a synopsis, a central sentence or the like for describing a multimedia text, and the other training sample can be an article, news information or the like. The sample label is used to indicate the semantic similarity between the two training samples contained in the training sample pair, and the sample label is mainly manually labeled. The sample label can be a numerical value, such as 90%, 0.4, etc., or a text label, such as strong correlation, weak correlation, negative correlation, etc., which is not limited here.

[0064] Then, the training sample pairs are input into the pre-training network model. The pre-training first feature extraction module extracts the feature vector of each training sample in the training sample pair. The pre-training similarity calculation module then calculates the similarity between the feature vectors of the two training samples, thus outputting the predicted similarity between the training sample pairs. Finally, based on the difference between the predicted similarity and the semantic similarity indicated by the sample labels, the model parameters of the network model are continuously updated until the difference between the predicted similarity and the semantic similarity indicated by the sample labels is small or less than a preset difference. At this point, training stops, and the trained Siamese network model is obtained.

[0065] Optional, please see the appendix. Figure 5 This diagram illustrates a method for determining feature similarity. First, the title of short video W (i.e., the media description information of short video W) is feature-encoded by a first feature extraction module to obtain a feature vector T. Then, the target text A in the content library is segmented to obtain a text segmentation set 502, which includes Sentences 1 to 24. Next, the second feature extraction module performs feature encoding on Sentences 1 to 24 respectively, resulting in a feature set 501 corresponding to text segmentation set 502. Feature set 501 includes feature vectors 1 to 24, and feature set 501 represents the feature information of target text A. Finally, the similarity calculation module calculates the first feature correlation degree between feature vector T and each feature vector in feature set 501, and uses this first feature correlation degree as the feature similarity between the feature information of short video W (i.e., feature vector T) and the feature information of target text A.

[0066] Furthermore, to improve efficiency in generating media features for the target multimedia, the feature set corresponding to each text in the content library can be pre-determined through the second feature extraction module in the Siamese network model. In other words, the feature information of each text in the content library is pre-obtained. Then, after obtaining the media description information of the target multimedia, the feature information of the target multimedia can be determined based on the media description information, and the feature similarity between the feature information of any text in the content library and the feature information of the target multimedia can be directly calculated.

[0067] S305, when the similarity between the textual semantics of the target text and the topical semantics of the target multimedia meets the preset similarity level, the target text is used as a reference content that matches the topical semantics of the target multimedia, until one or more reference contents that match the topical semantics are obtained from the content library.

[0068] In the embodiments of the present application, since the text semantics of the target text is composed of the text semantics of each segmented text in the target text, the feature similarity can be used to determine whether the similarity between the text semantics of the target text and the theme semantics of the target multimedia meets the preset similarity.

[0069] Specifically, the text features corresponding to the first feature correlation greater than the correlation threshold can be selected from the feature set of the target text according to the first feature correlation between the feature information of the target multimedia and each text feature in the feature set of the target text. Then, the number of the selected text features corresponding to the first feature correlation greater than the correlation threshold is determined. Finally, when the determined number of features is greater than or equal to the preset number, it is determined that the similarity between the text semantics of the target text and the theme semantics of the target multimedia meets the preset similarity. The correlation threshold can be a specific numerical value such as a percentage, a decimal, etc., such as 80%, 0.6, etc. In addition, the preset number can be a preset fixed number value, such as 4, 8, 15, etc. Alternatively, the preset number can also be a number value calculated according to the total number of segmented texts of the target text, for example, if the preset number is set to one fifth of the total number of segmented texts of the target text, when the total number of segmented texts of the target text is 50, the preset number should be 10.

[0070] For example, the correlation threshold is set to 0.6 and the preset number is set to 5. After the target text E is segmented, 10 segmented texts are obtained. After feature extraction processing is performed on the 10 segmented texts, text feature 1 to text feature 10 of the target text E can be obtained. Then, the feature information of the target multimedia R is obtained, and the first feature correlation between the feature information of the target multimedia R and each text feature of the target text E is calculated. The first feature correlation between the feature information of the target multimedia R and text feature 1 to text feature 10 is 0.8, 0.4, 0.2, 0.7, 0.3, 0.5, 0.5, 0.4, 0.7, and 0.8, respectively. As can be seen, the number of text features with a first feature correlation greater than 0.6 is 4, which is less than the preset number 5, so it can be determined that the feature similarity between the feature information of the target text E and the feature information of the target multimedia R does not meet the preset similarity between the text semantics of the target text and the theme semantics of the target multimedia.

[0071] Through the media description information explained in step S201, it can be known that the title of the target multimedia can be included in the media description information, and the title of the target multimedia is often the sentence that can most briefly and concisely summarize the central idea of the target multimedia. Therefore, optionally, it can be determined that the title feature of the title of the target multimedia is included in the feature information of the target multimedia, so that when the number of the text features corresponding to the first feature association degree greater than the association degree threshold value selected is greater than or equal to the preset number, the second feature association degree between each text feature in the feature set and the title feature of the target multimedia is obtained, and the text feature corresponding to the second feature association degree greater than the target threshold value is selected from the feature set. Then, when the number of the selected text features corresponding to the second feature association degree greater than the target threshold value is greater than or equal to the reference number, it is determined that the similarity between the text semantics of the target text and the theme semantics of the target multimedia satisfies the preset similarity degree.

[0072] The second feature association degree refers to the feature association degree between each text feature in the feature set and the title feature of the target multimedia. The reference number can be a preset fixed number value, such as 4, 8, 15, etc. The target threshold value can be a specific numerical value such as a percentage, a decimal, etc., such as 80%, 0.6, etc. Meanwhile, the target threshold value can be the same as the association degree threshold value, or can be different from the association degree threshold value, which is not limited here.

[0073] For example, the correlation threshold is set as 0.6, the preset number is 5, the target threshold is 0.7, and the reference number is 2. After the segmentation processing of the target text U, 10 segmented texts are obtained. After the feature extraction processing of the 10 segmented texts, the text features 1-10 of the target text U can be obtained. Then, the feature information of the target multimedia D is obtained, and the feature information of the target multimedia D includes the title feature and the introduction feature of the target multimedia D. The feature correlation of the title feature of the target multimedia D and each text feature of the target text E is calculated, and the feature correlation of the title feature of the target multimedia D and the text features 1-10 is 0.8, 0.4, 0.8, 0.9, 0.3, 0.5, 0.4, 0.6, 0.8, and 0.3, respectively. The feature correlation of the introduction feature of the target multimedia D and the text features 1-10 is 0.7, 0.5, 0.8, 0.7, 0.3, 0.8, 0.8, 0.6, 0.6, and 0.2, respectively. Then, the average feature correlation of the calculated title feature and introduction feature and the feature correlation of the text features 1-10 is taken as the first feature correlation of the feature information of the target multimedia D and the text features 1-10, so that the first feature correlation of the feature information of the target multimedia D and the text features 1-10 is 0.75, 0.45, 0.8, 0.8, 0.3, 0.65, 0.6, 0.6, 0.75, and 0.25, respectively. As can be seen, the number of text features with the first feature correlation greater than 0.6 is 5, which is equal to the preset number 5.

[0074] After the determined feature number is equal to the preset number, the second feature correlation between the title feature of the target multimedia D and the text features 1-10 is obtained, that is, 0.8, 0.4, 0.8, 0.9, 0.3, 0.5, 0.4, 0.6, 0.8, and 0.3. As can be seen, the number of second feature correlations greater than or equal to 0.7 is 4, which is greater than the reference number 2, so it can be determined that the similarity degree between the feature information of the target text E and the feature information of the target multimedia R indicated by the feature similarity of the text semantic of the target text and the theme semantic of the target multimedia does not satisfy the preset similarity degree.

[0075] Optionally, because the title is often the sentence that can most briefly and concisely summarize the central idea of the text, when the number of text features corresponding to the first feature correlation degree greater than the correlation degree threshold is greater than or equal to the preset number, the title feature of the title (one of the segmented texts of the target text) of the target text in the feature set can also be obtained, and the second feature correlation degree between the title feature of the target text and the title feature of the target multimedia. If the second feature correlation degree between the title feature of the target text and the title feature of the target multimedia is greater than the target threshold, it is determined that the similarity degree between the text semantics of the target text and the theme semantics of the target multimedia satisfies the preset similarity degree.

[0076] In a possible implementation, there can be multiple target texts in the content library, and the similarity degree between the text semantics of the target texts and the theme semantics of the target multimedia satisfies the preset similarity degree. Although the text semantics of the reference content obtained from the content library is similar to the theme semantics of the target multimedia, the reference content also has some sentences with low similarity degree to the theme semantics of the target multimedia. The more reference content obtained, the more sentences with low similarity degree to the theme semantics of the target multimedia are introduced, which may cause the semantic deviation of the media features of the target multimedia generated subsequently.

[0077] Therefore, it can be set that the reference content is stopped being obtained from the content library after a target number of reference contents are obtained from the content library, wherein the target number is a fixed number value, and is generally a positive integer, such as 1, 2, 3, etc. Optionally, after all texts in the content library are traversed and multiple reference contents are selected, the feature similarity between the feature information of each reference content and the feature information of the target multimedia can be sorted in descending order, and the reference contents with the top feature similarities are selected as target reference contents. Finally, the media features of the target multimedia are generated according to the feature information of the target multimedia and the feature information of each target reference content.

[0078] S306, generating the media features of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content.

[0079] The specific implementation of step S306 can refer to the specific implementation of step S203, which will not be repeated here.

[0080] In an embodiment, after the media features of the target multimedia are generated, the multimedia can also be recommended according to the media features. The specific process includes: first, generating the media label of the target multimedia according to the media features of the target multimedia; then, when detecting the demand for multimedia recommendation for the target user, obtaining the user demand information of the target user; finally, recommending the multimedia matching the demand information of the target user to the target user according to the media labels of different multimedia.

[0081] The specific way of generating the media tag of the target multimedia according to the media feature of the target multimedia can be inputting the media feature of the target multimedia into a machine learning model such as a softmax (a regression model for multi-classification) or a decision tree model to obtain the media tag of the target multimedia. The training process of the machine learning model such as the softmax and the decision tree model is a common technical means for those skilled in the art, and will not be described here. It should be noted that the media tag is used to indicate the media type of the target multimedia. The media type can be a keyword briefly describing the semantics represented by the media feature of the target multimedia. For example, when the semantics represented by the media feature of the target multimedia is "correct feeding method of 3-month-old kittens", the media tag can be "cute pets", "animal feeding", "cat", etc. Alternatively, the media type can also be the type of the multimedia, such as text type, video type, image type, etc.

[0082] In addition, the target user refers to a user who has a multimedia browsing and viewing demand, that is, a user who needs and allows multimedia recommendation. The specific way of recommending the multimedia matching the demand information of the target user to the target user according to the media tag of different multimedia can be: first, performing feature extraction processing on the demand information of the target user to obtain the demand feature of the demand information; then, determining the demand tag of the demand information according to the demand feature of the demand information; and finally, determining the multimedia matching the demand information of the target user as the multimedia matching the demand information of the target user by matching the media tag and the demand tag. The specific implementation process of the feature extraction processing and the determination of the demand tag of the demand information can be referred to steps S302 and S604, which will not be described here.

[0083] In another embodiment, steps S301 to S306 can be performed by a server in the multimedia processing system shown in Figure 1 After the server generates the media feature of the target multimedia, the server can also perform multimedia recommendation for a target user using a terminal device in the multimedia processing system shown in Figure 1 The specific process includes steps S3701-S3704:

[0084] S3701, the server generates a media tag of the target multimedia according to the media feature of the target multimedia.

[0085] The specific implementation process of step S3701 can be referred to the specific way of generating the media tag of the target multimedia in step S306, which will not be described here.

[0086] S3702, when the server detects the demand for multimedia recommendation for the target user, the server obtains the user demand information of the target user.

[0087] In this embodiment, the target user refers to a user using a terminal device. The detection of whether the target user has a demand for multimedia recommendation can be specifically: when a search operation of the target user is detected, it is determined that the target user has a demand for multimedia recommendation; when it is detected that the target user is browsing multimedia such as short videos and images, it is determined that the target user has a demand for multimedia recommendation. Alternatively, whether the target user has a demand for multimedia recommendation can also be determined by other manners, which are not limited herein. In addition, the demand information refers to the recommendation demand of the target user, such as the preference of watching short videos, etc. The demand information can be generated by historical browsing data, search data, historical operation data about multimedia, etc. of the target user.

[0088] S3703, the server recommends multimedia matching the demand information of the target user to the target user according to the media tags of different multimedia.

[0089] S3704, the server sends the matched multimedia to the client.

[0090] In the embodiments of the present application, since whether the one or more reference contents matched with the theme semantics are obtained from the content library is judged by whether the meaning expressed by the central idea or main content of the reference content (i.e., the text semantics of the reference content) is similar or identical to the theme semantics of the target multimedia, and the feature information of the target multimedia can represent the theme semantics of the target multimedia, and the feature information of the reference content can represent the text semantics of the reference content, therefore, by calculating the feature similarity between the feature information of the target text and the feature information of the target multimedia, the similarity between the text semantics of the target text and the theme semantics of the target multimedia can be conveniently judged, so as to determine whether the target text is the reference content matched with the theme semantics. Meanwhile, since the feature information of the target text is composed of the text features of each segmented text in the target text, when calculating the feature similarity between the feature information of the target text and the feature information of the target multimedia, the similarity between the text features of each segmented text in the target text and the feature information of the target multimedia needs to be calculated first, and then the feature similarity is further determined. Finally, according to the feature information of the target multimedia and the semantic represented by the media features of the target multimedia generated based on the feature information of each reference content, the semantic represented by the media features of the target multimedia not only contains the theme semantics of the target multimedia, but also contains the semantic of the reference content matched with the theme semantics of the target multimedia, so that the media features can accurately and sufficiently represent the media content of the target multimedia. In addition, in the embodiments of the present application, it is mentioned that the media tag of the target multimedia can be generated based on the media features of the target multimedia, so as to realize that when the demand for multimedia recommendation for the target user is detected, the multimedia matched with the demand information of the target user can be directly recommended to the target user according to the media tags of different multimedia, so as to further classify the target multimedia based on the media tags in advance on the basis of that the more accurate multimedia can be recommended to the target user based on the media features representing the semantic accurately and sufficiently, so as to improve the efficiency of multimedia recommendation, and further improve the user experience.

[0091] Based on the related description of the multimedia processing method, the present application further discloses a multimedia processing device. The multimedia processing device can run a computer program (including program code) in one of the above-mentioned computer devices. The multimedia processing device can perform the multimedia processing method as shown in Figure 2 and Figure 3 , please refer to Figure 7 , the multimedia processing device can at least include an acquisition unit 701 and a processing unit 702.

[0092] The acquisition unit 701 is configured to acquire the media description information of the target multimedia, and determine the feature information of the target multimedia according to the media description information, wherein the feature information of the target multimedia is used to describe the theme semantics of the target multimedia.

[0093] The acquisition unit 701 is further configured to acquire one or more reference contents matching the theme semantics from the content library according to the theme semantics described by the feature information of the target multimedia; any one of the one or more reference contents corresponds to a corresponding feature information.

[0094] The processing unit 702 is configured to generate the media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content.

[0095] In an embodiment, the content library includes N texts, N being a positive integer greater than a preset threshold; the acquisition unit 701 can specifically perform:

[0096] For a target text included in the content library, the target text is subjected to text segmentation processing to obtain a text segmentation set corresponding to the target text, the text segmentation set including one or more segmented texts obtained by subjecting the target text to text segmentation processing; the target text is any one of the N texts.

[0097] The feature information of the target text is generated according to the text segmentation set of the target text, and a feature similarity between the feature information of the target text and the feature information of the target multimedia is calculated; the feature similarity is used to indicate a similarity degree between the text semantics of the target text and the theme semantics of the target multimedia.

[0098] When the similarity degree indicated by the feature similarity satisfies a preset similarity degree, the target text is taken as one reference content matching the theme semantics of the target multimedia, until one or more reference contents matching the theme semantics are acquired from the content library.

[0099] In another embodiment, the acquisition unit 701 can specifically perform:

[0100] Each segmented text included in the text segmentation set is subjected to feature extraction processing to obtain a text feature corresponding to each segmented text respectively; a feature set composed of the text features corresponding to each segmented text is a feature set corresponding to the text segmentation set.

[0101] The feature set corresponding to the text segmentation set is taken as the feature information of the target text.

[0102] In another embodiment, the feature information of the target text is a feature set composed of the text features of the segmented texts of the target text; the acquisition unit 701 can specifically perform:

[0103] perform feature relevance calculation on the feature information of the target multimedia and any text feature in the feature set, to obtain a first feature relevance degree between the feature information of the target multimedia and any text feature;

[0104] The first feature relevance degree between the feature information of the target multimedia and each text feature is taken as a feature similarity between the feature information of the target text and the feature information of the target multimedia.

[0105] In yet another implementation, the acquisition unit 701 can be specifically further configured to perform:

[0106] According to the first feature relevance degree between the feature information of the target multimedia and each text feature, the text features corresponding to the first feature relevance degree greater than the relevance threshold are selected from the feature set;

[0107] The number of the selected text features corresponding to the first feature relevance degree greater than the relevance threshold is determined;

[0108] When the determined number of features is greater than or equal to the preset number, it is determined that the similarity degree between the text semantics of the target text and the theme semantics of the target multimedia satisfies the preset similarity degree;

[0109] The preset number is a preset fixed number value, or the preset number is a number value calculated according to the total number of segmented texts of the target text.

[0110] In yet another implementation, the media description information includes a title of the target multimedia, and the feature information of the target multimedia includes a title feature of the title; the acquisition unit 701 can be specifically further configured to perform:

[0111] When the determined number of features is greater than or equal to the preset number, the second feature relevance degree between each text feature in the feature set and the title feature of the target multimedia is acquired, and the text features corresponding to the second feature relevance degree greater than the target threshold are selected from the feature set;

[0112] When the number of the selected text features corresponding to the second feature relevance degree greater than the target threshold is greater than or equal to a reference number, it is determined that the similarity degree between the text semantics of the target text and the theme semantics of the target multimedia satisfies the preset similarity degree.

[0113] In yet another implementation, the determination process of the feature information and the process of performing feature similarity calculation on different feature information are both calculated by the acquisition unit 701 calling a twin network model;

[0114] The twin network model comprises a first feature extraction module and a second feature extraction module. The media description information is input into the first feature extraction module to obtain feature information of the target multimedia. The N texts contained in the content library are input into the second feature extraction module to obtain feature information of each text.

[0115] The twin network model further comprises a similarity calculation module. The similarity calculation module is configured to calculate similarities between the feature information of the target multimedia and the feature information of the texts in the content library, so as to obtain one or more reference contents from the content library.

[0116] In another embodiment, the twin network model is a trained network model. The obtaining unit 701 can further be configured to train the twin network model. The specific training process comprises:

[0117] obtaining a training sample pair and a sample label of the training sample pair. The training sample pair comprises two training samples. The sample label is used to indicate a semantic similarity between the two training samples in the training sample pair;

[0118] inputting the training sample pair into the network model before training, and outputting a predicted similarity between the training sample pair;

[0119] training the twin network model according to a difference between the predicted similarity and the semantic similarity indicated by the sample label, to obtain the trained twin network model.

[0120] According to one embodiment of the present application, Figure 2 and Figure 3 the steps involved in the method shown in Figure 7 may be performed by the units in the multimedia processing apparatus shown in Figure 2 For example, Figure 7 the steps S201 and S202 shown in may be performed by the obtaining unit 701 in the multimedia processing apparatus shown in Figure 7 Step S203 can be performed by the processing unit 702 in the multimedia processing apparatus shown in Figure 3 For example, Figure 7 the steps S301 to S305 shown in may be performed by the obtaining unit 701 in the multimedia processing apparatus shown in Figure 7 Step S306 can be performed by the processing unit 702 in the multimedia processing apparatus shown in Figure 6 For example, Figure 7 the steps S601 to S602 shown in may be performed by the obtaining unit 701 in the multimedia processing apparatus shown in Figure 7 Steps S603 to S607 can be performed by the processing unit 702 in the multimedia processing apparatus shown in

[0121] According to another embodiment of the present application, Figure 7 The units in the multimedia processing apparatus shown are divided based on logical functions. The units can be combined into one or several other units respectively or all, or some of the units can be further split into multiple units with smaller functions to implement the same operation without affecting the implementation of the technical effects of the embodiments of the present application. In other embodiments of the present application, the multimedia processing apparatus can include other units. In actual applications, the functions can be assisted by other units, and can be implemented by multiple units in cooperation.

[0122] According to another embodiment of the present application, the multimedia processing apparatus shown in Figure 2 Figure 3 or Figure 6 and the multimedia processing method of the embodiments of the present application can be constructed by running a computer program (including program codes) capable of executing the steps involved in the method shown in Figure 7 The computer program can be recorded on a computer storage medium, loaded into the computer device mentioned above through the computer storage medium, and run in the computer device.

[0123] In the embodiments of the present application, in the process of obtaining the media feature of the target multimedia, the media description information of the target multimedia is obtained first, and the feature information of the target multimedia is determined according to the media description information, wherein the feature information of the target multimedia is used to describe the theme semantics of the target multimedia; then, one or more reference contents matching the theme semantics are obtained from the content library, so as to finally generate the media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content. It can be seen that in the embodiments of the present application, without relying on optical character recognition and automatic speech recognition technologies, one or more reference contents matching the theme semantics of the target multimedia are obtained from the content library, so that more text information of the media content of the target multimedia can be obtained, so that the media feature of the target multimedia generated according to the feature information of the target multimedia and the feature information of each reference content can accurately and sufficiently represent the media content of the target multimedia. Therefore, the embodiments of the present application can accurately obtain the media feature of the multimedia.

[0124] Based on the related descriptions of the method embodiments and the apparatus embodiments, the embodiments of the present application further provide a computer device, please refer to Figure 8 ​The computer device at least includes a processor 801 and a computer storage medium 802, and the processor 801 and the computer storage medium 802 of the computer device can be connected through a bus or other means.

[0125] The computer storage medium 802 mentioned above is a memory device in the computer device, used for storing computer programs and data. It can be understood that the computer storage medium 802 herein can include an internal storage medium in the computer device, and of course can also include an extended storage medium supported by the computer device. The computer storage medium 802 provides a storage space that stores an operating system of the computer device. In addition, one or more computer programs suitable for being loaded and executed by the processor 801 are also stored in the storage space, and these computer programs can be one or more program codes. It should be noted that the computer storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory; optionally, it can also be at least one storage medium located away from the aforementioned processor. The processor 801 (or CPU (Central Processing Unit, Central Processor)) is the computing core and control core of the computer device, which is suitable for implementing one or more computer programs, and is specifically suitable for loading and executing one or more computer programs to implement corresponding method processes or corresponding functions.

[0126] In one embodiment, one or more computer programs stored in the computer storage medium 802 can be loaded and executed by the processor 801 to implement the corresponding method steps in the method embodiments shown in Figure 2 、 Figure 3 and Figure 6 . In a specific implementation, one or more computer programs in the computer storage medium 802 can be loaded and executed by the processor 801 to perform the following steps:

[0127] Obtain media description information of the target multimedia, and determine feature information of the target multimedia according to the media description information, the feature information of the target multimedia being used to describe the theme semantics of the target multimedia; obtain the theme semantics described by the feature information of the target multimedia, and obtain one or more reference contents matching the theme semantics from the content library; any reference content in the one or more reference contents corresponds to a corresponding feature information; generate the media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content.

[0128] In an embodiment, the content library contains N texts, N being a positive integer greater than a preset threshold; the processor 801 can be configured to load and execute: performing text segmentation processing on a target text contained in the content library to obtain a text segmentation set corresponding to the target text, the text segmentation set containing one or more segmented texts obtained by performing text segmentation processing on the target text; wherein the target text is any one of the N texts; generating feature information of the target text according to the text segmentation set of the target text, and calculating a feature similarity between the feature information of the target text and feature information of the target multimedia; the feature similarity is used to indicate a similarity between a text semantic of the target text and a theme semantic of the target multimedia; when the similarity between the text semantic of the target text and the theme semantic of the target multimedia indicated by the feature similarity satisfies a preset similarity degree, the target text is taken as a reference content matching the theme semantic of the target multimedia, until one or more reference contents matching the theme semantic are obtained from the content library.

[0129] In another embodiment, the processor 801 can be configured to load and execute: performing feature extraction processing on each segmented text contained in the text segmentation set respectively to obtain text features respectively corresponding to each segmented text; a feature set composed of the text features respectively corresponding to each segmented text is a feature set corresponding to the text segmentation set; and taking the feature set corresponding to the text segmentation set as the feature information of the target text.

[0130] In another embodiment, the feature information of the target text is a feature set composed of text features of each segmented text of the target text; the processor 801 can be configured to load and execute: performing feature association calculation on the feature information of the target multimedia and any text feature in the feature set to obtain a first feature association degree between the feature information of the target multimedia and any text feature; and taking the first feature association degree between the feature information of the target multimedia and each text feature as the feature similarity between the feature information of the target text and the feature information of the target multimedia.

[0131] In another embodiment, the processor 801 can be configured to load and execute: selecting, according to the first feature association degree between the feature information of the target multimedia and each text feature, a text feature corresponding to a first feature association degree greater than an association degree threshold from the feature set; determining a feature quantity of the text feature corresponding to the first feature association degree greater than the association degree threshold selected; and determining that the similarity between the text semantic of the target text and the theme semantic of the target multimedia satisfies the preset similarity degree when the determined feature quantity is greater than or equal to a preset quantity; wherein the preset quantity is a preset fixed quantity value, or the preset quantity is a quantity value calculated according to a total number of segmented texts of the target text.

[0132] In yet another implementation, the media description information includes a title of the target multimedia, and the feature information of the target multimedia includes title features of the title; the processor 801 can be configured to load and execute: when the determined number of features is greater than or equal to the preset number, obtaining a second feature correlation degree between each text feature in the feature set and the title features of the target multimedia, and selecting, from the feature set, a text feature corresponding to a second feature correlation degree greater than a target threshold; and when the number of selected text features corresponding to a second feature correlation degree greater than the target threshold is greater than or equal to a reference number, determining that the similarity between the text semantics of the target text and the theme semantics of the target multimedia satisfies the preset similarity.

[0133] In yet another implementation, the determination of the feature information and the feature similarity calculation of different feature information are both obtained by the processor 801 calling a twin network model.

[0134] The twin network model includes a first feature extraction module and a second feature extraction module, and the processor 801 can be configured to load and execute: inputting the media description information into the first feature extraction module to obtain the feature information of the target multimedia, and inputting the N texts included in the content library into the second feature extraction module to obtain the feature information of each text.

[0135] The twin network model further includes a similarity calculation module, and the processor 801 can be configured to load and execute: performing similarity calculation on the feature information of the target multimedia and the feature information of the texts in the content library by the similarity calculation module, to obtain one or more reference contents from the content library.

[0136] In yet another implementation, the twin network model is a trained network model, and the training process of the twin network model by the processor 801 includes: obtaining a training sample pair and a sample label of the training sample pair, the training sample pair including two training samples, and the sample label being used to indicate the semantic similarity between the two training samples included in the training sample pair; inputting the training sample pair into a network model before training, to output a predicted similarity between the training sample pair; and training the twin network model according to the difference between the predicted similarity and the semantic similarity indicated by the sample label, to obtain the trained twin network model.

[0137] In the embodiment of the present application, in the process of acquiring the media feature of the target multimedia, the media description information of the target multimedia can be acquired first, and the feature information of the target multimedia is determined according to the media description information, wherein the feature information of the target multimedia is used to describe the theme semantics of the target multimedia; then, one or more reference contents matching the theme semantics are acquired from the content library, so as to finally generate the media feature of the target multimedia according to the feature information of the target multimedia and the feature information of each reference content. It can be seen that in the embodiment of the present application, the optical character recognition and automatic speech recognition technologies are not relied on, but one or more reference contents matching the theme semantics of the target multimedia are acquired from the content library, so that more text information of the media content of the target multimedia can be acquired, thereby the media feature of the target multimedia generated according to the feature information of the target multimedia and the feature information of each reference content can accurately and sufficiently represent the media content of the target multimedia. Therefore, the embodiment of the present application can accurately acquire the media feature of the multimedia.

[0138] The present application further provides a computer storage medium, which stores one or more computer programs corresponding to the multimedia processing method described above. When one or more processors load and execute the one or more computer programs, the description of the multimedia processing method in the embodiment can be implemented, which will not be described herein again. The beneficial effects of the same method will not be described herein again. It can be understood that the computer program can be deployed on one or more devices capable of communicating with each other to execute.

[0139] It should be noted that according to one aspect of the present application, a computer program product or a computer program is also provided, which includes a computer program stored in a computer storage medium. The processor in the computer device reads the computer program from the computer storage medium, and then executes the computer program, so that the computer device can execute the multimedia processing method described above. Figure 2 and Figure 3 the multimedia processing method embodiments shown in the various optional manners.

[0140] It can be understood that in the specific embodiments of the present application, some embodiments relate to the data related to the user information such as the demand information; therefore when the above embodiments of the present application are applied to specific products or technologies, the user permission or consent needs to be obtained, and the collection, use and processing of the related data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0141] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiment of the data processing method. The computer storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.

[0142] The above only discloses part of the embodiments of the present application, and of course cannot limit the scope of the rights of the present application. Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiments can be implemented, and equivalent changes made according to the claims of the present application still fall within the scope of the present application.

Claims

1. A multimedia processing method, characterized in that, include: The media description information of the target multimedia is obtained, and the feature information of the target multimedia is determined based on the media description information. The feature information of the target multimedia is used to describe the thematic semantics of the target multimedia. The process involves obtaining the topic semantics described by the feature information of the target multimedia, and retrieving one or more reference contents from a content library that match the topic semantics. Each of the one or more reference contents corresponds to a specific feature information. The content library contains N texts, where N is a positive integer greater than a preset threshold. For the target texts contained in the content library, text segmentation is performed on the target texts to obtain a text segmentation set corresponding to the target texts. Feature information of the target texts is generated based on the text segmentation set of the target texts, and the feature similarity between the feature information of the target texts and the feature information of the target multimedia is calculated. The feature information of the target texts is derived from the feature information of the target multimedia. The feature set is composed of the text features of each segmented text; based on the first feature correlation degree between the feature information of the target multimedia and each text feature, text features with a first feature correlation degree greater than a correlation degree threshold are selected from the feature set; the number of selected text features with a first feature correlation degree greater than the correlation degree threshold is determined; when the number of determined features is greater than or equal to a preset number, the similarity between the text semantics of the target text and the topic semantics of the target multimedia is determined to meet a preset similarity degree, and the target text is used as a reference content that matches the topic semantics of the target multimedia, until one or more reference contents that match the topic semantics are obtained from the content library; Based on the feature information of the target multimedia and the feature information of each reference content, media features of the target multimedia are generated; the media features integrate the semantics of the feature information and the semantics of the reference content.

2. The method according to claim 1, characterized in that, The text segmentation set contains one or more segmented texts obtained by performing text segmentation processing on the target text; wherein, the target text is any one of the N texts; The feature similarity is used to indicate the degree of similarity between the textual semantics of the target text and the thematic semantics of the target multimedia.

3. The method according to claim 2, characterized in that, The step of generating feature information of the target text based on the text segmentation set of the target text includes: Each segmented text in the text segmentation set is subjected to feature extraction processing to obtain the text features corresponding to each segmented text; the feature set composed of the text features corresponding to each segmented text is the feature set corresponding to the text segmentation set. The feature set corresponding to the text segmentation set is used as the feature information of the target text.

4. The method according to claim 2, characterized in that, The calculation of the feature similarity between the feature information of the target text and the feature information of the target multimedia includes: The feature correlation between the feature information of the target multimedia and any text feature in the feature set is calculated to obtain the first feature correlation degree between the feature information of the target multimedia and any text feature. The first feature correlation degree between the feature information of the target multimedia and each text feature is used as the feature similarity between the feature information of the target text and the feature information of the target multimedia.

5. The method according to claim 4, characterized in that, in, The preset quantity is a preset fixed quantity value, or the preset quantity is a quantity value calculated based on the total number of segmented texts of the target text.

6. The method according to claim 5, characterized in that, The media description information includes the title of the target multimedia, and the feature information of the target multimedia includes the title features of the title; When the number of determined features is greater than or equal to a preset number, determining that the similarity between the textual semantics of the target text and the topical semantics of the target multimedia satisfies a preset similarity includes: When the number of determined features is greater than or equal to the preset number, the second feature correlation degree between each text feature in the feature set and the title feature of the target multimedia is obtained, and text features with a second feature correlation degree greater than the target threshold are selected from the feature set. When the number of selected corresponding second feature correlation degrees greater than the target threshold is greater than or equal to the reference number, it is determined that the similarity between the textual semantics of the target text and the topical semantics of the target multimedia satisfies the preset similarity.

7. The method according to any one of claims 1 to 6, characterized in that, The process of determining feature information, as well as the process of calculating feature similarity for different feature information, are both obtained by calling the Siamese network model. The Siamese network model includes a first feature extraction module and a second feature extraction module. The media description information is input into the first feature extraction module to obtain the feature information of the target multimedia. The N texts contained in the content library are input into the second feature extraction module to obtain the feature information of each text. The twin network model also includes a similarity calculation module, which is used to calculate the similarity between the feature information of the target multimedia and the feature information of the text in the content library, so as to obtain one or more reference contents from the content library.

8. The method according to claim 7, characterized in that, The Siamese network model is a trained network model, and the training process for the Siamese network model includes: Obtain training sample pairs and sample labels for the training sample pairs, wherein each training sample pair contains two training samples and the sample labels are used to indicate the semantic similarity between the two training samples contained in the training sample pair; The training sample pairs are input into the network model before training, and the predicted similarity between the training sample pairs is output. The Siamese network model is trained based on the difference between the predicted similarity and the semantic similarity indicated by the sample label, resulting in a trained Siamese network model.

9. The method according to claim 1, characterized in that, The method further includes: A media tag for the target multimedia is generated based on the media characteristics of the target multimedia, and the media tag is used to indicate the media type of the target multimedia; When a need for multimedia recommendations for a target user is detected, the user's need information is obtained, and multimedia that matches the user's need information is recommended based on the media tags of different multimedia.

10. A multimedia processing device, characterized in that, The multimedia processing device includes an acquisition unit and a processing unit, wherein: The acquisition unit is used to acquire media description information of the target multimedia and determine the feature information of the target multimedia based on the media description information. The feature information of the target multimedia is used to describe the thematic semantics of the target multimedia. The acquisition unit is further configured to acquire the topic semantics described by the feature information of the target multimedia, and acquire one or more reference contents that match the topic semantics from the content library; each of the one or more reference contents corresponds to a corresponding feature information; the content library contains N texts, where N is a positive integer greater than a preset threshold; for the target text contained in the content library, the target text is processed by text segmentation to obtain a text segmentation set corresponding to the target text; feature information of the target text is generated based on the text segmentation set of the target text, and feature similarity between the feature information of the target text and the feature information of the target multimedia is calculated; wherein, the feature information of the target text is composed of... The target text is composed of text features of each segmented text. Based on the first feature correlation degree between the feature information of the target multimedia and each text feature, text features with a first feature correlation degree greater than a correlation degree threshold are selected from the feature set. The number of selected text features with a first feature correlation degree greater than the correlation degree threshold is determined. When the number of determined features is greater than or equal to a preset number, the similarity between the text semantics of the target text and the topic semantics of the target multimedia is determined to meet a preset similarity degree, and the target text is used as a reference content that matches the topic semantics of the target multimedia, until one or more reference contents that match the topic semantics are obtained from the content library. The processing unit is configured to generate media features of the target multimedia based on the feature information of the target multimedia and the feature information of each reference content; the media features integrate the semantics of the feature information and the semantics of the reference content.

11. A computer device, characterized in that, include: A processor, the processor being adapted to implement one or more computer programs; A computer storage medium storing one or more computer programs, said one or more computer programs being adapted to be loaded by the processor and executed as described in any one of claims 1-9.

12. A computer storage medium, characterized in that, The computer storage medium stores one or more computer programs, which are adapted to be loaded by a processor and executed by the multimedia processing method as described in any one of claims 1-9.

13. A computer program product, characterized in that, The computer program product includes a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Theme type mining method and device, equipment and storage medium

    CN113392315A

  • System and method for providing supplemental information relevant to selected content in media

    US8799401B1