Multimedia retrieval method, device, computing device and storage medium

By performing fine-grained hierarchical semantic attribute extraction and matching semantic labels with multimedia labels on the search text, the problems of semantic confusion and low recognition in the existing search solutions are solved, and higher retrieval accuracy and recognition are achieved.

CN116304120BActive Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211434753.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2025-05-06
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

The existing multimedia video retrieval scheme cannot effectively distinguish the relationship between attributes and entity objects, resulting in confusion in search semantics and lack of association and sequential relationships in keyword information, which may lead to incorrect recognition or missed recognition.

Method used

By receiving the search text, semantic analysis is performed to obtain sub-attributes and correlations, input it into the pre-trained semantic matching model, obtain the label of the sub-attributes, and determine the similarity value between the label and the multimedia subclass label to match the multimedia subclass of the search text.

Benefits of technology

The relationship recognition problem of multiple different entity objects in the search text is solved, the accuracy and recognition of the search is improved, and the relationship between the semantics of the search text and multimedia tags can be better explored.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116304120B_ABST
    Figure CN116304120B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimedia retrieval method, device, computing equipment and storage medium, wherein the method comprises: receiving a retrieval text, performing semantic analysis on the retrieval text, obtaining at least one sub-attribute and the correlation between the sub-attributes; inputting each sub-attribute and the correlation into a pre-trained semantic matching model to obtain a label of each sub-attribute; determining the similarity value between the label of each sub-attribute and the pre-obtained multimedia sub-class label, and determining the multimedia sub-class matching with the retrieval text according to the size of each similarity value. The present invention realizes a multi-level matching mechanism between the semantic label of the retrieval text and the multimedia label by means of hierarchical semantic matching, can better mine the relationship between the semantics of the retrieval text and the multimedia label, and has a higher recognition degree and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimedia retrieval method, device, computing equipment and storage medium. Background Art

[0002] With the popularization of 5G and the reduction of network bandwidth costs, video data has gradually become one of the mainstream data in the fields of social networking, security, etc. Due to the explosive growth in the number of videos, the diversity of multimedia video content and the complexity of data structure, how to quickly and effectively retrieve the videos that users want from the video library has become a difficult problem.

[0003] There are currently two retrieval schemes: one is based on user keywords, and the other is based on preset optional features. However, when multiple different entity objects appear in the retrieval text (i.e., the question sentences used by the user when searching), the above retrieval schemes cannot distinguish the relationship between attributes and entity objects, which will lead to confusion between search semantics; and, since the current keyword information has no association relationship and no order relationship, it may lead to problems such as misidentification or missed recognition. Summary of the invention

[0004] In view of the above problems, the present invention is proposed to provide a multimedia retrieval method, apparatus, computing device and storage medium that overcome the above problems or at least partially solve the above problems.

[0005] According to one aspect of the present invention, a multimedia retrieval method is provided, the method comprising:

[0006] receiving a search text, performing semantic analysis on the search text, and obtaining at least one sub-attribute and a correlation between the sub-attributes;

[0007] Inputting each of the sub-attributes and the correlation into a pre-trained semantic matching model to obtain a label for each of the sub-attributes;

[0008] Determine the similarity value between the label of each of the sub-attributes and the pre-obtained multimedia sub-category label, and determine the multimedia sub-category matching the search text according to the magnitude of each of the similarity values.

[0009] Optionally, the receiving of the search text, performing semantic analysis on the search text, and obtaining at least one sub-attribute and the correlation between the sub-attributes includes:

[0010] Analyze the search text to obtain the word segmentation and part of speech of the search text;

[0011] Performing dependency syntactic analysis on the search text in combination with the segmented words and their parts of speech to obtain dependency relations between the segmented words;

[0012] According to the segmented words, their parts of speech and dependency relationships, sub-attributes of the search text and correlations between the sub-attributes are determined.

[0013] Optionally, the determining the sub-attributes of the search text and the correlation between the sub-attributes according to the segmentation words, their parts of speech, and the dependency relationship includes:

[0014] Encode the segmented words and their parts of speech to obtain a segmented word digital code, input the segmented word digital code into an encoder for conversion to obtain a segmented word vector; perform syntactic analysis on the search text to obtain a syntactic feature vector;

[0015] The word segmentation vector and the syntactic feature vector are concatenated and fused to obtain a first latent vector;

[0016] Processing the first latent vector through a multilayer perceptron to obtain a second latent vector;

[0017] A dependency matrix is ​​constructed according to the search text and the dependency, a product value of each position in the dependency matrix is ​​determined according to the dependency matrix, the dependency initialization vector and the second latent vector, and a correlation between the sub-attributes is determined according to the size of the product value and a first loss function.

[0018] Optionally, the first loss function includes a first cross entropy loss function and a KL divergence loss function;

[0019] Among them, the first cross entropy loss function takes the normalized exponential function and the actual relationship value as independent variables, and the KL divergence loss function takes the average value of the first latent vector and the average value of the target sub-attribute latent vector as independent variables.

[0020] Optionally, the training step of the semantic matching model includes:

[0021] Extracting sub-attributes from the retrieved text sample according to the semantic analysis result, and generating positive sample vectors, retrieved text vectors and negative sample vectors according to the sub-attribute conversion;

[0022] The positive sample vector, the search text vector and the negative sample vector are respectively subjected to word segmentation and feature encoding by an encoder, and then respectively input into respective bidirectional long short-term memory network models to obtain a positive ratio encoding vector, a text ratio encoding vector and a negative ratio encoding vector;

[0023] After the positive ratio coding vector, the text ratio coding vector and the negative ratio coding vector are processed by a normalized exponential function and a second loss function, a label of a sub-attribute in the retrieved text sample is obtained;

[0024] The labels of the sub-attributes are compared with the manual labels to obtain the corrected semantic matching model.

[0025] Optionally, the second loss function includes a second cross entropy loss function and a binary classification loss function;

[0026] Among them, the second cross entropy loss function takes the normalized exponential function and the actual label value as independent variables, and the binary classification loss function takes the similarity between the positive ratio encoding vector, the text ratio encoding vector and the negative ratio encoding vector as independent variables.

[0027] Optionally, the step of acquiring the multimedia subcategory label includes:

[0028] Extract key frames from multimedia, perform target detection on the key frames, analyze them according to preset hierarchical categories, and obtain labels of corresponding hierarchical categories.

[0029] According to another aspect of the present invention, a multimedia retrieval device is provided, the device comprising:

[0030] A semantic analysis module, adapted to receive a search text, perform semantic analysis on the search text, and obtain at least one sub-attribute and a correlation between the sub-attributes;

[0031] A label acquisition module, adapted to input each of the sub-attributes and the correlation into a pre-trained semantic matching model to obtain a label for each of the sub-attributes;

[0032] The search and matching module is adapted to determine the similarity value between the label of each sub-attribute and the pre-obtained multimedia sub-category label, and determine the multimedia sub-category matching the search text according to the size of each similarity value.

[0033] According to another aspect of the present invention, there is provided a computing device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus;

[0034] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute operations corresponding to the above-mentioned multimedia retrieval method.

[0035] According to another aspect of the present invention, a computer storage medium is provided, wherein the storage medium stores at least one executable instruction, and the executable instruction enables a processor to execute operations corresponding to the above-mentioned multimedia retrieval method.

[0036] According to the multimedia retrieval solution of the present invention, the problem that the existing keyword extraction method may cause confusion between search semantics is solved by extracting fine-grained hierarchical semantic attributes of the retrieval text and matching the retrieval text semantic tags with multimedia tags; thereby, the relationship between the retrieval text semantics and multimedia tags can be better explored, and multiple target retrieval can be achieved using one retrieval text, with higher recognition and accuracy.

[0037] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0039] Figure 1 A flowchart of a multimedia retrieval method provided by an embodiment of the present invention is shown;

[0040] Figure 2 A schematic diagram of a multimedia retrieval scenario provided by an embodiment of the present invention is shown;

[0041] Figure 3 A flowchart of semantic analysis of search text provided by an embodiment of the present invention is shown;

[0042] Figure 4 A flowchart of extracting fine-grained hierarchical semantic attributes of search text provided by an embodiment of the present invention is shown;

[0043] Figure 5 A schematic diagram of a process for unified identification of multiple subjects, multiple sub-attributes, and coverage attributes provided by an embodiment of the present invention is shown;

[0044] Figure 6 A flowchart of searching for text semantic matching provided by an embodiment of the present invention is shown;

[0045] Figure 7 A schematic diagram showing the structure of a multimedia retrieval device provided by an embodiment of the present invention is shown;

[0046] Figure 8 A flow chart of performing video retrieval using a device provided by an embodiment of the present invention is shown;

[0047] Fig. 9A schematic diagram of the structure of a computing device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0048] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.

[0049] Figure 1 The flowchart of the multimedia retrieval method embodiment of the present invention is shown, and the method is applied to a computing device. The computing device includes various computers, smart terminals, and tablet computers. Figure 1 As shown, the method comprises the following steps:

[0050] Step 110: receiving a search text, performing semantic analysis on the search text, and obtaining at least one sub-attribute and correlations between the sub-attributes.

[0051] Combination Figure 2 As shown in the schematic diagram of the multimedia retrieval scenario, the embodiment of the present invention can be used to retrieve corresponding multimedia resources by receiving search questions and other terms input by the user, wherein the multimedia includes video, sound, image collection and animation, etc.

[0052] according to Figure 2 In the multimedia retrieval scenario shown, the user enters the terms, questions, etc. that he wants to search for in the search box to obtain the corresponding retrieval text, and obtains the sub-attributes through the semantic analysis method. Among them, the sub-attribute is not just a single word, but a fragment sequence, which has better recognition ability than a single participle or a word with a state (such as a negative sentence: "not wearing glasses"). In addition, step 110 can also identify the correlation between sub-attributes, such as parallel relationships and subordinate relationships, and can also identify key parent attributes from sub-attributes, so that different retrieval requirements of multiple targets can be identified from a sentence of the retrieval text.

[0053] Step 120: Input each of the sub-attributes and the correlation into a pre-trained semantic matching model to obtain a label for each of the sub-attributes.

[0054] In order to retrieve the corresponding multimedia resources according to the search text, it is necessary to match the sub-attribute information obtained from the search text with the label of the multimedia resource. This step uses the pre-trained semantic matching model to obtain the sub-attribute label with the highest matching degree to improve the matching accuracy.

[0055] Step 130: Determine the similarity value between the label of each of the sub-attributes and the pre-obtained multimedia sub-category label, and determine the multimedia sub-category matching the search text according to the size of each of the similarity values.

[0056] Continue to see Figure 2 As shown, multimedia data such as videos are saved in a database after corresponding tags are extracted through analysis, wherein the multimedia subclass is used to represent the hierarchical structure of multimedia and can represent any category thereof. During retrieval, a match is performed between the tags in the retrieval text and the tags of the corresponding multimedia category, and the corresponding multimedia is determined or located based on the matching results.

[0057] It should be noted that step 120 and step 130 can be combined, for example, the multimedia resources matching the user's search can be obtained by combining a semantic matching model with a similarity value analysis. Preferably, by sorting the similarity values, multimedia that exceeds a preset threshold can be determined as meeting the user's needs.

[0058] In summary, according to Figure 1 The embodiment shown solves the problem that the existing keyword extraction method may cause confusion between search semantics by performing fine-grained hierarchical semantic attribute extraction on the search text and a matching mechanism between the search text semantic tags and multimedia tags; thereby, it is able to better explore the relationship between the search text semantics and multimedia tags, and use one search text to achieve retrieval of multiple targets with higher recognition and accuracy.

[0059] In one or some embodiments, receiving a search text, performing semantic analysis on the search text, and obtaining at least one sub-attribute and the correlation between the sub-attributes includes:

[0060] Analyze the search text to obtain the word segmentation and part of speech of the search text;

[0061] Performing dependency syntactic analysis on the search text in combination with the segmented words and their parts of speech to obtain dependency relations between the segmented words;

[0062] According to the segmented words, their parts of speech and dependency relationships, sub-attributes of the search text and correlations between the sub-attributes are determined.

[0063] In order to make the expression of the search text more accurate, the search text is also preprocessed, including uppercase and lowercase conversion, traditional Chinese and simplified Chinese conversion, and wrong character correction.

[0064] In one or some embodiments, determining the sub-attributes of the search text and the correlation between the sub-attributes according to the segmentation words, their parts of speech, and dependency relationships includes:

[0065] Encode the segmented words and their parts of speech to obtain a segmented word digital code, input the segmented word digital code into an encoder for conversion to obtain a segmented word vector; perform syntactic analysis on the search text to obtain a syntactic feature vector;

[0066] The word segmentation vector and the syntactic feature vector are concatenated and fused to obtain a first latent vector;

[0067] Processing the first latent vector through a multilayer perceptron to obtain a second latent vector;

[0068] A dependency matrix is ​​constructed according to the search text and the dependency, a product value of each position in the dependency matrix is ​​determined according to the dependency matrix, the dependency initialization vector and the second latent vector, and a correlation between the sub-attributes is determined according to the size of the product value and a first loss function.

[0069] In a specific example, combining Figure 3 The flowchart of semantic analysis of the search text shown in FIG. 1 further explains the semantic analysis process of the search text. The overall process of extracting the core sub-attributes based on dependency relationships is as follows:

[0070] (1) Text preprocessing: mainly preprocessing the search text. The operations that can be performed include: case conversion, traditional Chinese and simplified Chinese conversion, and spelling error correction.

[0071] (2) Text lexical analysis: Perform word segmentation and part-of-speech analysis on the search text. For example, "a man wearing glasses" can be segmented into "a / m wearing / v glasses / n / u man / n". In this example, the part-of-speech classification standards of ansj, CTB, and PKU can be used. Of course, other word segmentation and part-of-speech tagging standards can also be used.

[0072] (3) Text syntactic analysis: Combine the word segmentation results and part of speech to perform dependency syntactic analysis on the search text. Through dependency syntactic analysis, the analysis of the associated attributes between words is realized. The dependency relationship attributes used in this example include but are not limited to: ATT (attributive), DE (of), SBV (subject), VOB (object), COO (general parallel), COS (shared parallel), ADV (adverbial), HED (core), etc.; for example, "a man wearing glasses" can be parsed into the following attribute relationship table 1 through word segmentation, part of speech analysis, and dependency syntactic analysis.

[0073] Table 1 Attribute splitting and its correlation

[0074] sequence Participle Part of Speech Dependency number Dependencies 1 one m 5 NUM 2 Wear v 4 DE 3 Glasses n 2 VOB 4 of u 5 ATT 5 man n 0 HED

[0075] Preferably, the dependency relationship in this example can adopt existing standards such as PKU Multi-view Chinese Treebank.

[0076] (4) Fine-grained hierarchical semantic attribute extraction: Fine-grained attribute extraction can be achieved, and the dependency relationship between attributes and core entities can be identified. The specific process is as follows: Figure 4 The process of extracting fine-grained hierarchical semantic attributes from retrieved text is shown in FIG.

[0077] Compared with the attribute extraction method of the existing dependency extraction model, the method used in this example has higher generalization ability and accuracy. The overall solution steps are as follows:

[0078] Step 1: Convert the result of lexical analysis into digital code X(x1,x2,x3,...,xn) in combination with the dictionary and input it into the encoder.

[0079] Step 2: Through the encoder (the encoder can choose common encoders such as word2vec, BERT, RoBERTa, etc.), the digital code is converted into a vector feature H(x)(h1,h2,h3,...,hn)=Encoder(X);

[0080] Step 3: Concatenate the syntactic features D(x)(d1, d2, d3, ..., dn) obtained by syntactic analysis with H obtained by the encoder to obtain a new first latent vector feature H'(x)(h1⊕d1, h2⊕d2, ....hn⊕dn);

[0081] Step 4: Pass the first latent vector feature through a multi-layer perceptron layer (MLP) to obtain the second latent vector G(H'(x)) of the new dimension, where G is the mapping function representation of the MLP layer;

[0082] Step 5: Combine Figure 5 The flowchart of unified identification of multiple subjects, multiple sub-attributes, and covering attributes shown in the figure first constructs the corresponding dependency matrix for the search text, and sets six relationships between different positions: H (indicating the head), E (indicating the tail), L (indicating the connection relationship), C (indicating the subordinate relationship), O (indicating the parallel relationship), and N (no relationship). Each relationship is initialized separately. q d q , W k d k , used to score the relationship between positions i and j in the matrix, where q i =W q G(H'(x i ))+d q , k j =Wk G(H'(x j )), the final score is S(i,j)=q i Tk j The above method is used to score the upper triangular matrix. Finally, the relationship with the highest score is determined by the Softmax function.

[0083] See also Figure 5 As shown, by combining sequence tagging and dependency parsing, fine-grained attribute recognition can be achieved, and the correlation between attributes can be identified. For example, in a search question, different corresponding attributes of multiple entities (men, children) can be identified, and then accurate retrieval of multiple targets can be achieved based on a search text.

[0084] The decoding method of the search text is: first find the H attribute character as the first character, then find the L attribute character in the same row, then find the next L or E attribute character in the column where the next L character is located, and end with the E attribute character as the segment. In addition, there are also subordinate relationships (C) and parallel relationships (O) between the E character attributes (the attribute segments they represent) and the E character attributes (the attribute segments they represent).

[0085] In one or some embodiments, the first loss function includes a first cross entropy loss function and a KL divergence loss function;

[0086] Among them, the first cross entropy loss function takes the normalized exponential function and the actual relationship value as independent variables, and the KL divergence loss function takes the average value of the first latent vector and the average value of the target sub-attribute latent vector as independent variables.

[0087] In order to make the above model converge, the softmax function can be used to convert the output results of the neural network, express the output results in the form of probability, and then use the loss function to calculate the gap with the actual classification, so as to optimize the model parameters through iteration and other methods.

[0088] Specifically, we can set loss = CrossEntroyLoss (Softmax (QK), Y) to obtain the result of the correlation between sub-attributes. In addition, considering that the overlay of sub-attributes cannot change the original semantics, it is necessary to constrain the recognition effect of sub-attributes. Loss sub-attribute constraint = KL (Avg (H' (X)), Avg (H' (X 子属性1 ,X 子属性2 ,...,X 子属性N ))), where KL represents KL distance, Avg represents vector averaging, and since the input of the sub-attribute is unknown at the beginning, the sequence of decoded predicted sub-attributes is used.

[0089] The final model loss can be expressed as follows (where α is the weight factor):

[0090] Loss=α·CrossEntroyLoss(Softmax(QK),Y)+(1-α)·KL(Avg(H'(X)),Avg(H'(X 子属性1 ,X 子属性2 ,...,X 子属性N )))

[0091] In short, through the semantic analysis module, we can get the sub-attribute list of the user's question and the correlation between the sub-attributes.

[0092] In one or some embodiments, the step of training the semantic matching model includes:

[0093] Extracting sub-attributes from the retrieved text sample according to the semantic analysis result, and generating positive sample vectors, retrieved text vectors and negative sample vectors according to the sub-attribute conversion;

[0094] The positive sample vector, the search text vector and the negative sample vector are respectively subjected to word segmentation and feature encoding by an encoder, and then respectively input into respective bidirectional long short-term memory network models to obtain a positive ratio encoding vector, a text ratio encoding vector and a negative ratio encoding vector;

[0095] After the positive ratio coding vector, the text ratio coding vector and the negative ratio coding vector are processed by a normalized exponential function and a second loss function, a label of a sub-attribute in the retrieved text sample is obtained;

[0096] The labels of the sub-attributes are compared with the manual labels to obtain the corrected semantic matching model.

[0097] Furthermore, in one or some optional embodiments, the second loss function includes a second cross entropy loss function and a binary classification loss function;

[0098] Among them, the second cross entropy loss function takes the normalized exponential function and the actual label value as independent variables, and the binary classification loss function takes the similarity between the positive ratio encoding vector, the text ratio encoding vector and the negative ratio encoding vector as independent variables.

[0099] Specific, combined Figure 6The retrieval text semantic matching flowchart shown in the figure is used to achieve the matching and retrieval of semantic analysis results and video analysis results. The process mainly includes two parts: model training and decoding retrieval. The training process is mainly used to train the adapted decoder, which can achieve the alignment of the decoding of the video analysis results and the decoding of the semantic analysis results. In the decoding retrieval process, it is only necessary to match and retrieve the labels (decoding features) of the semantic analysis with the labels (decoding features) of the pre-stored video analysis results.

[0100] The video analysis features are mainly three-level features: attribute category-attribute subcategory-attribute entity. If the decoding is performed directly according to the conventional method, and then matched by similarity, first, there are too many matching categories (generally dozens of second-level categories and thousands of third-level categories), and the screening information cannot be fully utilized; and the text similarity between attribute entities is too high, and the conventional semantics are too close (for example: "pedestrian structure-whether to carry an umbrella-yes" and "pedestrian structure-whether to carry an umbrella-no"), which leads to a large matching error. Therefore, in this embodiment, by combining classification and contrast learning, the similarity optimization and strengthening within the sub-classification is realized, and a better hierarchical matching effect is achieved.

[0101] See also Figure 6 As shown, the semantic matching process in this embodiment is as follows:

[0102] First, during the training process, for the query (that is, the user's question (retrieval text)), a large number of sub-attributes are extracted according to the semantic analysis module. For example, "a man with an umbrella and glasses" extracts "a man with an umbrella", "a man with glasses" and "a man", three sub-attribute fragments (if a sub-attribute has a parent structure, it needs to be combined with the parent node as a fragment). For example, the video structured positive and negative samples corresponding to "a man" are: "human structured-gender-male" and "human structured-gender-female". The video structured positive and negative samples corresponding to "a man with an umbrella" are: "pedestrian structured-whether to carry an umbrella-yes" and "pedestrian structured-whether to carry an umbrella-no". The negative sample can be one or more, and can also be sampled. For example, the positive sample of "a man wearing long clothes" is "pedestrian structured-top-long sleeves", and its negative sample can be "pedestrian structured-top-short sleeves" or "pedestrian structured-top-vest", etc.

[0103] The positive and negative samples and the attribute fragments of the text are segmented and encoded as input X, X 正 , X 负 . And input the digitized encoding information into the neural network encoder (such as BERT, RoBERTa, etc.). Encoder(X), Encoder(X positive), Encoder(X negative) are obtained respectively.

[0104] The encoded information is passed through the BiLSTM network respectively to obtain the semantic information BiLSTM (Encoder (X)) with further fusion of information.

[0105] The sub-attributes of the question, positive and negative samples should belong to the same secondary sub-class label. To simplify the representation, F(X)=Softmax(BiLSTM(Encoder(X)) is defined here, and the loss function is defined as follows:

[0106] Loss1=CrossEntryLoss(F(X),Y))+CrossEntryLoss(F(X positive),Y))+CrossEntryLoss(F(X negative),Y))

[0107] In addition, the sub-attribute labels of the question should be able to distinguish the difference between positive and negative samples by similarity. At the same time, considering that positive and negative samples belong to the same secondary classification, the corresponding discrimination can be strengthened. The loss function can be defined as follows:

[0108] Loss2=α·BinaryLoss(S(BiLSTM(Encoder(X),BiLSTM(Encoder(X positive)),1)+(1-α)·BinaryLoss(S(BiLSTM(Encoder(X),BiLSTM(Encoder(X negative))),0)

[0109] Here, S can be a similarity function of various text representations, such as cosine distance, etc. The closer to 0, the less similar it is, and the closer to 1, the more similar it is.

[0110] The final training module can express the loss function as: Loss = β·Loss1+(1-β)·Loss2.

[0111] In one or some embodiments, the step of obtaining the multimedia subcategory tag includes:

[0112] Extract key frames from multimedia, perform target detection on the key frames, analyze them according to preset hierarchical categories, and obtain labels of corresponding hierarchical categories.

[0113] Specifically, taking video as an example, the analysis of its content generally adopts the method of extracting key frames from the video, performing target detection on the key frames, general image recognition and other analysis methods to achieve labeling of key content in the key frames.

[0114] In the embodiment of the present invention, multimedia can be analyzed according to the tags of fixed hierarchical categories that have been set. The fixed hierarchical categories include: pedestrian structured-age group-elderly, pedestrian structured-top style-short sleeve, vehicle structured-general vehicle type-passenger bus, etc. The same key frame can identify multiple targets, and there can be multiple structured information for different targets.

[0115] In summary, the multimedia search execution process of the embodiment of the present invention is roughly divided into the following stages:

[0116] (1) During the reasoning and retrieval process, the semantic parsing module needs to be performed in real time to parse the user's search text into corresponding semantic sub-attributes.

[0117] (2) The parsed sub-attributes are bi-encoded using the pre-trained encoder model (i.e., BiLSTM (Encoder (X)) is obtained).

[0118] (3) The ratio encoding is input into a semantic matching model to predict the labels corresponding to the secondary subclasses of the retrieved text (i.e., the classification results obtained by Softmax).

[0119] (4) and all encoding results of the video subclass (i.e., the pre-obtained BiLSTM (Encoder (X 子类1 ))、BiLSTM(Encoder(X 子类2 ))、....、BiLSTM(Encoder(X 子类n ))) Perform similarity matching, and the matching method still uses the S function. Preferably, the class with a threshold greater than 0.5 and the most similar class is selected as the matching class, so as to locate the video resource.

[0120] A user search question (text) can have multiple parent attributes of the same level, and a parent attribute can have multiple child attributes. Through the above steps, a search question can be matched with multiple video subcategories of the same level.

[0121] Figure 7 FIG. 2 shows a schematic diagram of the structure of an embodiment of a multimedia retrieval device according to the present invention. Figure 7 As shown, the device 700 includes:

[0122] The semantic analysis module 710 is adapted to receive a search text, perform semantic analysis on the search text, and obtain at least one sub-attribute and a correlation between the sub-attributes;

[0123] A label acquisition module 720, adapted to input each of the sub-attributes and the correlation into a pre-trained semantic matching model to obtain a label of each of the sub-attributes;

[0124] The search matching module 730 is adapted to determine the similarity value between the label of each of the sub-attributes and the pre-obtained multimedia sub-category label, and determine the multimedia sub-category matching the search text according to the size of each of the similarity values.

[0125] For details, see Figure 8 The connection structure diagram of the video retrieval module shown in the figure shows that in specific operations, it is first necessary to analyze the multimedia (taking video as an example) to obtain the label classification information of the multimedia; then analyze the search text to obtain the corresponding sub-attributes and related relationships; finally, match the search text and multimedia resources through the semantic matching model to obtain the best matching result, so as to locate the multimedia resources desired by the user according to the search text.

[0126] In one or some embodiments, the semantic analysis module 710 is further adapted to:

[0127] Analyze the search text to obtain the word segmentation and part of speech of the search text;

[0128] Performing dependency syntactic analysis on the search text in combination with the segmented words and their parts of speech to obtain dependency relations between the segmented words;

[0129] According to the segmented words, their parts of speech and dependency relationships, sub-attributes of the search text and correlations between the sub-attributes are determined.

[0130] In one or some embodiments, the semantic analysis module 710 is further adapted to:

[0131] Encode the segmented words and their parts of speech to obtain a segmented word digital code, input the segmented word digital code into an encoder for conversion to obtain a segmented word vector; perform syntactic analysis on the search text to obtain a syntactic feature vector;

[0132] The word segmentation vector and the syntactic feature vector are concatenated and fused to obtain a first latent vector;

[0133] Processing the first latent vector through a multilayer perceptron to obtain a second latent vector;

[0134] A dependency matrix is ​​constructed according to the search text and the dependency, a product value of each position in the dependency matrix is ​​determined according to the dependency matrix, the dependency initialization vector and the second latent vector, and a correlation between the sub-attributes is determined according to the size of the product value and a first loss function.

[0135] In one or some embodiments, the first loss function includes a first cross entropy loss function and a KL divergence loss function;

[0136] Among them, the first cross entropy loss function takes the normalized exponential function and the actual relationship value as independent variables, and the KL divergence loss function takes the average value of the first latent vector and the average value of the target sub-attribute latent vector as independent variables.

[0137] In one or some embodiments, the training step of the semantic matching model in the tag acquisition module 720 includes:

[0138] Extracting sub-attributes from the retrieved text sample according to the semantic analysis result, and generating positive sample vectors, retrieved text vectors and negative sample vectors according to the sub-attribute conversion;

[0139] The positive sample vector, the search text vector and the negative sample vector are respectively subjected to word segmentation and feature encoding by an encoder, and then respectively input into respective bidirectional long short-term memory network models to obtain a positive ratio encoding vector, a text ratio encoding vector and a negative ratio encoding vector;

[0140] After the positive ratio coding vector, the text ratio coding vector and the negative ratio coding vector are processed by a normalized exponential function and a second loss function, a label of a sub-attribute in the retrieved text sample is obtained;

[0141] The labels of the sub-attributes are compared with the manual labels to obtain the corrected semantic matching model.

[0142] In one or some embodiments, the second loss function includes a second cross entropy loss function and a binary classification loss function;

[0143] Among them, the second cross entropy loss function takes the normalized exponential function and the actual label value as independent variables, and the binary classification loss function takes the similarity between the positive ratio encoding vector, the text ratio encoding vector and the negative ratio encoding vector as independent variables.

[0144] In one or some embodiments, the step of obtaining multimedia subcategory tags in the search and matching module 730 includes:

[0145] Extract key frames from multimedia, perform target detection on the key frames, analyze them according to preset hierarchical categories, and obtain labels of corresponding hierarchical categories.

[0146] The key points and beneficial effects of the multimedia retrieval solution disclosed in the present invention include:

[0147] First, the embodiment of the present invention first proposes a fine-grained hierarchical semantic attribute extraction method. Compared with the traditional semantic analysis method, it can realize the extraction of sub-attributes, where the sub-attribute is not just a single word, but a sequence of fragments, and has better recognition capabilities for new words and words with states (such as negative sentences: "not wearing glasses"). In addition, this method can also identify the association relationship between sub-attributes, such as parallel relationships and subordinate relationships. The key parent attribute can be identified from the sub-attributes, so that different retrieval requirements of multiple targets can be identified by searching the text.

[0148] Second, an embodiment of the present invention proposes a semantic encoding method of multi-level comparison and matching, which can realize category classification and recognition under secondary classification, and can better complete the semantic encoding of multi-level classification labels through training of positive and negative sample comparison, so as to more accurately realize the matching of semantic sub-attributes and video labels.

[0149] Third, the embodiments of the present invention overcome the confusion between search terms that may be caused by keyword extraction in existing solutions (for example, if one wants to search for "a man wearing a hat and a woman drinking", it is easy to search for the woman wearing a hat and the man drinking through the existing solutions). The embodiments of the present invention can realize the retrieval of multiple targets in one search question, thereby overcoming the defects of the existing solutions and improving the matching accuracy.

[0150] An embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores at least one executable instruction, and the computer executable instruction can execute the multimedia retrieval method in any of the above method embodiments.

[0151] Fig. 9 The schematic diagram of the structure of the computing device embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the computing device.

[0152] like Fig. 9 As shown, the computing device may include: a processor (processor) 902 , a communication interface (Communications Interface) 904 , a memory (memory) 906 , and a communication bus 908 .

[0153] The processor 902, the communication interface 904, and the memory 906 communicate with each other via a communication bus 908. The communication interface 904 is used to communicate with other devices such as a client or other server network elements. The processor 902 is used to execute a program 910, which can specifically execute the relevant steps in the above-mentioned multimedia retrieval method embodiment for a computing device.

[0154] Specifically, the program 910 may include program codes, which include computer operation instructions.

[0155] The processor 902 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the computing device may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.

[0156] The memory 906 is used to store the program 910. The memory 906 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0157] The program 910 can be specifically used to enable the processor 902 to perform operations corresponding to the above-mentioned multimedia retrieval method embodiment.

[0158] The algorithm or display provided herein is not inherently related to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious to construct the structure required for this type of system. In addition, the embodiment of the present invention is not directed to any specific programming language yet. It should be understood that various programming languages ​​can be utilized to realize the content of the present invention described herein, and the description made to specific languages ​​above is for disclosing the best mode of the present invention.

[0159] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0160] Similarly, it should be understood that in order to streamline the present invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, the various features of the embodiments of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the following intention: that the claimed invention requires more features than the features explicitly recited in each claim. More specifically, as reflected in the claims below, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Therefore, the claims that follow the specific embodiment are hereby expressly incorporated into the specific embodiment, with each claim itself serving as a separate embodiment of the present invention.

[0161] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0162] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.

[0163] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all functions of some or all components according to embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., computer program and computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0164] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be understood as limitations on the order of execution.

Claims

1. A multimedia retrieval method, characterized in that: The method comprises: Analyze the search text to obtain the segmentation words and their parts of speech of the search text; perform dependency syntactic analysis on the search text in combination with the segmentation words and their parts of speech to obtain dependency relations between the segmentation words; encode the segmentation words and their parts of speech to obtain segmentation digital codes, and input the segmentation digital codes into the encoder for conversion to obtain segmentation vectors; perform syntactic analysis on the search text to obtain syntactic feature vectors; concatenate and fuse the segmentation vectors and the syntactic feature vectors to obtain a first latent vector; process the first latent vector through a multilayer perceptron to obtain a second latent vector; construct a dependency matrix according to the search text and the dependency relations, determine the product value of each position in the dependency matrix according to the dependency matrix, the dependency initialization vector and the second latent vector, and determine the correlation between sub-attributes according to the size of the product value and a first loss function; Inputting each of the sub-attributes and the correlation into a pre-trained semantic matching model to obtain a label for each of the sub-attributes; Determine the similarity value between the label of each of the sub-attributes and the pre-obtained multimedia sub-category label, and determine the multimedia sub-category matching the search text according to the magnitude of each of the similarity values.

2. The method according to claim 1, characterized in that The first loss function includes a first cross entropy loss function and a KL divergence loss function; Among them, the first cross entropy loss function takes the normalized exponential function and the actual relationship value as independent variables, and the KL divergence loss function takes the average value of the first latent vector and the average value of the target sub-attribute latent vector as independent variables.

3. The method according to any one of claims 1 to 2, characterized in that: The training steps of the semantic matching model include: Extracting sub-attributes from the retrieved text sample according to the semantic analysis result, and generating positive sample vectors, retrieved text vectors and negative sample vectors according to the sub-attribute conversion; The positive sample vector, the search text vector and the negative sample vector are respectively subjected to word segmentation and feature encoding by an encoder, and then respectively input into respective bidirectional long short-term memory network models to obtain a positive ratio encoding vector, a text ratio encoding vector and a negative ratio encoding vector; After the positive ratio coding vector, the text ratio coding vector and the negative ratio coding vector are processed by a normalized exponential function and a second loss function, a label of a sub-attribute in the retrieved text sample is obtained; The labels of the sub-attributes are compared with the manual labels to obtain the corrected semantic matching model.

4. The method according to claim 3, characterized in that The second loss function includes a second cross entropy loss function and a binary classification loss function; Among them, the second cross entropy loss function takes the normalized exponential function and the actual label value as independent variables, and the binary classification loss function takes the similarity between the positive ratio encoding vector, the text ratio encoding vector and the negative ratio encoding vector as independent variables.

5. The method according to any one of claims 1 to 2, characterized in that: The step of obtaining the multimedia subcategory label includes: Extract key frames from multimedia, perform target detection on the key frames, analyze them according to preset hierarchical categories, and obtain labels of corresponding hierarchical categories.

6. A multimedia retrieval device, characterized in that: The device comprises: A semantic analysis module, adapted to analyze a search text to obtain the segmentation words and their parts of speech of the search text; perform dependency syntactic analysis on the search text in combination with the segmentation words and their parts of speech to obtain dependency relations between the segmentation words; encode the segmentation words and their parts of speech to obtain segmentation digital codes, and input the segmentation digital codes into an encoder for conversion to obtain segmentation vectors; perform syntactic analysis on the search text to obtain syntactic feature vectors; concatenate and fuse the segmentation vectors and the syntactic feature vectors to obtain a first latent vector; process the first latent vector through a multi-layer perceptron to obtain a second latent vector; construct a dependency matrix according to the search text and the dependency relations, determine the product values ​​of each position in the dependency matrix according to the dependency matrix, the dependency initialization vector and the second latent vector, and determine the correlation between sub-attributes according to the size of the product values ​​and a first loss function; A label acquisition module, adapted to input each of the sub-attributes and the correlation into a pre-trained semantic matching model to obtain a label for each of the sub-attributes; The search and matching module is adapted to determine the similarity value between the label of each of the sub-attributes and the pre-obtained multimedia sub-category label, and determine the multimedia sub-category matching the search text according to the size of each of the similarity values.

7. A computing device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the multimedia retrieval method according to any one of claims 1-5.

8. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and the executable instruction enables a processor to execute an operation corresponding to the multimedia retrieval method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Information acquisition method and device, computer equipment and storage medium

    CN109815333A

  • Intelligent full-text retrieval method and system based on semantic understanding

    CN112883165A