Multimedia data search control method and device

By using multi-head attention mechanism and preset attention representation in multimedia data search, the problems of inefficiency and insufficient accuracy in processing multi-type multimedia data are solved, and deeper semantic understanding and personalized search results are achieved.

CN119938956APending Publication Date: 2025-05-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411983624.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional search technology is inefficient and inaccurate in processing large-scale, multi-type multimedia data, and fails to fully utilize the intrinsic connections and complementarity between different types of data.

Method used

By obtaining user input information and preset attention representation in the multimedia database, the target attention representation is determined using the multi-head attention mechanism, and the matching multimedia data is determined as search results based on the representation and mapping relationship.

Benefits of technology

It realizes deeper semantic understanding and information retrieval, improves the intelligence of the search system, enables it to respond more accurately to users' query needs, and provides richer and more personalized search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938956A_ABST
    Figure CN119938956A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia data search control method and device.The method comprises the steps that input information of a user and a multimedia database are obtained, multiple pieces of multimedia data in the multimedia database have preset attention representations, and the preset attention representations represent features of the multimedia data; the preset attention representation comprises a fusion attention representation of the first-level multimedia data and the second-level multimedia data and an attention representation of the second-level multimedia data, and the first-level multimedia data comprises the second-level multimedia data; determining a target attention representation matched with the input information according to the multi-head attention mechanism, the input information and a plurality of preset attention representations; and according to the target attention representation and the mapping relationship, determining multimedia data matched with the input information, and taking the multimedia data as a search result. According to the method and the device, the problems of low efficiency and insufficient accuracy when large-scale and multi-type multimedia data is processed by a traditional search technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of information retrieval, and in particular to a method for controlling multimedia data search, a control device for multimedia data search, a computer-readable storage medium, and an electronic device. Background Art

[0002] With the rapid development of information technology, the amount of multimedia data has increased dramatically, which has not only brought unprecedented opportunities but also challenges to information retrieval and analysis. Traditional search methods often face the problems of low efficiency and insufficient accuracy when processing large-scale multimedia data. The complexity and diversity of multimedia data make it difficult for a single search technology to meet the growing demand. Multimedia data usually includes multiple types such as text, images, audio and video, each of which has unique characteristics and information. For example, text data contains keywords and semantic information, image data contains visual features such as color, texture and shape, audio data contains sound features such as frequency, rhythm and timbre, and video data combines the dynamic characteristics of images and time series.

[0003] Existing technologies often process these data types independently and lack effective fusion mechanisms, which limits the comprehensiveness and accuracy of search results. For example, a search engine based solely on text may not be able to effectively index and retrieve image or video content, while an image search engine that focuses on visual features may not be able to understand text descriptions or audio annotations associated with the image. This separate processing approach ignores the inherent connections and complementarities between different types of data, thereby affecting the quality of search results and user experience. Summary of the invention

[0004] The main purpose of the present application is to provide a control method for multimedia data search, a control device for multimedia data search, a computer-readable storage medium and an electronic device, so as to at least solve the problems of low efficiency and insufficient accuracy encountered by traditional search technologies when processing large-scale, multi-type multimedia data.

[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a control method for multimedia data search is provided, including: obtaining user input information and a multimedia database, wherein multiple multimedia data in the multimedia database all have preset attention representations, and the preset attention representations are characterized as features of the multimedia data, and the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of the secondary multimedia data, and the primary multimedia data include the secondary multimedia data; according to a multi-head attention mechanism, the input information and multiple preset attention representations, determining a target attention representation that matches the input information; according to the target attention representation and the mapping relationship, determining the multimedia data that matches the input information, and using the multimedia data as search results.

[0006] Optionally, the control method further includes: performing format conversion on the input information; and determining keyword information of the input information according to the input information after the format conversion.

[0007] Optionally, the above control method also includes: acquiring initial multimedia data, the initial multimedia data including picture data, text data and audio data; preprocessing the initial multimedia data, the preprocessing including format conversion and feature extraction, the features of the feature extraction including at least: color and shape of picture information, spectrogram and spectrum of audio information.

[0008] Optionally, the preprocessing of the initial multimedia data includes: classifying the initial multimedia data once according to the data format to obtain a plurality of the first-level multimedia data, and performing a first feature extraction on the first-level multimedia data according to a feature extraction algorithm to obtain a first target feature; classifying the first-level multimedia data twice according to the type of data parameters to obtain a plurality of the second-level multimedia data, and performing a second feature extraction on the second-level multimedia data according to the feature extraction algorithm to obtain a second target feature.

[0009] Optionally, the above-mentioned control method also includes: determining a first attention representation of the secondary multimedia data according to the second target feature and the multi-head attention mechanism; determining a second attention representation of the primary multimedia data and the secondary multimedia data according to the first target feature, the second target feature and the multi-head attention mechanism; determining the third attention representation according to the first attention representation and the second attention representation; determining the preset attention representation according to the first attention representation, the second attention representation and the third attention representation.

[0010] Optionally, the feature extraction technology includes a short-time Fourier transform algorithm, a recurrent neural network algorithm and a convolutional neural network algorithm.

[0011] Optionally, the above-mentioned control method also includes: determining the matching degree between multiple preset attention representations and the input information according to the multi-head attention mechanism, the input information and the multiple preset attention representations; and determining that the multimedia data matching the input information is output in sequence as search results according to the matching degree, the preset attention representations and the mapping relationship.

[0012] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a control device for multimedia data search is provided, including: an acquisition module, used to acquire user input information and a multimedia database, wherein multiple multimedia data in the multimedia database all have preset attention representations, and the preset attention representations are characterized as features of the multimedia data, and the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of the secondary multimedia data, and the primary multimedia data include the secondary multimedia data; a first determination module, used to determine a target attention representation that matches the input information based on a multi-head attention mechanism, the input information and multiple preset attention representations; a second determination module, used to determine the multimedia data that matches the input information based on the target attention representation and a mapping relationship, and use the multimedia data as search results.

[0013] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the control method for searching multimedia data.

[0014] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a control method for performing the above-mentioned multimedia data search.

[0015] By applying the technical solution of the present application, by obtaining the user's input information and a multimedia database, multiple multimedia data in the multimedia database all have preset attention representations, the preset attention representations are characterized as features of the multimedia data, the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of secondary multimedia data, the primary multimedia data includes secondary multimedia data; according to the multi-head attention mechanism, input information and multiple preset attention representations, the target attention representation matching the input information is determined; according to the target attention representation and the mapping relationship, the multimedia data matching the input information is determined, and the multimedia data is used as the search result, and the control method of the multimedia data search can automatically identify and integrate the key features of various media data, achieve a deeper level of semantic understanding and information retrieval, improve the intelligence of the search system, enable it to respond to the user's query needs more accurately, and provide richer and more personalized search results. Unifying the feature representations of multiple multimedia data into the target attention representation enhances the intrinsic connection and complementarity between different types of data and improves the accuracy of the search. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings constituting part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 A hardware structure block diagram of a mobile terminal for executing a control method for multimedia data search provided in an embodiment of the present application is shown;

[0018] Figure 2 A schematic diagram of a process flow of a multimedia data search control method provided according to an embodiment of the present application is shown;

[0019] Figure 3 A schematic structural diagram of a multimedia data search control device provided according to an embodiment of the present application is shown.

[0020] The above drawings include the following reference numerals:

[0021] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION

[0022] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] As introduced in the background technology, traditional search technologies in the prior art often independently process data in different formats such as text, images, audio and video when processing large-scale multimedia data, ignoring the intrinsic connections and complementarities between different types of data, and failing to fully utilize the intrinsic connections and complementarities between these data, which leads to a lack of an effective fusion mechanism, thereby limiting the comprehensiveness and accuracy of search results, and further affecting the quality of search results and user experience. In order to solve the problems of inefficiency and insufficient accuracy encountered by traditional search technologies when processing large-scale, multi-type multimedia data, embodiments of the present application provide a control method for multimedia data search, a control device for multimedia data search, a computer-readable storage medium and an electronic device.

[0026] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0027] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure block diagram of a mobile terminal of a multimedia data search control method according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0028] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the control method for multimedia data search in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0029] In this embodiment, a control method for multimedia data search running on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0030] Figure 2 FIG. 1 is a flow chart of a method for controlling multimedia data search according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0031] Step S101, obtaining user input information and a multimedia database, wherein multiple multimedia data in the multimedia database all have preset attention representations, and the preset attention representations are characterized as features of the multimedia data. The preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of secondary multimedia data, and the primary multimedia data include secondary multimedia data.

[0032] Specifically, the user's input information and preset attention representation can be used to effectively filter and sort multimedia data, improve the relevance and accuracy of search results, and better capture the important information and relevance of multimedia data through the features of preset attention representation, so as to recommend and sort search results according to user needs and interests. By fusing the attention representation of primary and secondary multimedia data, the quality of search results and user experience can be further improved, making it easier for users to find multimedia data of interest.

[0033] Specifically, the above-mentioned multimedia data (MultimediaData) is a combination of data in various media formats, which may include text, images, audio, video and other data. However, it should be noted that the multimedia data in this application is not limited to the above-mentioned types, and the technical personnel of this application do not make specific limitations.

[0034] Step S102: Determine a target attention representation that matches the input information based on the multi-head attention mechanism, the input information, and multiple preset attention representations.

[0035] Specifically, through the multi-head attention mechanism, the system can focus on different parts of the input information at the same time, improve the processing efficiency and accuracy of multimedia data, and through multiple preset attention representations, adjust the generated target attention representation according to different task requirements, and integrate multiple types of multimedia data information. It is no longer a single search for a certain type, making the search results more in line with user needs, thereby more accurately matching the target information of the user's query.

[0036] Step S103, determining multimedia data matching the input information according to the target attention representation and the mapping relationship, and using the multimedia data as search results.

[0037] Specifically, by determining the matching multimedia data based on the user's target attention representation and mapping relationship, the content related to the user's needs can be effectively screened out, reducing the user's time and energy consumption in the search process. The accuracy and relevance of search results can also be improved, making it easier for users to find the information they need, thereby improving user satisfaction and search experience.

[0038] In the specific implementation process, the above method also includes the following steps: converting the format of the input information; and determining keyword information of the input information according to the input information after the format conversion.

[0039] In this solution, by converting the format of input information, input information of different formats can be processed uniformly, so that the system can better understand and analyze user input. By determining the keyword information of the input information, it can help the system accurately identify the user's search intent, thereby more accurately matching the user's needs and providing relevant multimedia data, thereby improving the accuracy and efficiency of the system search.

[0040] For example, for image or video data, it can be converted into a specific encoding format or data structure so that the system can analyze and search it, and use text mining technology or natural language processing technology to extract keyword information from the input information.

[0041] During the specific implementation process, the above method also includes the following steps: obtaining initial multimedia data, which includes picture data, text data and audio data; preprocessing the initial multimedia data, which includes format conversion and feature extraction, and the features of the feature extraction include at least: the color and shape of the picture information, and the spectrogram and spectrum of the audio information.

[0042] In this solution, the initial multimedia data is obtained and formatted so that the data can be correctly read and processed by the system. By extracting information such as color and shape from image data, and extracting information such as spectrogram and spectrum from audio data, the multimedia data can be converted into a vector representation with specific features, which is convenient for similarity calculation and search. Through preprocessing and feature extraction, the search scope can be greatly narrowed, the amount of calculation can be reduced, and the search efficiency can be improved. By extracting the key features of multimedia data, data matching and retrieval can be performed more accurately, improving the accuracy of search results.

[0043] Specifically, the above preprocessing includes format conversion and feature extraction of image, audio, video and other data, wherein the feature extraction may include but is not limited to: color, texture, shape, spectrogram, spectrum, key frame and the like.

[0044] In the specific implementation process, the initial multimedia data is preprocessed through the following steps: the initial multimedia data is classified once according to the data format to obtain multiple first-level multimedia data, and the first-level multimedia data is subjected to a first feature extraction according to the feature extraction algorithm to obtain a first target feature; the first-level multimedia data is classified twice according to the data parameter type to obtain multiple second-level multimedia data, and the second-level multimedia data is subjected to a second feature extraction according to the feature extraction algorithm to obtain a second target feature.

[0045] In this solution, through classification and feature extraction, the search scope can be reduced and the search space can be narrowed, thereby improving the search speed and efficiency. By extracting key features, the user's search needs can be more accurately matched, and the relevance and accuracy of search results can be improved. By finely processing and classifying multimedia data, and fusing feature representation of classified and processed multimedia data, users can be provided with more targeted and personalized search results, thereby improving the user experience.

[0046] Specifically, the initial multimedia data is classified into a plurality of first-level multimedia data, including images, texts and audio data, and the first-level multimedia data is subjected to first feature extraction according to a primary feature extractor (including a feature extraction algorithm), wherein the image feature extractor can use a convolutional neural network to extract basic visual features from the image, the text feature extractor can use a transformer to extract semantic features from the text, and the audio feature extractor can use a convolutional neural network or a recurrent neural network to extract acoustic features from the audio. The specific steps are as follows:

[0047] For image I, text T, and audio A, primary features (i.e., first target features) are obtained through their respective feature extractors:

[0048] F I =f CNN (I; θ I ), F T =f Transformer (T; θ T ), F A =f A (A; θ A ),

[0049] Among them, f CNN is the function of the convolutional neural network, f Transformer is a function of the converter model, θ I ,θ T and θ A are the feature parameters of image I, text T, and audio A, respectively. These parameters are learned during the model training process and are used to adjust the weights and biases within the primary feature vector function model of image I, text T, and audio A to optimize the performance of the model. I is the primary feature vector of the image, F T is the primary feature vector of the text, F A is the primary feature vector of the audio, F I 、F T and F A It is the output of the original data after being processed by the feature extractor in the respective field (i.e., the first target feature). The output data can be text or audio transcription text or text in the picture.

[0050] Specifically, according to the type of data parameters, the primary multimedia data is secondary classified to obtain multiple secondary multimedia data, wherein the secondary features of the image include pixel features, background text features in the video and subtitle text features, and the secondary features of the audio include audio spectrum features and audio text features. According to the secondary feature extractor (including feature extraction algorithm), the secondary multimedia data is subjected to second feature extraction, and the secondary feature extractor includes a secondary feature extractor of the image and a secondary feature extractor of the audio. The secondary feature extractor of the image may include a pixel feature extractor, a background text feature extractor and a subtitle text feature extractor, and the secondary feature extractor of the audio may include an audio spectrum feature extractor and an audio text feature extractor, wherein the pixel feature extractor is directly extracted from the original image, the background text feature extractor uses OCR (Optical Character Recognition) technology to recognize the text in the image, and then uses a transformer to extract features, the subtitle text feature extractor uses a transformer to extract features from the subtitle text, and the audio spectrum feature extractor may use a short-time Fourier transform (STFT) or a Mel Frequency Cepstral coefficient (Mel Frequency Cepstral Coefficients, MFCC) extracts the spectral features of the audio. The audio text feature extractor uses speech recognition technology to convert the audio into text, and then uses the transformer to extract features. The specific steps are as follows:

[0051] 1) The following formula is used for the secondary features of the image:

[0052] F pix =f CNN (I; θ pix ), F Bt =f Transformer (f OCR (I); θ Bt ), F St =f Transformer (S t θ St ),

[0053] Among them, θ pix is the pixel feature, θ Bt is the background text feature in the video, θ St is the subtitle text feature, f OCR It is the function of optical character recognition (OCR), which is used to extract text from images. Transformer is a function of the converter model. pix ,θ Btand θ St Represent the model parameters of pixel features, background text features in the video, and subtitle text features respectively.

[0054] 2) The following formula is used for the secondary features of audio:

[0055] F spec =f STFT (A; θ spec ), F At =f Transformer (f ASR (A); θ At ),

[0056] Among them, θ spec is the audio spectrum feature, θ At is the audio text feature, f STFT Represents short-time Fourier transform, which is used to extract frequency domain features from time domain signals and is often used in audio processing. ASR is the function of automatic speech recognition (ASR) to extract text from audio, Transformer is a function of the transformer model.

[0057] It should be noted that the above F pix 、F Bt 、F St 、F spec and F At is a higher-level feature for images and audio (i.e., the second target feature), for example, F pix Refers to the pixel-level features extracted from the image, F Bt is the feature of the text content extracted from the image, F St are features of scene description extracted from images.

[0058] In the specific implementation process, the above method also includes the following steps: determining a first attention representation of secondary multimedia data based on the second target feature and the multi-head attention mechanism; determining a second attention representation of primary multimedia data and secondary multimedia data based on the first target feature, the second target feature and the multi-head attention mechanism; determining a third attention representation based on the first attention representation and the second attention representation; determining and pre-setting the attention representation based on the first attention representation, the second attention representation and the third attention representation.

[0059] In this scheme, by utilizing the multi-head attention mechanism combined with different target features to determine different levels of attention representation, the association and importance between multimedia data can be better captured, thereby improving the accuracy and effectiveness of search results. Through multi-level attention representation, the user's query intention and the content of multimedia data can be more comprehensively understood, improving the relevance of search results and user experience. By comparing and matching with the preset attention representation, the user's search needs can be better met and search results that are more in line with user expectations can be provided.

[0060] It should be noted that the above-mentioned multi-head attention mechanism is a technology used to enhance the performance of neural networks, especially in natural language processing tasks. By paying attention to different parts of the input sequence at the same time and learning different representations in different attention heads, the network's expressiveness and generalization capabilities are improved. In the multi-head attention mechanism, the input sequence is first mapped to multiple queries, keys, and numeric vectors. Each attention head calculates an attention distribution for weighted summation of different parts of the input sequence. The outputs of multiple heads are linearly transformed and concatenated and then projected to the final output space again. By using the multi-head attention mechanism, multiple information can be paid attention to simultaneously during the search process, the correlation between different media data can be effectively utilized, and the quality and relevance of search results can be improved.

[0061] Specifically, a multi-head attention mechanism is applied at each level to calculate the correlation between different features. In a hierarchical manner, attention is first calculated between the secondary features, and then these features are aggregated to the primary feature level for recalculation. The specific steps are as follows:

[0062] 1) Calculate the attention between the secondary features of the image. You can use the multi-head attention mechanism to strengthen the connection between different features and get the first attention representation:

[0063] H I,pix =MultiHeadAttention(Q=F pix ,K=F pix ,V=F pix ; W Q ,W K ,W V ),

[0064] H I,Bt =MultiHeadAttention(Q=F pix ,K=F Bt ,V=F Bt ; W Q ,W K ,W V ),

[0065] HI,St =MultiHeadAttention(Q=F pix ,K=F St ,V=F St ; W Q ,W K ,W V ),

[0066] 2) For the secondary features of the audio, the second attention representation is obtained:

[0067] H A,spec =MultiHeadAttention(Q=F spec ,K=F spec ,V=F St ; W Q ,W K ,W V ),

[0068] H A,At =MultiHeadAttention(Q=F spec ,K=F At ,V=F At ; W Q ,W K ,W V ),

[0069] 3) Calculate cross-modal attention between the primary and secondary features of different modalities to obtain the third attention representation:

[0070] H cross,IT =MultiHeadAttention(Q=F I ,K=F T ,V=F T ; W Q ,W K ,W V ),

[0071] H cross,IA =MultiHeadAttention(Q=F I ,K=F A ,V=F A ; W Q ,W K ,W V ),

[0072] H cross,TA =MultiHeadAttention(Q=F T ,K=F A ,V=F A ; W Q ,WK ,W V ),

[0073] 4) All features processed by the attention mechanism are combined to form the final comprehensive features (i.e., the preset attention representation). The steps are as follows:

[0074] H=g(H I,pix ,H I,Bt ,H I,St ,H A,spec ,H A,At ,H cross,IT ,H cross,IA ,H cross,TA ; Θ g ),

[0075] Among them, W Q ,W K ,W V is the weight matrix in the multi-head attention mechanism, that is, the weight matrix of query, key, and value in the multi-head attention mechanism, which is used to transform the original feature vector to facilitate the calculation of the attention score, H I,pix , H I,Bt , H I,St , H A,spec and H A,At It is the feature after applying the multi-head attention mechanism, reflecting the interaction and importance between different features. cross,IT , H cross,IA , H cross,TA is the result of cross-modal attention between features of different modalities, used to capture information interaction between different modalities, H is the final comprehensive feature vector, obtained by fusing all previously calculated features, g is a nonlinear transformation function used to integrate features from different sources, and can be a neural network composed of several fully connected layers, such as a neural network composed of one or more fully connected layers, Θ g is the set of parameters of the function g.

[0076] In the specific implementation process, the above feature extraction technologies include short-time Fourier transform algorithm, recurrent neural network algorithm and convolutional neural network algorithm.

[0077] In this solution, feature extraction technology can more effectively identify and extract important features in the data, thereby improving the accuracy and efficiency of search results. The short-time Fourier transform algorithm can be used for frequency domain feature extraction of audio and video data, the recurrent neural network algorithm can be used for feature extraction of sequence data, and the convolutional neural network algorithm can be used for feature extraction of image data. By combining the above algorithms, multimedia data can be analyzed and understood more comprehensively, thereby improving the performance of the search system and user experience.

[0078] Specifically, the Short-time Fourier Transform (STFT) algorithm is a commonly used signal processing technology for analyzing signals in time and frequency. The STFT algorithm can be used to perform spectrum analysis on multimedia data, thereby realizing retrieval and search of these data. The Recurrent Neural Network (RNN) algorithm is a neural network algorithm that can process sequence data. It has memory ability when processing sequence data and can capture the time dependency in sequence data. It usually uses back propagation algorithm and gradient descent algorithm to update the parameters of the neural network by minimizing the loss function, thereby improving the accuracy of the model. The Convolutional Neural Network (CNN) is a commonly used deep learning algorithm. By constructing a CNN model suitable for multimedia data search tasks, it can extract features and learn representations of multimedia data, thereby realizing effective search and retrieval of data. The accuracy and efficiency of the search can be improved by adjusting the structure and parameters of the CNN model to meet the needs of multimedia data search.

[0079] During the specific implementation process, the above method also includes the following steps: determining the matching degree between multiple preset attention representations and the input information based on the multi-head attention mechanism, input information and multiple preset attention representations; determining that the multimedia data matching the input information is output in sequence as search results based on the matching degree, preset attention representation and mapping relationship.

[0080] In this scheme, through the matching degree of the multi-head attention mechanism and the preset attention representation, the relevance of multimedia data to input information can be determined more effectively, thereby improving the accuracy and relevance of search results. By evaluating the matching degree of multiple preset attention representations, multimedia data with a high degree of matching with the input information can be better selected for output, thereby improving user experience and search results, and can effectively improve the accuracy and relevance of search results and improve user satisfaction.

[0081] Specifically, one or more fully connected layers are used to map the comprehensive features to the final output space:

[0082] O=h(H;Θ h ),

[0083] Among them, O is the final output of the model (which can be a classification label, regression value or other form of prediction), h is the function that maps the comprehensive features to the output space, usually a fully connected layer, Θ h is the set of parameters of the function h.

[0084] The training goal of the model is to maximize the performance of the specific task and the accuracy of sentiment classification:

[0085] max∑log P(y|I,T,A,S t ,B t ),

[0086] Where y is the true label of a single sample, log P(y|I,T,A,S t ,B t ) is the logarithm of the probability that the model predicts the correct label given an input, and is an important indicator for evaluating model performance.

[0087] Specifically, O is the final output of the model, which can be a classification label, regression value, or other forms of prediction results. max includes all parameters that need to be learned. During the training process, these parameters are continuously updated through methods such as gradient descent to minimize the loss function. Its training data set includes input data and corresponding labels. The above model structure can better capture the complex relationship between different modalities and different levels of features, thereby achieving better performance in multimodal tasks.

[0088] Specifically, the above-mentioned data fusion (Data Fusion) is the process of combining data or information from multiple sources to obtain more accurate, comprehensive and reliable results or conclusions, and the above-mentioned multimedia data fusion search (Multimedia Data Fusion Search) is a search method involving information fusion of multiple media forms (such as text, images, audio, video, etc.).

[0089] Exemplarily, the control method for multimedia data search in the embodiment of the present application includes: first, preprocessing the input multimedia data, including format conversion and feature extraction of image, audio, video and other data; fusing the preprocessed multimedia data to form a unified multimedia data representation, which not only includes various features of the multimedia data, but also includes the intrinsic connection and complementarity between the multimedia data; receiving the user's search request, parsing and processing the search request, and extracting the user's query intention and keywords; calculating the similarity between the user's query intention and keywords and the multimedia data representation, obtaining a list of multimedia data related to the user's needs, and sorting them according to the similarity; outputting the sorted multimedia data list as the search result.

[0090] Specifically, first create the basic framework of the search system, then configure the multimedia data processing module through the graphical interface, define the multimedia data source, data structure and data processing flow, and finally, deploy the search system to the target environment, and perform testing and optimization.

[0091] In addition, the following functions may also be included but are not limited to:

[0092] 1) Support the association of multiple data sources, integrate data from different sources, and provide a unified search interface;

[0093] 2) Support visual display of search results, such as charts, maps, etc.;

[0094] 3) Access data sources and analyze the characteristics of each data source to determine the most suitable access method (direct database connection, Application Programming Interfac (API) call, etc.);

[0095] 4) Develop or configure ETL (Extract-Transform-Load) processes to regularly synchronize data to centralized storage;

[0096] 5) Data standardization and cleaning, implementation of data cleaning rules, such as removing duplicate records, filling missing values, etc.;

[0097] 6) Define a set of standard data models to ensure consistent data formats imported from different data sources;

[0098] 7) Build a search engine to index the standardized data into a full-text search engine;

[0099] 8) Configure the search engine to optimize search performance, such as setting a reasonable word segmenter, adjusting the relevance algorithm, etc.;

[0100] 9) Develop a unified search interface and design a simple and intuitive user interface to ensure that users can easily enter search criteria;

[0101] 10) Implement the front-end and back-end interaction logic. When a user submits a search request, the front-end application initiates an API request to the back-end service.

[0102] 11) The backend service is responsible for processing the request, calling the search engine API to obtain the results, and returning the results to the front end.

[0103] Through the above-mentioned multimedia data search control method, multimedia data can be integrated to search, identify and integrate the key features of various media data, provide richer and more personalized search results, integrate the features of different media types into a unified feature representation, consider the correlation and complementarity between features, and in the process of optimizing the recommendation algorithm, since different types of data can provide different information dimensions, it is necessary to consider the characteristics of multimedia data, which in turn affects the accuracy and personalization of the recommendation. The close relationship between the features is fully considered to form a comprehensive feature space, and the ability of the search system to handle cross-media queries is improved. The main process steps of this method include multimedia data preprocessing, multimedia data fusion, search request processing, similarity calculation and sorting, and detection result output. It can be mainly applied to various public security search scenarios that require efficient search functions, such as police personnel data search, police event data search, police item data search, and police organization data search.

[0104] The technical solution provided by the above-mentioned embodiments of the present invention solves the efficiency and accuracy problems encountered by traditional search technologies when processing large-scale, multi-type multimedia data. Existing search systems often independently process data in different formats such as text, images, audio and video, and fail to fully utilize the inherent connections and complementarities between these data, resulting in limitations in search results. The present invention aims to achieve deeper semantic understanding and information retrieval by integrating data features of different media types. It can automatically identify and integrate key features of various media data, and improve the intelligence of the search system by constructing a unified feature representation, so that it can more accurately respond to user query needs, while providing richer and more personalized search results. It can identify and utilize the correlation between different data types, enhance the semantic understanding ability and accuracy of the search, reduce false detections and missed detections, and thus provide users with a richer, more accurate and personalized search experience.

[0105] The embodiment of the present application also provides a control device for multimedia data search. It should be noted that the control device for multimedia data search in the embodiment of the present application can be used to execute the control method for multimedia data search provided in the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions thereof will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0106] The following is an introduction to the control device for multimedia data search provided in an embodiment of the present application.

[0107] Figure 3Schematic diagram of a control device for multimedia data search according to an embodiment of the present application. Figure 3 As shown, the device comprises:

[0108] The acquisition module 100 is used to acquire user input information and a multimedia database, wherein the plurality of multimedia data in the multimedia database all have preset attention representations, the preset attention representations are characterized as features of the multimedia data, the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of secondary multimedia data, and the primary multimedia data include secondary multimedia data;

[0109] A first determination module 200, configured to determine a target attention representation matching the input information according to the multi-head attention mechanism, the input information and a plurality of preset attention representations;

[0110] The second determination module 300 is used to determine multimedia data matching the input information according to the target attention representation and the mapping relationship, and use the multimedia data as the search result.

[0111] Through this embodiment, the user's input information and multimedia database can be obtained through the acquisition module 100, and multiple multimedia data in the multimedia database all have preset attention representations, and the preset attention representations are characterized as features of the multimedia data. The preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of secondary multimedia data, and the primary multimedia data include secondary multimedia data; the first determination module 200 is used to determine the target attention representation matching the input information according to the multi-head attention mechanism, input information and multiple preset attention representations; the second determination module 300 is used to determine the multimedia data matching the input information according to the target attention representation and the mapping relationship, and use the multimedia data as the search result; the control device for searching multimedia data can automatically identify and integrate the key features of various media data, and achieve deeper semantic understanding and information retrieval by integrating data features of different media types. By constructing a unified feature representation, the intelligence of the search system is improved, so that it can respond to user query needs more accurately, and provide richer and more personalized search results.

[0112] Specifically, the acquisition module 100 can effectively filter and sort multimedia data through the user's input information and preset attention representation, improve the relevance and accuracy of search results, and better capture the important information and relevance of multimedia data through the characteristics of the preset attention representation, so as to recommend and sort search results according to the user's needs and interests. By integrating the attention representation of primary and secondary multimedia data, the quality of search results and user experience can be further improved, making it easier for users to find multimedia data of interest.

[0113] Specifically, the first determination module 200 uses a multi-head attention mechanism, so that the system can focus on different parts of the input information at the same time, thereby improving the processing efficiency and accuracy of multimedia data. Through multiple preset attention representations, the generated target attention representation can be adjusted according to different task requirements, so that the search results are more in line with user needs, thereby more accurately matching the target information of the user query.

[0114] Specifically, the second determination module 300 can effectively screen out content related to user needs by determining matching multimedia data based on the user's target attention representation and mapping relationship, reduce the user's time and energy consumption in the search process, and improve the accuracy and relevance of search results, making it easier for users to find the information they need, thereby improving user satisfaction and search experience.

[0115] In some optional implementations, the control device in the embodiment of the present application further includes: a format conversion module for converting the format of the input information; and a third determination module for determining keyword information of the input information based on the input information after format conversion.

[0116] In some optional embodiments, the control device in the embodiment of the present application also includes: a second acquisition module, used to acquire initial multimedia data, the initial multimedia data includes picture data, text data and audio data; a preprocessing module, used to preprocess the initial multimedia data, the preprocessing includes format conversion and feature extraction, the features of the feature extraction include at least: color and shape of the picture information, spectrogram and spectrum of the audio information.

[0117] In some optional embodiments, the above-mentioned preprocessing module includes: a first classification unit, used to classify the initial multimedia data once according to the data format to obtain multiple first-level multimedia data, and perform a first feature extraction on the first-level multimedia data according to a feature extraction algorithm to obtain a first target feature; a second classification unit, used to classify the first-level multimedia data twice according to the data parameter type to obtain multiple second-level multimedia data, and perform a second feature extraction on the second-level multimedia data according to a feature extraction algorithm to obtain a second target feature.

[0118] In some optional embodiments, the control device in the embodiment of the present application also includes: a first determination unit, used to determine a first attention representation of the secondary multimedia data based on the second target feature and the multi-head attention mechanism; a second determination unit, used to determine the second attention representation of the primary multimedia data and the secondary multimedia data based on the first target feature, the second target feature and the multi-head attention mechanism; a third determination unit, used to determine the third attention representation based on the first attention representation and the second attention representation; and a fourth determination unit, used to determine and preset the attention representation based on the first attention representation, the second attention representation and the third attention representation.

[0119] In some optional implementations, the above-mentioned feature extraction technology includes a short-time Fourier transform algorithm, a recurrent neural network algorithm, and a convolutional neural network algorithm.

[0120] In some optional embodiments, the control device in the embodiment of the present application also includes: a fourth determination module, used to determine the matching degree of multiple preset attention representations with the input information based on the multi-head attention mechanism, the input information and the multiple preset attention representations; a fifth determination module, used to determine that the multimedia data matching the input information is output in sequence as a search result based on the matching degree, the preset attention representation and the mapping relationship.

[0121] The control device for multimedia data search may include a processor and a memory. The acquisition module, the first determination module, the second determination module, etc. are all stored in the memory as program units, and the processor executes the program units stored in the memory to implement corresponding functions. The modules are all located in the same processor; or, the modules are located in different processors in any combination. The processor includes a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels may be provided.

[0122] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0123] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the control method for searching multimedia data.

[0124] Specifically, the control method for multimedia data search includes:

[0125] Step S101, obtaining user input information and a multimedia database, wherein the plurality of multimedia data in the multimedia database all have preset attention representations, the preset attention representations are characterized as features of the multimedia data, the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of secondary multimedia data, and the primary multimedia data include secondary multimedia data;

[0126] Specifically, the user's input information and preset attention representation can be used to effectively filter and sort multimedia data, improve the relevance and accuracy of search results, and better capture the important information and relevance of multimedia data through the features of preset attention representation, so as to recommend and sort search results according to user needs and interests. By fusing the attention representation of primary and secondary multimedia data, the quality of search results and user experience can be further improved, making it easier for users to find multimedia data of interest.

[0127] Step S102, determining a target attention representation matching the input information according to the multi-head attention mechanism, the input information and a plurality of preset attention representations;

[0128] Specifically, through the multi-head attention mechanism, the system can focus on different parts of the input information at the same time, improve the processing efficiency and accuracy of multimedia data, and through multiple preset attention representations, adjust the generated target attention representation according to different task requirements, so that the search results are more in line with user needs, thereby more accurately matching the target information of the user query.

[0129] Step S103, determining multimedia data matching the input information according to the target attention representation and the mapping relationship, and using the multimedia data as search results.

[0130] Specifically, by determining the matching multimedia data based on the user's target attention representation and mapping relationship, the content related to the user's needs can be effectively screened out, reducing the user's time and energy consumption in the search process. The accuracy and relevance of search results can also be improved, making it easier for users to find the information they need, thereby improving user satisfaction and search experience.

[0131] An embodiment of the present invention provides an electronic device, the device includes a processor, a memory, and a program stored in the memory and executable on the processor, and when the processor executes the program, at least the following steps are implemented: obtaining user input information and a multimedia database, wherein multiple multimedia data in the multimedia database all have preset attention representations, and the preset attention representations are characterized as features of the multimedia data, and the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of secondary multimedia data, and the primary multimedia data include secondary multimedia data; determining a target attention representation that matches the input information based on a multi-head attention mechanism, input information, and multiple preset attention representations; determining multimedia data that matches the input information based on the target attention representation and the mapping relationship, and using the multimedia data as search results. The electronic device in this article may be a server, a PC, a PAD, a mobile phone, etc.

[0132] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0133] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0134] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0135] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0137] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0138] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0139] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0140] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0141] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A control method for multimedia data search, characterized in that: include: Acquire user input information and a multimedia database, wherein the plurality of multimedia data in the multimedia database all have preset attention representations, the preset attention representations are characterized as features of the multimedia data, the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of the secondary multimedia data, and the primary multimedia data include the secondary multimedia data; Determine a target attention representation matching the input information according to the multi-head attention mechanism, the input information and the plurality of preset attention representations; According to the target attention representation and the mapping relationship, the multimedia data matching the input information is determined, and the multimedia data is used as the search result.

2. The control method according to claim 1, characterized in that: The control method further comprises: Converting the input information into a new format; According to the input information after the format conversion, keyword information using the input information is determined.

3. The control method according to claim 1, characterized in that: The control method further comprises: Acquiring initial multimedia data, wherein the initial multimedia data includes picture data, text data, and audio data; The initial multimedia data is preprocessed, the preprocessing comprising format conversion and feature extraction, the features of the feature extraction comprising at least: color and shape of picture information, spectrogram and spectrum of audio information.

4. The control method according to claim 3, characterized in that: The preprocessing of the initial multimedia data includes: Classifying the initial multimedia data once according to the data format to obtain the plurality of primary multimedia data, and extracting the first feature of the primary multimedia data according to the feature extraction algorithm to obtain the first target feature; According to the data parameter types, the primary multimedia data is secondary classified to obtain a plurality of secondary multimedia data, and according to the feature extraction algorithm, the secondary multimedia data is subjected to second feature extraction to obtain a second target feature.

5. The control method according to claim 4, characterized in that: The control method further comprises: Determining a first attention representation of the secondary multimedia data according to the second target feature and the multi-head attention mechanism; Determining a second attention representation of the primary multimedia data and the secondary multimedia data according to the first target feature, the second target feature, and the multi-head attention mechanism; Determining a third attention representation according to the first attention representation and the second attention representation; According to the first attention representation, the second attention representation and the third attention representation, determine the preset attention representation.

6. The control method according to claim 4, characterized in that: The feature extraction technology includes a short-time Fourier transform algorithm, a recurrent neural network algorithm and a convolutional neural network algorithm.

7. The control method according to claim 1, characterized in that: The control method further comprises: Determining, according to the multi-head attention mechanism, the input information and the plurality of preset attention representations, a degree of matching between the plurality of preset attention representations and the input information; According to the matching degree, the preset attention representation and the mapping relationship, it is determined that the multimedia data matching the input information is output in sequence as a search result.

8. A control device for multimedia data search, characterized in that: include: an acquisition module, configured to acquire user input information and a multimedia database, wherein the plurality of multimedia data in the multimedia database all have preset attention representations, wherein the preset attention representations are characterized as features of the multimedia data, wherein the preset attention representations include fused attention representations of primary multimedia data and secondary multimedia data, and attention representations of the secondary multimedia data, wherein the primary multimedia data include the secondary multimedia data; A first determination module is used to determine a target attention representation matching the input information according to the multi-head attention mechanism, the input information and the plurality of preset attention representations; The second determination module is used to determine the multimedia data matching the input information according to the target attention representation and the mapping relationship, and use the multimedia data as the search result.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the control method for multimedia data search according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a control method for executing the multimedia data search described in any one of claims 1 to 7.