Picture sound retrieval method and device, equipment and storage medium

By constructing a graph-sound retrieval model trained on data and text data based on multiple speech images, the problem of low accuracy in the prior art image retrieval is solved, and more efficient image and speech data matching is achieved.

CN120086406APending Publication Date: 2025-06-03HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411979281.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing image and audio search methods have shortcomings in retrieval accuracy, and it is difficult to effectively match image and voice data.

Method used

By constructing a graph-sound search model, the model trains the data and related text data through multiple speech images, and can match the speech data to be retrieved with the related image data or the image data to be retrieved with the related speech data.

Benefits of technology

Improve the matching accuracy between image data and speech data and enhance the accuracy of image sound retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086406A_ABST
    Figure CN120086406A_ABST
Patent Text Reader

Abstract

The invention provides an image and sound retrieval method and device, equipment and a storage medium, and the method comprises the steps: obtaining to-be-retrieved data, the to-be-retrieved data comprises to-be-retrieved voice data and / or to-be-retrieved image data; inputting the to-be-retrieved voice data into the image and sound retrieval model to obtain image data related to the to-be-retrieved voice data, and / or inputting the to-be-retrieved image data into the image and sound retrieval model to obtain voice data related to the to-be-retrieved image data; wherein the image and voice retrieval model is obtained by training multiple pieces of voice image pair data and multiple pieces of text data related to the multiple pieces of voice image pair data, and any target voice image pair data in the multiple pieces of voice image pair data comprises target voice data and target image data which are consistent in corresponding content, so that the accuracy of data retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to an image-audio retrieval method, device, equipment, and storage medium. Background Art

[0002] Image-audio retrieval refers to analyzing and understanding the correlation between audio and image data to retrieve another modality data (e.g., speech modality data or image modality data) that matches a given modality data (e.g., image modality data or speech modality data). For example, for a given audio clip, image-audio retrieval can find the corresponding image, or for a given image, image-audio retrieval can find the relevant audio. Currently, the similarity between the features of images and speech is often calculated for matching. However, this method has the problem of low retrieval accuracy. Summary of the Invention

[0003] The present application provides an image-audio retrieval method, device, equipment, and storage medium, which can improve the accuracy of data retrieval.

[0004] In a first aspect, an image-audio retrieval method is provided, including: obtaining data to be retrieved, where the data to be retrieved includes speech data to be retrieved and / or image data to be retrieved; inputting the speech data to be retrieved into an image-audio retrieval model to obtain image data related to the speech data to be retrieved, and / or inputting the image data to be retrieved into the image-audio retrieval model to obtain speech data related to the image data to be retrieved; wherein the image-audio retrieval model is trained by a plurality of speech-image pair data and a plurality of text data related to the plurality of speech-image pair data, and any target speech-image pair data in the plurality of speech-image pair data includes: target speech data and target image data with consistent corresponding content.

[0005] In a second aspect, an image-audio retrieval device is provided, including: a first acquisition module, configured to obtain data to be retrieved, where the data to be retrieved includes speech data to be retrieved and / or image data to be retrieved; an image-audio retrieval module, configured to input the speech data to be retrieved into an image-audio retrieval model to obtain image data related to the speech data to be retrieved, and / or input the image data to be retrieved into the image-audio retrieval model to obtain speech data related to the image data to be retrieved; wherein the image-audio retrieval model is trained by a plurality of speech-image pair data and a plurality of text data related to the plurality of speech-image pair data, and any target speech-image pair data in the plurality of speech-image pair data includes: target speech data and target image data with consistent corresponding content.

[0006] In a third aspect, an electronic device is provided, including: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementation manners.

[0007] In a fourth aspect, a computer-readable storage medium is provided, which is used to store a computer program, and the computer program causes a computer to execute the method in the first aspect or its various implementation manners.

[0008] In a fifth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions cause a computer to execute the method in the first aspect or its various implementation manners.

[0009] In a sixth aspect, a computer program is provided, and the computer program causes a computer to execute the method in the first aspect or its various implementation manners.

[0010] In the technical solution of this application, image data related to the speech data to be retrieved and / or speech data related to the image data to be retrieved can be determined through the speech-image retrieval model. The speech-image retrieval model is trained based on speech-image pair data (including speech data and image data with consistent corresponding contents) and the related text data. Therefore, when determining the image data related to the speech data to be retrieved and / or the speech data related to the image data to be retrieved through the speech-image retrieval model, not only the speech data to be retrieved and / or the image data to be retrieved are considered, but also the related text data is considered, which can achieve more accurate matching and alignment between the image data and the speech data, thereby improving the accuracy of speech-image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings required for use in the following description of the embodiments will be introduced below.

[0012] Figure 1 It is a flowchart of a speech-image retrieval method provided by an embodiment of this application;

[0013] Figure 2 It is a schematic diagram of a speech-image retrieval method provided by an embodiment of this application;

[0014] Figure 3 It is a schematic diagram of another speech-image retrieval method provided by an embodiment of this application;

[0015] Figure 4 It is a schematic diagram of yet another speech-image retrieval method provided by an embodiment of this application;

[0016] Figure 5 It is a schematic diagram of a speech-image retrieval device 500 provided by an embodiment of this application;

[0017] Figure 6It is a schematic diagram of an electronic device 600 provided by an embodiment of the present application. Detailed implementation manners

[0018] Next, the technical solution of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application.

[0019] It should be noted that the information, data (including, but not limited to: data for analysis, stored data, displayed data, etc., such as voice data, image data, and text data, etc.) and signals involved in the present application are all authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the data to be retrieved and the operations performed on the data to be retrieved involved in the present application are all obtained under full authorization.

[0020] In one embodiment, the technical solution of the present application can be used in the scenario of image-audio retrieval. Specifically, it can be applied to the scenario of retrieving image data using voice data or retrieving voice data using image data, but not limited thereto.

[0021] In one embodiment, the solution provided by the present application can be executed by any electronic device with data processing capabilities. For example, the electronic device can be a server, specifically an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Again, the electronic device can be a terminal device, specifically a tablet computer, a laptop computer, or a desktop computer, etc. Further, the electronic device can be a combination of a server and a terminal device, where the server and the terminal device in the combination can communicate wirelessly or wiredly, and the present application does not make specific limitations on the electronic device.

[0022] Figure 1 It is a flowchart of an image-audio retrieval method provided by an embodiment of the present application. This method can be executed by the electronic device in the above embodiment, as Figure 1 shown, this method includes:

[0023] S110: Obtain the data to be retrieved, where the data to be retrieved includes the voice data to be retrieved and / or the image data to be retrieved;

[0024] S120: Input the voice data to be retrieved into the image-audio retrieval model to obtain the image data related to the voice data to be retrieved, and / or input the image data to be retrieved into the image-audio retrieval model to obtain the voice data related to the image data to be retrieved.

[0025] Among them, the graphophone retrieval model is trained by multiple voice-image pair data and multiple text data related to the multiple voice-image pair data. Any target voice-image pair data in the multiple voice-image pair data includes: target voice data and target image data with consistent corresponding content.

[0026] Among them, if the data to be retrieved includes the voice data to be retrieved, the voice data to be retrieved can be input into the graphophone retrieval model to obtain the image data related to the voice data to be retrieved; if the data to be retrieved includes the image data to be retrieved, the image data to be retrieved is input into the graphophone retrieval model to obtain the voice data related to the image data to be retrieved; if the data to be retrieved includes the voice data to be retrieved and the image data to be retrieved, the voice data to be retrieved is input into the graphophone retrieval model to obtain the image data related to the voice data to be retrieved, and the image data to be retrieved is input into the graphophone retrieval model to obtain the voice data related to the image data to be retrieved.

[0027] It should be noted that the voice-image pair data in this application refers to a pair of voice data and image data with consistent corresponding content. The voice data and image data with consistent corresponding content can be understood as: the content represented by the voice data is the same as the content represented by the image data, or the voice data is used to describe the image data, or the image content shown in the image data is the same as the information conveyed by the voice data.

[0028] Exemplarily, assuming the image data is an image of a white dog carrying a stick in its mouth on a beach, the voice data with consistent corresponding content to the image data can be the Chinese voice "A white dog on a beach is carrying a stick in its mouth" or the English voice "A white dog on a beach is carrying a stick in its mouth."

[0029] In addition, this application does not limit the languages of the voice data and text data involved.

[0030] In one embodiment, the user can input or select the data to be retrieved based on the client. The client can obtain the data to be retrieved and send it to the electronic device, so that the electronic device can obtain the data to be retrieved.

[0031] Next, the training process of the graphophone retrieval model will be introduced first:

[0032] In one embodiment, the training process of the graphophone retrieval model may include the following steps:

[0033] S1: Obtain multiple voice-image pair data.

[0034] S2: For the target voice-image pair data, obtain the voice features of the target voice data and the image features of the target image data through the graph-audio retrieval model.

[0035] Among them, the voice features include: voice global features and / or voice local features; the image features include: image global features and / or image local features.

[0036] Specifically, the audio encoder of the graph-audio retrieval model is used to extract features from the target voice data to obtain the voice features of the target voice data; the image encoder of the graph-audio retrieval model is used to extract features from the target image data to obtain the image features of the target image data.

[0037] Exemplarily, the audio encoder may include HuBERT and Transformer. Among them, HuBERT is an audio feature extraction model based on self-supervised learning, which can learn the target voice data by predicting the quantization representation of masked audio segments and extract the feature representation of the target voice data. Transformer is a neural network structure based on the self-attention mechanism, which can further process the features extracted by HuBERT to capture the temporal and context information in the target voice data.

[0038] For example, the voice features may include at least one of voice global features and voice local features. Or, the voice features may also be features obtained by splicing the voice global features and the voice local features. For example, they may be features obtained by splicing the voice global features and the voice local features at the beginning and end or by weighted splicing. Among them, the voice global features may be at least one of the vector corresponding to the average energy of the entire voice segment of the target voice data, the vector corresponding to the fundamental frequency, the vector corresponding to the spectral envelope, the Mel Frequency Cepstrum Coefficient (MFCC) feature, the log-mel, and the Linear Predictive Coding (LPC). Similarly, the voice local features may be at least one of the vector corresponding to the average energy of each voice segment in the target voice data, the vector corresponding to the fundamental frequency, the vector corresponding to the spectral envelope, the MFCC feature, the log-mel, and the LPC.

[0039] Exemplarily, the image encoder may be an encoder based on Contrastive Language–Image Pre-training (CLIP).

[0040] For example, the image features may include at least one of global image features and local image features. Alternatively, the image features may also be features obtained by splicing the global image features and the local image features. For example, they may be features obtained by splicing the global image features and the local image features end to end or by weighted splicing. The global image features may be at least one of the color features of the entire image of the target image data (specifically, vectors corresponding to the color distributions in the image), texture features (specifically, gray-level co-occurrence matrices), shape features (specifically, edge features and contour features of each object in the image), and spatial relationship features (specifically, vectors or matrices corresponding to the relative position relationships between different objects or regions in the image). Similarly, the local image features may be at least one of the color features of each image block in the target image data (specifically, vectors corresponding to the color distributions in the image block), texture features (specifically, gray-level co-occurrence matrices of the image block), shape features (specifically, edge features and contour features of each object in the image block), and spatial relationship features (specifically, vectors or matrices corresponding to the relative position relationships between different objects or regions in the image block).

[0041] S3: Generate target text data related to the target speech-image pair data according to the target speech-image pair data.

[0042] Specifically, the target text data may be generated according to the target speech data in the target speech-image pair data. Alternatively, the target text data may also be generated by describing the target image data in the target speech-image pair data based on a large model.

[0043] Exemplarily, the target speech data may be subjected to speech recognition to obtain the target text data.

[0044] For example, the target speech data may be subjected to speech recognition based on Automatic Speech Recognition (ASR) to obtain the target text data.

[0045] For example, the probability distribution of each vector in the speech features of the target speech data corresponding to each character in the dictionary may be determined; then, according to the probability distribution, the speech features are mapped into multiple phonemes, where the phonemes include characters or words; and the multiple phonemes are combined to obtain the target text data.

[0046] Exemplarily, the graph-audio retrieval model may further include a text processing module, and the text processing module may generate the target text data according to the target speech data, that is, the target text data may be generated through the text processing module of the graph-audio retrieval model according to the target speech data.

[0047] In the above embodiments, by performing speech recognition on the speech data in the speech-image pair data to obtain text data, not only can the efficiency of generating text data be improved, but also a high correlation between the text data and the speech data can be ensured.

[0048] S4: Train the graph-audio retrieval model according to the speech features, image features, and target text data.

[0049] Specifically, S4 includes the following steps:

[0050] S4-1: Determine the first loss according to the speech features and image features;

[0051] S4-2: Determine the second loss according to the target text data and the target speech-image pair data;

[0052] S4-3: Train the graph-audio retrieval model according to the first loss and / or the second loss.

[0053] The following will introduce S4-1, S4-2, and S4-3 in turn:

[0054] In one embodiment, S4-1 can be implemented through the following steps, that is, determine the first loss according to the speech features and image features:

[0055] S4-1-1: For the speech features, calculate the similarity degree value of the speech features relative to the image features to obtain the audio-image similarity.

[0056] Specifically, for the speech features, based on the image features of each image data in multiple speech-image pair data, calculate the similarity degree value of the speech features relative to the image features to obtain the audio-image similarity.

[0057] Exemplarily, for the speech features, calculate the first sub-audio-image similarity between the speech features and the image features, determine the sum of the similarities between the speech features and the image features of each image data in multiple speech-image pair data to obtain the second sub-audio-image similarity, and calculate the ratio of the first sub-audio-image similarity and the second sub-audio-image similarity to obtain the audio-image similarity.

[0058] Among them, the first sub-audio-image similarity can be obtained by calculating the vector inner product corresponding to the speech features and the image features; correspondingly, the similarity between the speech features and the image features of each image data in multiple speech-image pair data is: the vector inner product corresponding to the speech features and the image features of each image data in multiple speech-image pair data.

[0059] S4-1-2: For the image features, calculate the similarity degree value of the image features relative to the speech features to obtain the image-audio similarity.

[0060] Specifically, for image features, based on the speech features of each speech data in multiple speech-image pair data, the similarity degree value of the image features relative to the speech features can be calculated to obtain the image-speech similarity.

[0061] Exemplarily, for image features, calculate the first sub-image-speech similarity between the image features and the speech features, determine the sum of the similarities between the image features and the speech features of each speech data in the multiple speech-image pair data to obtain the second sub-image-speech similarity, and calculate the ratio of the first sub-image-speech similarity to the second sub-image-speech similarity to obtain the image-speech similarity.

[0062] Among them, the vector inner product corresponding to the image features and the speech features can be calculated to obtain the first sub-image-speech similarity; correspondingly, the similarity between the image features and the speech features of each speech data in the multiple speech-image pair data is: the vector inner product corresponding to the image features and the speech features of each speech data in the multiple speech-image pair data.

[0063] S4-1-3: Determine the first sub-loss according to the speech-image similarity, and determine the second sub-loss according to the image-speech similarity.

[0064] Specifically, the speech-image similarity can be calculated based on the logarithmic function to obtain the initial first sub-loss corresponding to the target speech-image pair data; sum the initial first sub-losses corresponding to each of the multiple speech-image pair data to obtain the first sub-loss; the image-speech similarity can be calculated based on the logarithmic function to obtain the initial second sub-loss corresponding to the target speech-image pair data; sum the initial second sub-losses corresponding to each of the multiple speech-image pair data to obtain the second sub-loss.

[0065] S4-1-4: Determine the first loss according to the first sub-loss and the second sub-loss.

[0066] Specifically, the average value of the first sub-loss and the second sub-loss can be calculated to obtain the first loss.

[0067] In one embodiment, assume that the speech characteristic is the speech global feature The image feature is the global feature of the image The first loss can be the contrastive loss between the speech global feature and the image global feature.

[0068] First, the similarity between the i-th speech data and the j-th image data in the N speech-image pair data can be calculated through formula (1) (Specifically, the vector inner product of the corresponding features); calculate the similarity between the i-th image data and the j-th speech data through formula (2) (Specifically, it is the inner product of the vectors of the corresponding features), where N is a positive integer, i = 1,......, N, and j = 1,......, N.

[0069]

[0070] Next, the contrast loss function L between the i-th speech data and the i-th image data can be determined according to formula (3) based on the above similarity, through the speech-to-image similarity between the i-th speech data and the i-th image data. s2i , that is, the first sub-loss; the contrast loss function L between the i-th image data and the i-th speech data can be determined according to formula (4) through the image-to-speech similarity between the i-th image data and the i-th speech data. i2s , that is, the second sub-loss, where τ is a parameter.

[0071]

[0072]

[0073] Finally, the contrast loss L between the speech feature and the image feature can be determined through formula (5). sic , that is, the first loss.

[0074]

[0075] In the above embodiments, the accuracy of aligning the speech feature and the image feature can be improved multi-dimensionally and multi-levelly through the similarity of the image data relative to the speech data and the similarity of the speech data relative to the image data (similarities from two perspectives) in the speech-image pair data. Specifically, it is to better shorten the distance between positive samples (such as the i-th speech data and the i-th image data, that is, the image data and the speech data with corresponding consistent content), and increase the differentiation between negative samples (such as the i-th speech data and the j-th image data, i≠j, that is, the image data and the speech data with corresponding inconsistent content), so as to better realize the alignment and matching between the image data and the speech data, and further improve the accuracy of model training and the accuracy of image-audio retrieval.

[0076] In one embodiment, S4-2 can be implemented through the following steps, that is, determining the second loss according to the target text data and the target speech-image pair data:

[0077] S4-2-1: Determine the second loss according to the target speech data in the target text data and the target speech-image pair data.

[0078] Exemplarily, based on the speech features of the target text data and the target speech data, a target probability distribution can be determined, which is used to represent the probability of each vector in the speech features corresponding to each character in the dictionary; then, based on the target probability distribution and the target text data, a second loss is determined.

[0079] For example, as shown in formula (6), through a linear layer, a linear mapping can be performed on the target text data and the speech features to obtain a linear output Linear dm →|V|(S) (where the speech features can be frame-level features, that is, the speech local features S = {S1, S2,..., Sn}); through a normalization layer, the linear output is normalized to obtain the target probability distribution q.

[0080] q = Softmax(Linear dm →|V|(S)) Formula (6)

[0081] Among them, the graph-audio retrieval model further includes a text processing module, and the text processing module includes: a linear layer and a normalization layer. That is, the above target probability distribution can be calculated through the text processing module of the graph-audio retrieval model.

[0082] For example, as shown in formula (7), based on the Connectionist Temporal Classification (CTC) loss function, according to the target probability distribution and the target text data Y, the CTC loss L between the speech features and the text data can be determined ctc , that is, the second loss.

[0083]

[0084] In addition, the second loss can also be determined based on the target text data and the target image data. The content is similar to the process of determining the second loss based on the target text data and the target speech data above.

[0085] In the above embodiments, through the speech recognition auxiliary task of audio-to-text transcription (corresponding to generating text data from speech data and determining the second loss accordingly), not only can the ability to extract semantic features, that is, text features, from speech data be improved, and the ability to extract fine-grained semantic information be enhanced, thereby further assisting in improving the graph-audio retrieval performance, but also the corresponding relationship between images and speech can be implicitly connected through the generated text data, achieving better alignment between images and audio, and improving the accuracy of model training and the accuracy of graph-audio retrieval.

[0086] In one embodiment, S4-3 can be performed through the following steps, that is, training the graph-audio retrieval model according to the first loss and / or the second loss: determining a balance factor; calculating the product of the balance factor and the second loss to obtain a weighted second loss; calculating the sum of the weighted second loss and the first loss to obtain a third loss; and training the graph-audio retrieval model according to the third loss.

[0087] For example, as shown in formula (8), the contrast loss L between the speech feature and the image feature sic and the CTC loss L between the speech feature and the text data ctc of the weighted sum can be determined as the third loss, γ is the harmonic coefficient, that is, the balance factor, and the balance factor can be any value.

[0088] L=(L sic +γL ctc ) Formula (8)

[0089] Specifically, the training of the graph-audio retrieval model can be completed when the third loss, the second loss, or the first loss is less than a preset loss threshold.

[0090] Among them, the above training process is to narrow the distance between positive samples (such as the i-th speech data and the i-th image data), and at the same time increase the differentiation between negative samples (such as the i-th speech data and the j-th image data, i≠j). The i-th speech data and the i-th image data belong to the same speech-image pair data, that is, the i-th speech-image pair data.

[0091] In the above embodiment, the comprehensive and weighted consideration of the first loss and the second loss can be realized through the balance factor, that is, the comprehensive and weighted consideration of the respective corresponding losses between the image data and the speech data and between the text data and the speech data can be realized, and then a more reasonable final loss, that is, the third loss, can be determined, thereby improving the accuracy of model training and the accuracy of graph-audio retrieval.

[0092] In one embodiment, as Figure 2 shown, the training data for training the graph-audio retrieval model consists of speech-image pair data (image-speech pairs), where the speech content of the speech data is a description of the image information of the image data. An automatic speech recognition model (for example, ASR module) can be used to perform automatic speech recognition on the speech data to generate text data for assisting the graph-audio retrieval model.

[0093] Combined with the above content, the system architecture diagram of the graph-audio retrieval model is as Figure 3As shown in the figure, a multi-task training framework can be used to train the image-audio retrieval model. The loss function includes an image-audio contrast loss function, i.e., the first loss, and a speech recognition loss function, i.e., the second loss. The image-audio retrieval model can include: an image encoder, a speech encoder, and a speech recognition module (which can be the above-mentioned speech recognition model, i.e., the text processing module, including a Linear layer and a Softmax normalization layer). Among them, the image encoder uses the CLIP image encoder model to extract high-dimensional image features from the input image data. The parameters of this module are frozen during the training process and not updated. The speech encoder can extract high-dimensional audio features from the input speech data, including a frozen Hubert model and several learnable Transformer encoding layers. In addition, the speech recognition module can transcribe the frame-level speech features into a text sequence to obtain text data.

[0094] The retrieval method based on the image-audio retrieval model will be further introduced below. The corresponding content and effects between the retrieval method and the above model training method can be referred to each other.

[0095] In one embodiment, the above step of inputting the speech data to be retrieved into the image-audio retrieval model to obtain the image data related to the speech data to be retrieved may include: determining the text data related to the speech data to be retrieved through the image-audio retrieval model, and determining the image data related to the speech data to be retrieved from multiple preset image data according to the speech data to be retrieved and the text data related to the speech data to be retrieved.

[0096] Specifically, through the image-audio retrieval model, the similarity between the multiple preset image data and the text data to be retrieved (specifically, the similarity between corresponding features) and the similarity between the speech data to be retrieved and the multiple preset image data (specifically, the similarity between corresponding features) can be calculated; the image data related to the speech data to be retrieved is determined from the multiple preset image data according to the above similarities. That is to say, the final retrieval result (the image data related to the speech data to be retrieved, and / or, the speech data related to the image data to be retrieved) can be determined by calculating two similarities (the similarity between the image data and the speech data, and the similarity between the image data and the text data). Among them, the similarity between the image data and the speech data can implicitly connect the corresponding relationship between the image and the speech, realize better matching between the image and the speech, and thus improve the performance of image-audio cross-modal retrieval.

[0097] Inputting the image data to be retrieved into the image-audio retrieval model to obtain the audio data related to the image data to be retrieved may include: determining, through the image-audio retrieval model, the text data related to multiple preset audio data, and determining, from the multiple preset audio data, the audio data related to the image data to be retrieved according to the image data to be retrieved and the text data related to the multiple preset audio data.

[0098] Specifically, through the image-audio retrieval model, the similarity between the image data to be retrieved and the text data (specifically, the similarity between corresponding features) and the similarity between the image data to be retrieved and the multiple preset audio data (specifically, the similarity between corresponding features) can be calculated; and the audio data related to the image data to be retrieved is determined from the multiple preset audio data according to the above similarities.

[0099] Exemplarily, the image-audio retrieval model includes: an audio encoder, an image encoder, a text processing module, a matching value calculation module, and an output module. Among them:

[0100] Determining, through the image-audio retrieval model, the text data related to the audio data to be retrieved, and determining, from the multiple preset image data, the image data related to the audio data to be retrieved according to the audio data to be retrieved and the text data related to the audio data to be retrieved includes:

[0101] Through the audio encoder, feature extraction is performed on the audio data to be retrieved to obtain the audio features to be retrieved; through the image encoder, feature extraction is performed on the multiple preset image data to obtain the respective preset image features of the multiple preset image data; through the text processing module, speech recognition is performed on the audio data to be retrieved to obtain the text data; through the text feature extraction module (which can be a text encoder based on CLIP. Additionally, it can be a module in the image-audio retrieval model or not), feature extraction is performed on the text data to obtain the text features; through the matching value calculation module, according to the multiple preset image features, the audio features to be retrieved, and the text features, the matching value between the multiple preset image features and the audio features to be retrieved is determined; through the output module, according to the matching value between the multiple preset image features and the audio features to be retrieved, the image data related to the audio data to be retrieved is determined and output from the multiple preset image data.

[0102] Determining, through the image-audio retrieval model, the text data related to the multiple preset audio data, and determining, from the multiple preset audio data, the audio data related to the image data to be retrieved according to the image data to be retrieved and the text data related to the multiple preset audio data includes:

[0103] Through an audio encoder, perform feature extraction on multiple preset speech data to obtain the preset speech features of each of the multiple preset speech data; through an image encoder, perform feature extraction on the image data to be retrieved to obtain the image feature to be retrieved; through a text processing module, perform speech recognition on the multiple preset speech data respectively to obtain multiple text data; through a text feature extraction module, perform feature extraction on the multiple text data respectively to obtain multiple text features; through a matching value calculation module, determine the matching value between the multiple preset speech features and the image feature to be retrieved according to the multiple preset speech features, the image feature to be retrieved, and the multiple preset text features; through an output module, determine and output the speech data related to the image data to be retrieved from the multiple preset speech data according to the matching value between the multiple preset speech features and the image feature to be retrieved.

[0104] For example, the first matching value between each of the multiple preset image features and the speech feature to be retrieved can be calculated; the second matching value between each of the multiple preset image features and the text feature can be calculated; according to the multiple first matching values and the multiple second matching values, determine the matching value between the multiple preset image features and the speech feature to be retrieved. Specifically, the sum of the corresponding multiple first matching values and the multiple second matching values can be determined as the matching value between the multiple preset image features and the speech feature to be retrieved.

[0105] Similarly, the third matching value between each of the multiple preset speech features and the image feature to be retrieved can be calculated; the fourth matching value between each of the multiple preset text features and the image feature to be retrieved can be calculated; according to the multiple third matching values and the multiple fourth matching values, determine the matching value between the multiple preset speech features and the image feature to be retrieved. Specifically, the sum of the corresponding multiple third matching values and the multiple fourth matching values can be determined as the matching value between the multiple preset speech features and the image feature to be retrieved.

[0106] Among them, the above matching values can all be determined through similarity algorithms, such as, but not limited to, cosine similarity, vector dot product, Euclidean distance, Manhattan distance, Jaccard similarity coefficient, Pearson correlation coefficient.

[0107] It should be noted that the image-audio retrieval model can also be understood as or referred to as an image-audio retrieval system. In addition, the above division of each unit or module in the image-audio retrieval model is only an exemplary logical function division, and there can be other division methods, which are not limited in this application.

[0108] Exemplarily, the above audio encoder and image encoder can be divided into the same module: an encoder module, which is used to execute the corresponding content of the above audio encoder and image encoder. For example, through the encoder module, feature extraction can be performed on the speech data to be retrieved to obtain the speech features to be retrieved, and feature extraction can be performed on multiple preset image data to obtain the preset image features of each of the multiple preset image data.

[0109] Alternatively, the matching value calculation module and the output module can also be divided into the same module: a matching output module, which is used to execute the corresponding content of the above matching value calculation module and output module. For example, through the matching output module, according to multiple preset image features, the speech features to be retrieved, and text features, the matching value between the multiple preset image features and the speech features to be retrieved can be determined, and based on the matching value between the multiple preset image features and the speech features to be retrieved, the image data related to the speech data to be retrieved can be determined and output from the multiple preset image data.

[0110] In one embodiment, as Figure 4 shown, the graphophone retrieval performance and the graphophone retrieval accuracy can be improved through an integration mechanism. First, the audio encoder and the image encoder of the graphophone retrieval model can be used to perform feature extraction on speech data and image data respectively (if the data to be retrieved is speech data, the image data can be multiple preset image data, that is, the image data prepared in the database to be matched with the speech data; if the data to be retrieved is image data, the speech data can be multiple preset speech data, that is, the speech data prepared in the database to be matched with the image data) to obtain the speech representation Scls and the image representation Icls under the global feature; then, the speech-image match score, that is, the matching value between the speech data and the image data, can be calculated through the cosine similarity according to the above global vectors; in addition, through the ASR module, the probability distribution of the frame-level speech representation and the speech local features in the speech features can be calculated, and it can be decoded to obtain the text sequence, that is, the text data. Then, the text encoder based on CLIP is used to extract the text data to obtain the text representation (which can be under the global feature); then, the matching value between the text representation Tcls and the image representation Icls is calculated to obtain the image-text match score; finally, the sum of these two scores can be used as the final similarity score, and the graphophone retrieval result can be obtained according to the high and low similarity scores. For example, the speech data or image data corresponding to the highest similarity score can be determined as the retrieval result.

[0111] Through the technical solution of this application, text data related to multiple voice-image pair data can be used to assist in the training and use of the graph-audio retrieval model, thereby reducing the gap between the image and speech semantic spaces, and improving the accuracy of graph-audio retrieval.

[0112] Among them, in the model training stage, in addition to adopting the image-audio alignment task (corresponding to determining the first loss according to the image data and speech data, specifically the speech feature and the image feature), it also includes the speech recognition auxiliary task of transcribing speech into text (corresponding to generating text data according to the speech data and determining the second loss), which can improve the ability of the audio encoder to extract fine-grained semantic information, enhance the ability of the speech encoder to extract semantic features from the speech modality, and implicitly connect the corresponding relationship between the image and the speech, realize better alignment between the image and the audio, improve the accuracy between the aligned speech and image feature spaces, and improve the performance of image-audio cross-modal retrieval.

[0113] In the model usage stage, the speech data can be decoded into text through the speech recognition module. By introducing an integration mechanism, the graph-audio matching score, that is, the matching value between the image feature and the speech feature, and the graph-text matching score, that is, the matching value between the image feature and the text feature, are fused to further improve the graph-audio retrieval accuracy.

[0114] It should be noted that all the above technical solutions can be combined arbitrarily to form alternative embodiments of this application, which will not be elaborated one by one here.

[0115] Figure 5 The following is a schematic diagram of a graph-audio retrieval device 500 provided by an embodiment of this application, as Figure 5 shown. The device 500 includes: a first acquisition module 501, a graph-audio retrieval module 502, a second acquisition module 503, a feature extraction module 504, a text generation module 505, and a model training module 506.

[0116] In one embodiment, the first acquisition module 501 is configured to acquire data to be retrieved, and the data to be retrieved includes data to be retrieved speech data and / or data to be retrieved image data; the graph-audio retrieval module 502 is configured to input the data to be retrieved speech data into the graph-audio retrieval model to obtain image data related to the data to be retrieved speech data, and / or input the data to be retrieved image data into the graph-audio retrieval model to obtain speech data related to the data to be retrieved image data; among them, the graph-audio retrieval model is trained by multiple voice-image pair data and multiple text data related to the multiple voice-image pair data, and any target voice-image pair data in the multiple voice-image pair data includes: target speech data and target image data with consistent corresponding contents.

[0117] Exemplarily, the second acquisition module 503 is configured to acquire a plurality of speech-image pair data; the feature extraction module 504 is configured to, for the target speech-image pair data, obtain the speech feature of the target speech data and the image feature of the target image data through the speech-image retrieval model; the text generation module 505 is configured to generate target text data related to the target speech-image pair data according to the target speech-image pair data; the model training module 506 is configured to train the speech-image retrieval model according to the speech feature, the image feature, and the target text data; wherein, the speech feature includes: speech global feature and / or speech local feature; the image feature includes: image global feature and / or image local feature.

[0118] Exemplarily, the feature extraction module 504 is specifically configured to: perform feature extraction on the target speech data through the audio encoder of the speech-image retrieval model to obtain the speech feature of the target speech data; perform feature extraction on the target image data through the image encoder of the speech-image retrieval model to obtain the image feature of the target image data.

[0119] Exemplarily, the model training module 506 is specifically configured to: determine a first loss according to the speech feature and the image feature; determine a second loss according to the target text data and the target speech-image pair data; train the speech-image retrieval model according to the first loss and / or the second loss.

[0120] Exemplarily, the model training module 506 is specifically configured to: for the speech feature, calculate the similarity degree value of the speech feature relative to the image feature to obtain the speech-image similarity; for the image feature, calculate the similarity degree value of the image feature relative to the speech feature to obtain the image-speech similarity; determine a first sub-loss according to the speech-image similarity, and determine a second sub-loss according to the image-speech similarity; determine a first loss according to the first sub-loss and the second sub-loss.

[0121] Exemplarily, the model training module 506 is specifically configured to: for the speech feature, calculate the similarity degree value of the speech feature relative to the image feature based on the image feature of each image data in the plurality of speech-image pair data to obtain the speech-image similarity; for the image feature, calculate the similarity degree value of the image feature relative to the speech feature based on the speech feature of each speech data in the plurality of speech-image pair data to obtain the image-speech similarity.

[0122] Exemplarily, the model training module 506 is specifically configured to: for the speech features, calculate the first sub-audio-graph similarity between the speech features and the image features, determine the sum of the similarities between the speech features and the image features of each image data in the multiple speech-image pair data to obtain the second sub-audio-graph similarity, and calculate the ratio of the first sub-audio-graph similarity to the second sub-audio-graph similarity to obtain the audio-graph similarity; for the image features, calculate the first sub-graph-audio similarity between the image features and the speech features, determine the sum of the similarities between the image features and the speech features of each speech data in the multiple speech-image pair data to obtain the second sub-graph-audio similarity, and calculate the ratio of the first sub-graph-audio similarity to the second sub-graph-audio similarity to obtain the graph-audio similarity.

[0123] Exemplarily, the model training module 506 is specifically configured to: calculate the vector inner product corresponding to the speech features and the image features to obtain the first sub-audio-graph similarity; correspondingly, the similarity between the speech features and the image features of each image data in the multiple speech-image pair data is: the vector inner product corresponding to the speech features and the image features of each image data in the multiple speech-image pair data; calculate the vector inner product corresponding to the image features and the speech features to obtain the first sub-graph-audio similarity; correspondingly, the similarity between the image features and the speech features of each speech data in the multiple speech-image pair data is: the vector inner product corresponding to the image features and the speech features of each speech data in the multiple speech-image pair data.

[0124] Exemplarily, the model training module 506 is specifically configured to: calculate the audio-graph similarity based on the logarithmic function to obtain the initial first sub-loss corresponding to the target speech-image pair data; sum up the initial first sub-losses corresponding to the multiple speech-image pair data respectively to obtain the first sub-loss; determine the second sub-loss according to the graph-audio similarity, including: calculate the graph-audio similarity based on the logarithmic function to obtain the initial second sub-loss corresponding to the target speech-image pair data; sum up the initial second sub-losses corresponding to the multiple speech-image pair data respectively to obtain the second sub-loss.

[0125] Exemplarily, the model training module 506 is specifically configured to: calculate the average value of the first sub-loss and the second sub-loss to obtain the first loss. Exemplarily, the model training module 506 is specifically configured to: determine the second loss according to the target text data and the target speech data in the target speech-image pair data.

[0126] Exemplarily, the model training module 506 is specifically configured to: determine the target probability distribution according to the target text data and the speech features of the target speech data, where the target probability distribution is used to represent the probability that each vector in the speech features corresponds to each character in the dictionary; determine the second loss according to the target probability distribution and the target text data.

[0127] Exemplarily, the model training module 506 is specifically configured to: perform a linear mapping on the target text data and the speech features through a linear layer to obtain a linear output; perform a normalization process on the linear output through a normalization layer to obtain a target probability distribution.

[0128] Exemplarily, the graph-audio retrieval model further includes a text processing module, and the text processing module includes: a linear layer and a normalization layer.

[0129] Exemplarily, the model training module 506 is specifically configured to: determine a second loss based on the connectionist temporal classification loss function according to the target probability distribution and the target text data.

[0130] Exemplarily, the model training module 506 is specifically configured to: determine a balance factor; calculate the product of the balance factor and the second loss to obtain a weighted second loss; calculate the sum of the weighted second loss and the first loss to obtain a third loss; and train the graph-audio retrieval model according to the third loss.

[0131] Exemplarily, the text generation module 505 is specifically configured to: generate target text data according to the target speech data in the target speech image pair data.

[0132] Exemplarily, the text generation module 505 is specifically configured to: perform speech recognition on the target speech data to obtain target text data.

[0133] Exemplarily, the text generation module 505 is specifically configured to: determine the probability distribution of each vector in the speech features of the target speech data corresponding to each character in the dictionary; map the speech features into multiple phonemes according to the probability distribution, where the phonemes include characters or words; and combine the multiple phonemes to obtain target text data.

[0134] Exemplarily, the text generation module 505 is specifically configured to: generate target text data according to the target speech data through the text processing module of the graph-audio retrieval model.

[0135] Exemplarily, the graph-audio retrieval module 502 is specifically configured to: determine text data related to the speech data to be retrieved through the graph-audio retrieval model, and determine image data related to the speech data to be retrieved from multiple preset image data according to the speech data to be retrieved and the text data related to the speech data to be retrieved; or, determine text data related to multiple preset speech data through the graph-audio retrieval model, and determine speech data related to the image data to be retrieved from multiple preset speech data according to the image data to be retrieved and the text data related to the multiple preset speech data.

[0136] Exemplarily, the image-audio retrieval model includes: an audio encoder, an image encoder, a text processing module, a matching value calculation module, and an output module; the image-audio retrieval module 502 is specifically configured to: extract features from the speech data to be retrieved through the audio encoder to obtain the speech features to be retrieved; extract features from multiple preset image data through the image encoder to obtain the respective preset image features of the multiple preset image data; perform speech recognition on the speech data to be retrieved through the text processing module to obtain text data; extract features from the text data through the text feature extraction module to obtain text features; determine the matching values between the multiple preset image features and the speech features to be retrieved according to the multiple preset image features, the speech features to be retrieved, and the text features through the matching value calculation module; determine and output the image data related to the speech data to be retrieved from the multiple preset image data according to the matching values between the multiple preset image features and the speech features to be retrieved through the output module; or extract features from multiple preset speech data through the audio encoder to obtain the respective preset speech features of the multiple preset speech data; extract features from the image data to be retrieved through the image encoder to obtain the image features to be retrieved; perform speech recognition on the multiple preset speech data respectively through the text processing module to obtain multiple text data; extract features from the multiple text data respectively through the text feature extraction module to obtain multiple text features; determine the matching values between the multiple preset speech features and the image features to be retrieved according to the multiple preset speech features, the image features to be retrieved, and the multiple preset text features through the matching value calculation module; determine and output the speech data related to the image data to be retrieved from the multiple preset speech data according to the matching values between the multiple preset speech features and the image features to be retrieved through the output module.

[0137] Exemplarily, the image-audio retrieval module 502 is specifically configured to: calculate the first matching values between the multiple preset image features and the speech features to be retrieved respectively; calculate the second matching values between the multiple preset image features and the text features respectively; determine the matching values between the multiple preset image features and the speech features to be retrieved according to the multiple first matching values and the multiple second matching values. Or calculate the third matching values between the multiple preset speech features and the image features to be retrieved respectively; calculate the fourth matching values between the multiple preset text features and the image features to be retrieved respectively; determine the matching values between the multiple preset speech features and the image features to be retrieved according to the multiple third matching values and the multiple fourth matching values.

[0138] Exemplarily, the image-audio retrieval module 502 is specifically configured to: determine the sum of a plurality of first matching values corresponding to a plurality of second matching values as the matching value between a plurality of preset image features and the speech feature to be retrieved; or determine the sum of a plurality of third matching values corresponding to a plurality of fourth matching values as the matching value between a plurality of preset speech features and the image feature to be retrieved.

[0139] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 5 The illustrated device 500 can execute the above method embodiments, and the foregoing and other operations and / or functions of each module in the device 500 respectively implement the corresponding processes in the above methods. For the sake of brevity, they will not be elaborated here.

[0140] The device 500 of the embodiments of the present application has been described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0141] Figure 6 It is a schematic diagram of an electronic device 600 provided by an embodiment of the present application.

[0142] As Figure 6 shown, the electronic device 600 may include:

[0143] A memory 610 and a processor 620. The memory 610 is used to store a computer program and transmit the program code to the processor 620. In other words, the processor 620 can call and run the computer program from the memory 610 to implement the method in the embodiments of the present application.

[0144] For example, the processor 620 can be used to execute the above method embodiments according to the instructions in the computer program.

[0145] In some embodiments of the present application, the processor 620 may include but is not limited to:

[0146] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0147] In some embodiments of the present application, the memory 610 includes, but is not limited to:

[0148] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0149] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 610 and executed by the processor 620 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0150] As Figure 6 shown, the electronic device may further include:

[0151] A transceiver 630 that can be connected to the processor 620 or the memory 610.

[0152] Among them, the processor 620 can control the transceiver 630 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 630 can include a transmitter and a receiver. The transceiver 630 can further include an antenna, and the number of antennas can be one or more.

[0153] It should be understood that the components in the electronic device are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.

[0154] This application also provides a computer storage medium with a computer program stored thereon. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Or rather, this application embodiment also provides a computer program product containing instructions. When the instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0155] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the computer can execute all or part of the corresponding processes in the methods in this application embodiment and generate the functions achievable by the methods in this application embodiment. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a Digital Video Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.

[0156] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0157] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems, devices, or modules can be electrical, mechanical, or other forms.

[0158] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

Claims

1. A method for image and sound retrieval, characterized in that: include: Acquiring data to be retrieved, wherein the data to be retrieved includes voice data to be retrieved and / or image data to be retrieved; Inputting the voice data to be retrieved into the image-sound retrieval model to obtain image data related to the voice data to be retrieved, and / or inputting the image data to be retrieved into the image-sound retrieval model to obtain voice data related to the image data to be retrieved; Among them, the image-to-speech retrieval model is trained by multiple speech-image pair data and multiple text data related to the multiple speech-image pair data, and any target speech-image pair data in the multiple speech-image pair data includes: target speech data and target image data with consistent corresponding content.

2. The method according to claim 1, characterized in that Before inputting the voice data to be retrieved into the image-sound retrieval model to obtain image data related to the voice data to be retrieved, and / or inputting the image data to be retrieved into the image-sound retrieval model to obtain voice data related to the image data to be retrieved, the method further includes: Acquire the plurality of speech-image pair data; For the target speech-image pair data, obtaining speech features of the target speech data and image features of the target image data through the image-sound retrieval model; Generating target text data related to the target voice-image pair data according to the target voice-image pair data; Training the image-sound retrieval model according to the speech features, the image features and the target text data; The speech features include: global speech features and / or local speech features; the image features include: global image features and / or local image features.

3. The method according to claim 2, characterized in that The obtaining of the speech features of the target speech data and the image features of the target image data by the image-sound retrieval model includes: Extracting features of the target speech data through the audio encoder of the image-sound retrieval model to obtain speech features of the target speech data; The image encoder of the image-to-sound retrieval model performs feature extraction on the target image data to obtain image features of the target image data.

4. The method according to claim 2, characterized in that: The step of training the image-sound retrieval model according to the speech features, the image features and the target text data comprises: Determining a first loss according to the speech feature and the image feature; Determining a second loss according to the target text data and the target speech-image pair data; The image-to-speech retrieval model is trained according to the first loss and / or the second loss.

5. The method according to claim 4, characterized in that The determining the first loss according to the speech feature and the image feature comprises: For the speech feature, calculating the similarity value of the speech feature to the image feature to obtain the sound-image similarity; For the image feature, calculating a similarity value of the image feature to the speech feature to obtain image-speech similarity; Determine a first sub-loss according to the sound-image similarity, and determine a second sub-loss according to the image-sound similarity; The first loss is determined according to the first sub-loss and the second sub-loss.

6. The method according to claim 5, characterized in that The step of calculating the similarity value of the speech feature to the image feature to obtain the sound-image similarity includes: For the speech feature, based on the image feature of each image data in the plurality of speech-image pair data, a similarity value of the speech feature relative to the image feature is calculated to obtain the speech-image similarity; The step of calculating the similarity value of the image feature relative to the speech feature to obtain the image-sound similarity includes: For the image feature, based on the voice feature of each voice data in the plurality of voice-image pair data, a similarity value of the image feature relative to the voice feature is calculated to obtain the image-sound similarity.

7. A picture and sound retrieval device, characterized in that: include: A first acquisition module, used to acquire data to be retrieved, wherein the data to be retrieved includes voice data to be retrieved and / or image data to be retrieved; An image-sound retrieval module, used for inputting the voice data to be retrieved into an image-sound retrieval model to obtain image data related to the voice data to be retrieved, and / or inputting the image data to be retrieved into the image-sound retrieval model to obtain voice data related to the image data to be retrieved; Among them, the image-to-speech retrieval model is trained by multiple speech-image pair data and multiple text data related to the multiple speech-image pair data, and any target speech-image pair data in the multiple speech-image pair data includes: target speech data and target image data with consistent corresponding content.

8. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising instructions, characterized in that When the computer program product runs on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 6.