Video music identification method and apparatus, storage medium, and electronic device

CN122802704APending Publication Date: 2026-09-22TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510331696.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种视频音乐的识别方法、装置和存储介质及电子设备,以至少解决视频音乐的识别准确性较低的技术问题

Benefits of technology

[0021]根据本申请实施例的又一方面,还提供了一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,其中,上述处理器通过计算机程序执行上述的视频音乐的识别方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802704A_ABST
    Figure CN122802704A_ABST
Patent Text Reader

Abstract

The application discloses a video music identification method and device, a storage medium and an electronic equipment. The method comprises the following steps: obtaining video data of music to be identified; obtaining a first text feature corresponding to the video data, and obtaining a first audio feature corresponding to the video data, wherein the first text feature is used for representing text information associated with the video data, and the first audio feature is used for representing audio information in the video data; performing feature fusion on the first text feature and the first audio feature to obtain a first multi-modal feature corresponding to the video data; determining audio data from a music database by using the first multi-modal feature, and taking music corresponding to the audio data as music identified in the video data. The application solves the technical problem of low identification accuracy of video music.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a method, apparatus, storage medium, and electronic device for identifying video music. Background Technology

[0002] In video music recognition scenarios, identification is typically performed using the video's audio. However, background music in videos is usually edited, and the audio may have variations or pitch shifts, as well as noise or poor sound quality. These issues can affect the music recognition process, leading to inaccurate results and consequently, low accuracy in video music recognition. Therefore, there is a problem with the accuracy of video music recognition.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method, apparatus, storage medium, and electronic device for identifying video music, in order to at least solve the technical problem of low accuracy in identifying video music.

[0005] According to one aspect of the embodiments of this application, a method for identifying music in a video is provided, comprising: acquiring video data of music to be identified; acquiring a first text feature corresponding to the video data and acquiring a first audio feature corresponding to the video data, wherein the first text feature is used to characterize text information associated with the video data, and the first audio feature is used to characterize audio information in the video data; performing feature fusion on the first text feature and the first audio feature to obtain a first multimodal feature corresponding to the video data; using the first multimodal feature to determine audio data from a music database, and using the music corresponding to the audio data as the music identified in the video data; wherein a second multimodal feature corresponding to the audio data matches the first multimodal feature, the second multimodal feature is obtained by feature fusion of the second text feature and the second audio feature, the second text feature is used to characterize text information associated with the audio data, and the second audio feature is used to characterize audio information in the audio data.

[0006] According to another aspect of the embodiments of this application, a video music recognition device is also provided, comprising: a first acquisition unit, configured to acquire video data of music to be recognized; a second acquisition unit, configured to acquire a first text feature corresponding to the video data and acquire a first audio feature corresponding to the video data, wherein the first text feature is used to characterize text information associated with the video data, and the first audio feature is used to characterize audio information in the video data; a fusion unit, configured to perform feature fusion on the first text feature and the first audio feature to obtain a first multimodal feature corresponding to the video data; and a first determination unit, configured to use the first multimodal feature to determine audio data from a music database, and to use the music corresponding to the audio data as the music identified in the video data; wherein a second multimodal feature corresponding to the audio data matches the first multimodal feature, the second multimodal feature is obtained by feature fusion of the second text feature and the second audio feature, the second text feature is used to characterize text information associated with the audio data, and the second audio feature is used to characterize audio information in the audio data.

[0007] As an optional solution, the second acquisition unit includes: a first extraction module, used to extract features from the video description information associated with the video data to obtain the first feature, wherein the first text feature includes the first feature.

[0008] As an optional solution, the first extraction module includes: a first extraction submodule, used to extract features from the video title information of the video data to obtain title text features, wherein the first feature includes the title text features and the video description information includes the video title information; a second extraction submodule, used to extract features from the video tag information of the video data to obtain tag text features, wherein the first feature includes the tag text features and the video description information includes the video tag information; and a third extraction submodule, used to extract features from the video source information of the video data to obtain source text features, wherein the first feature includes the source text features and the video description information includes the video source information.

[0009] As an optional solution, the second acquisition unit includes: a second extraction module, used to extract features from the playback data in the video data to obtain the second feature, wherein the first text feature includes the second feature.

[0010] As an optional solution, the second extraction module includes: a first recognition submodule, used to perform optical character recognition on the playback data in the video data to obtain the first text corresponding to the playback data; and a fourth extraction submodule, used to extract features from the first text to obtain the screen text features, wherein the second feature includes the screen text features.

[0011] As an optional solution, the second extraction module includes: a second recognition submodule, used to perform automatic speech recognition on the audio playback data in the video data to obtain the second text corresponding to the audio playback data; and a fifth extraction submodule, used to extract features from the second text to obtain audio text features, wherein the second features include the audio text features.

[0012] As an optional solution, the first determining unit includes: a first determining module, used to determine the music segment corresponding to the audio data as the music identified in the video data; and a second determining module, used to determine the music to which the music segment corresponding to the audio data belongs as the music identified in the video data.

[0013] As an optional solution, the first acquisition unit includes: a first acquisition module, used to acquire the first text feature, the first audio feature, and the identity feature corresponding to the video data, wherein the identity feature is used to represent the identity information; a first fusion module, used to perform feature fusion on the first text feature, the first audio feature, and the identity feature to obtain a third multimodal feature corresponding to the video data; and a third determination module, used to determine target audio data from the music database using the third multimodal feature, and to use the music corresponding to the target audio data as the music identified in the video data; wherein a fourth multimodal feature corresponding to the target audio data matches the first multimodal feature, the fourth multimodal feature is obtained by feature fusion of the third text feature and the third audio feature, the third text feature is used to represent the text information associated with the target audio data, and the third audio feature is used to represent the audio information in the target audio data.

[0014] As an optional solution, the first determining unit includes: a retrieval module, used to retrieve the music multimodal feature that has the closest spatial distance to the first multimodal feature in the music retrieval library, wherein the spatial distance is negatively correlated with the similarity between the features; and a fourth determining module, used to determine the data corresponding to the music multimodal feature as the audio data when the feature similarity between the music multimodal feature and the first multimodal feature is greater than a preset threshold.

[0015] According to another aspect of the embodiments of this application, another method for identifying music in video is provided, comprising: processing multiple audio reference data using an audio multimodal model to obtain multimodal features corresponding to each audio reference data in the multiple audio reference data; adding the multimodal features corresponding to each audio reference data to a music database; acquiring video data of music to be identified, wherein the multiple audio reference data includes the audio data; processing the video data using a video multimodal model to obtain a first multimodal feature corresponding to the video data, wherein the audio multimodal model and the video multimodal model are jointly trained; determining audio data from the music database using the first multimodal feature, and using the music corresponding to the audio data as the music identified in the video data; wherein a second multimodal feature corresponding to the audio data matches the first multimodal feature, the second multimodal feature is obtained by feature fusion of a second text feature and a second audio feature, the second text feature being used to characterize the text information associated with the audio data, and the second audio feature being used to characterize the audio information in the audio data.

[0016] According to another aspect of the embodiments of this application, another video music recognition device is also provided, comprising: a first processing unit, configured to process multiple audio reference data using an audio multimodal model to obtain multimodal features corresponding to each audio reference data in the multiple audio reference data; an adding unit, configured to add the multimodal features corresponding to each audio reference data to a music database; a third acquisition unit, configured to acquire video data of the music to be recognized, wherein the multiple audio reference data includes the audio data; and a second processing unit, configured to process the video data using a video multimodal model to obtain the first multimodal features corresponding to the video data. The multimodal features include an audio multimodal model and a video multimodal model that are jointly trained; a second determining unit is used to determine audio data from a music database using the first multimodal features and to use the music corresponding to the audio data as the music identified in the video data; wherein the second multimodal feature corresponding to the audio data matches the first multimodal feature, and the second multimodal feature is obtained by fusing the second text feature and the second audio feature, the second text feature being used to characterize the text information associated with the audio data, and the second audio feature being used to characterize the audio information in the audio data.

[0017] As an optional solution, the second processing unit includes: a first processing module for segmenting the video data to obtain multiple video segments; a third extraction module for extracting features from the multiple video segments using a first segment model in the video multimodal model to obtain text features and audio features corresponding to each video segment; a second fusion module for fusing the text features and audio features corresponding to each video segment using a first fusion model in the video multimodal model to obtain multimodal features corresponding to each video segment; and a first sequence model in the video multimodal model for temporally fusing the multimodal features corresponding to the multiple video segments to obtain the first multimodal features.

[0018] As an optional solution, the first processing unit includes: a second processing module for segmenting the audio data to obtain multiple audio segments; a fourth extraction module for extracting features from the multiple audio segments using the second segment model in the audio multimodal model to obtain text features and audio features corresponding to each audio segment; and a third fusion module for performing temporal fusion of the multimodal features corresponding to the multiple audio segments using the second sequence model in the audio multimodal model to obtain the second multimodal features.

[0019] As an optional solution, the second processing unit includes: a second acquisition module, used to acquire multiple positive and negative sample pairs, wherein the positive sample pairs in the positive and negative sample pairs consist of the same video and corresponding music, and the negative sample pairs in the positive and negative sample pairs consist of the same video and non-corresponding music; and a training module, used to jointly train the initial video multimodal model and the initial audio multimodal model using the multiple positive and negative sample pairs until the joint convergence condition is met, thereby obtaining the trained video multimodal model and the trained audio multimodal model.

[0020] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the video and music recognition method described above.

[0021] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described video music recognition method through the computer program.

[0022] In this embodiment, video data of the music to be identified is first acquired, and then a first text feature associated with the video data and a first audio feature from the video are extracted. The first text feature is used to characterize the text information related to the video, while the first audio feature directly reflects the audio information in the video.

[0023] Further feature fusion of the first text features and the first audio features yields the first multimodal features corresponding to the video data. By fusing information from both text and audio dimensions, a more comprehensive and accurate music representation is formed, thereby enhancing the ability to recognize music.

[0024] Furthermore, the first multimodal feature is used to search for matching audio data in the music database. Each audio data in the music database is also fused with its associated second text feature and its own second audio feature in a similar manner to form a second multimodal feature. By comparing the similarity between the first and second multimodal features, audio data matching the video music can be identified.

[0025] In the process of using the music corresponding to the matched audio data as the music identified in the video data, since the multimodal features integrate information from both text and audio dimensions, even if there are issues such as variations, pitch shifts, or noise in the audio of the video, the text information can still provide useful clues to assist in the recognition process. This reduces the possibility of deviations in the recognition results of video music, thereby improving the technical effect of video music recognition accuracy and solving the technical problem of low recognition accuracy of video music. Attached Figure Description

[0026] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0027] Figure 1 This is a schematic diagram of an application environment for an optional video music recognition method according to an embodiment of this application;

[0028] Figure 2 This is a schematic diagram of the flow of an optional video music recognition method according to an embodiment of this application;

[0029] Figure 3This is a schematic diagram of an optional video music recognition method according to an embodiment of this application;

[0030] Figure 4 This is a schematic diagram of an alternative video music recognition method according to an embodiment of this application;

[0031] Figure 5 This is a schematic diagram of another optional video music recognition method according to an embodiment of this application;

[0032] Figure 6 This is a schematic diagram of another optional video music recognition method according to an embodiment of this application;

[0033] Figure 7 This is a schematic diagram of another optional video music recognition method according to an embodiment of this application;

[0034] Figure 8 This is a schematic diagram of another optional video music recognition method according to an embodiment of this application;

[0035] Figure 9 This is a schematic diagram of another optional video music recognition method according to an embodiment of this application;

[0036] Figure 10 This is a schematic diagram of another optional video music recognition method according to an embodiment of this application;

[0037] Figure 11 This is a schematic diagram of an optional video music recognition device according to an embodiment of this application;

[0038] Figure 12 This is a schematic diagram of an alternative video music recognition device according to an embodiment of this application;

[0039] Figure 13 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0040] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0041] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0042] According to one aspect of the embodiments of this application, a method for identifying video music is provided. Optionally, as an optional implementation, the above-described method for identifying video music can be applied to, but is not limited to, [examples of other methods]. Figure 1 The environment shown may include, but is not limited to, user equipment 102 and server 112. User equipment 102 may include, but is not limited to, a display 104, a processor 106 and a memory 108. Server 112 includes a database 114 and a processing engine 116.

[0043] The specific process can be summarized in the following steps:

[0044] Step S102: User equipment 102 obtains a video music recognition request;

[0045] Step S104: Send the video music recognition request to the server 112 via network 110;

[0046] In steps S106-S112, server 112 obtains video data of the music to be identified through processing engine 116, then obtains the first text feature corresponding to the video data and the first audio feature corresponding to the video data, then performs feature fusion on the first text feature and the first audio feature to obtain the first multimodal feature corresponding to the video data, and finally uses the first multimodal feature to determine the audio data from the music database, and uses the music corresponding to the audio data as the music identified in the video data;

[0047] In step S114, the video music recognition result is sent to the user equipment 102 via the network 110. The user equipment 102 displays the video music recognition result on the display 104 via the processor 106 and stores the video music recognition result in the memory 108.

[0048] remove Figure 1Beyond the examples shown, the terminal devices described above can be terminal devices configured with a target client, including but not limited to at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, etc. The target client can be a video client, instant messaging client, browser client, educational client, etc. The networks described above can include, but are not limited to, wired networks and wireless networks. The wired networks include local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless networks include Bluetooth, Wi-Fi, and other networks that enable wireless communication. The server described above can be a single server, a server cluster consisting of multiple servers, or a cloud server. The above is merely an example, and no limitations are imposed in this embodiment.

[0049] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the video music recognition method can be performed by an electronic device, such as... Figure 1 The user equipment or server shown includes the following specific steps:

[0050] S202, Obtain video data of the music to be identified;

[0051] S204, obtain the first text feature corresponding to the video data and obtain the first audio feature corresponding to the video data, wherein the first text feature is used to characterize the text information associated with the video data and the first audio feature is used to characterize the audio information in the video data.

[0052] S206, perform feature fusion on the first text feature and the first audio feature to obtain the first multimodal feature corresponding to the video data;

[0053] S208: Using the first multimodal feature, determine the audio data from the music database, and use the music corresponding to the audio data as the music identified in the video data.

[0054] Among them, the second multimodal feature corresponding to the audio data is matched with the first multimodal feature. The second multimodal feature is obtained by feature fusion of the second text feature and the second audio feature. The second text feature is used to represent the text information associated with the audio data, and the second audio feature is used to represent the audio information in the audio data.

[0055] In optional embodiments, the video music recognition method can be applied to multiple scenarios such as short video applications, long video applications, and live streaming platforms.

[0056] When video music recognition methods are applied to short video applications, it can be used, but is not limited to, when a user encounters a short video that interests them and wants to know the specific information about the music in that video, they can obtain that music information by using video music recognition methods on that short video.

[0057] When video music recognition methods are applied to long-form video applications, this can be used, but is not limited to, when a user is watching a long video and becomes interested in a particular segment, wanting to know the music information within that segment. In this case, the segment containing the music to be identified can be extracted from the long video, and then the video music recognition method can be applied to that segment to derive the music information for that specific video segment.

[0058] When video music recognition methods are applied to live streaming platforms, it can be, but is not limited to, when a user hears a piece of music that interests them while watching a live stream. In this case, the user can record the live stream and apply the video music recognition method to the recorded video to obtain the music information from the live stream.

[0059] In an optional embodiment, the video data may include, but is not limited to, data containing video information extracted from the video stream.

[0060] In an optional embodiment, the first text feature may be, but is not limited to, text information used to characterize the association of video data.

[0061] To further illustrate, for a video on a short video application, the text information that can be used to characterize the video data association can include, but is not limited to, video source information, video title, video author name, video tags, video description, text information appearing in the video, text information corresponding to the audio information in the video, text information of bullet comments appearing in the video, text information in the comment section of the video, and other text information associated with the video.

[0062] In an optional embodiment, the first audio feature may be, but is not limited to, extracted from the audio information of the video data, used to characterize the audio data in the video, and may include, but is not limited to, information such as musical melody, musical rhythm, and musical spectrum.

[0063] In optional embodiments, feature fusion may be, but is not limited to, the process of fusing features from different modalities to generate a comprehensive feature representation.

[0064] In an optional embodiment, the first multimodal feature may be, but is not limited to, a feature obtained by feature fusion of the first text feature and the first audio feature.

[0065] In an optional embodiment, the music database may be, but is not limited to, a database for storing audio data or music data.

[0066] In an optional embodiment, the second multimodal feature may be, but is not limited to, a feature obtained by fusing the second text feature and the second audio feature, which can be used to represent a certain audio data in a music database.

[0067] In an optional embodiment, the second text feature may be, but is not limited to, text information associated with a certain audio data in the music database, and the second audio feature may be, but is not limited to, audio information associated with a certain audio data in the aforementioned music database.

[0068] To further illustrate, suppose there is audio data A. The second text feature of audio data A can be, but is not limited to, text information associated with audio data A, such as the author of the audio, the lyrics of the audio, and the time when the audio was released. The audio data feature of audio data A can be, but is not limited to, audio information associated with audio data A, such as music melody, music rhythm, and music spectrum.

[0069] In an optional embodiment, after receiving a request to identify music in a video, the video data of the music to be identified is obtained, and the video data includes image data, audio data, text data, and other data.

[0070] Furthermore, based on the acquired video data, the first text features are obtained by analyzing text information associated with the video data, such as video title, video author name, video tags, video description, text information appearing in the video, text information corresponding to audio information in the video, text information of bullet comments appearing in the video, and text information in the comment section of the video. Simultaneously, the first audio features are obtained by analyzing audio information in the video data, such as musical melody, musical rhythm, and musical spectrum.

[0071] Then, by fusing the acquired first text features and first audio features, a first multimodal feature corresponding to the video data and capable of representing the music information in the video is obtained.

[0072] Next, based on the first multimodal feature obtained by fusion, the audio multimodal feature corresponding to each audio data is retrieved in the music database. The audio multimodal feature is a feature fusion of the audio text feature and the audio feature of the audio data, which can be used to characterize the feature representation of the audio data. The first multimodal feature is then used to match multiple audio multimodal features.

[0073] Finally, if the audio multimodal feature corresponding to the audio data matches the first multimodal feature in the audio database, then the audio data is identified, and the audio multimodal feature of the audio data is the second multimodal feature. The music corresponding to the audio data is then used as the music identified in the video data.

[0074] It should be noted that accurate music recognition in videos is achieved through multimodal fusion of audio and text information. First, the system acquires video data from the video stream. Then, it extracts first text features and first audio features using text processing and audio processing models, respectively, representing the text and audio information of the video data. Next, the first text and first audio features are fused to obtain the first multimodal feature, which incorporates comprehensive information from both the video's audio and text. Then, the system uses the first multimodal feature to search a music database, finding the audio data with the highest matching degree. The music corresponding to this audio data is then used as the recognition result for the video's music. Each song or audio segment in the music database is pre-processed through feature extraction and fusion to obtain a second multimodal feature, enabling effective comparison between the music or audio in the database and the multimodal information in the video.

[0075] In addition, since the identified or matched audio data may be a musical segment from a song rather than a complete piece of music, when returning the identification results of the video music, the music corresponding to the audio data is returned as the music identified in the video data.

[0076] Further examples, such as Figure 3 As shown, after receiving the video music recognition request 304 from the user equipment 302, the video data 306 of the music to be recognized is obtained. Then, the multimodal feature 308 corresponding to the video data 306 is obtained based on the video data 306. Next, the audio data 312 is determined from the music database 310 based on the multimodal feature 308. Finally, the music 316 corresponding to the audio data 312 is determined based on the audio data 312. Among them, the multimodal feature 308 is a multimodal feature obtained by extracting and fusing features from the video data 306, and the multimodal feature 314 is a multimodal feature obtained by extracting and fusing features from the audio data 312. The multimodal feature 308 matches the multimodal feature 314.

[0077] In this embodiment, video data of the music to be identified is first acquired, and then a first text feature associated with the video data and a first audio feature from the video are extracted. The first text feature is used to characterize the text information related to the video, while the first audio feature directly reflects the audio information in the video.

[0078] Further feature fusion of the first text features and the first audio features yields the first multimodal features corresponding to the video data. By fusing information from both text and audio dimensions, a more comprehensive and accurate music representation is formed, thereby enhancing the ability to recognize music.

[0079] Furthermore, the first multimodal feature is used to search for matching audio data in the music database. Each audio data in the music database is also fused with its associated second text feature and its own second audio feature in a similar manner to form a second multimodal feature. By comparing the similarity between the first and second multimodal features, audio data matching the video music can be identified.

[0080] In the process of identifying the music corresponding to the matched audio data as the music in the video data, since the multimodal features integrate information from both text and audio dimensions, even if there are problems such as variations, pitch shifts, or noise in the audio of the video, the text information can still provide useful clues to assist the recognition process. This reduces the possibility of deviations in the recognition results of video music, thereby achieving the technical effect of improving the recognition accuracy of video music.

[0081] As an optional approach, the first text features corresponding to the video data are obtained, including:

[0082] Feature extraction is performed on the video description information associated with the video data to obtain the first feature, wherein the first text feature includes the first feature.

[0083] In optional embodiments, the video description information may be, but is not limited to, information used to describe video data, and may include, but is not limited to, the video title, video author name, video tags, video description, bullet screen text information appearing in the video, and text information in the comment section of the video.

[0084] It should be noted that by extracting features from the video description information associated with the video data, a first feature that can be used to characterize the video description information associated with the video data is obtained. This allows the video description information to be used as one of the features for music recognition, thereby improving the accuracy of video music recognition.

[0085] Through the embodiments of this application, feature extraction is performed on the video description information associated with video data to obtain a first feature, wherein the first text feature includes the first feature. By obtaining the first feature through the video description information associated with video data, the technical objective of using video description information as one of the features for music recognition is achieved, thereby improving the technical effect of improving the accuracy of video music recognition.

[0086] As an optional approach, feature extraction is performed on the video description information associated with the video data to obtain a first feature, which includes at least one of the following:

[0087] S1-1, extract features from the video title information of the video data to obtain title text features, wherein the first feature includes the title text features and the video description information includes the video title information;

[0088] S1-2, extract features from the video tag information of the video data to obtain tag text features, wherein the first feature includes tag text features and the video description information includes video tag information;

[0089] S1-3, extract features from the video source information of the video data to obtain source text features, wherein the first feature includes source text features and the video description information includes video source information.

[0090] In an optional embodiment, the video title information may include, but is not limited to, information that includes a video title.

[0091] In optional embodiments, the title text features may be, but are not limited to, features extracted from the video title, capable of representing information such as keywords and themes in the title content.

[0092] In optional embodiments, video tagging information may be, but is not limited to, tagging information added by the video author or platform for classifying and labeling videos.

[0093] To illustrate further, suppose a short video is about A's concert. The video tag information for this video data can include tags such as "music", "concert", and "A".

[0094] In optional embodiments, the tag text features may be, but are not limited to, feature representations obtained after feature extraction of video tag information.

[0095] In optional embodiments, the video source information may be, but is not limited to, the source information of the video data, and may include, but is not limited to, information such as the author of the video data, the publishing platform of the video data, and the publishing time of the video data.

[0096] In optional embodiments, the source text features may be, but are not limited to, feature information obtained after feature extraction of video source information.

[0097] In an optional embodiment, if the video description information includes video title information, the video title information is used to extract features to obtain title text features, and the title text features are used as part of the first feature.

[0098] To illustrate further, suppose there is a video with the title "Song XXX is amazing!", then information about music recognition can be obtained based on the title, such as "Song XXX" mentioned above.

[0099] In an optional embodiment, if the video description information includes video tag information, the video tag information is used to extract features to obtain tag text features, and the tag text features are used as part of the first feature.

[0100] To illustrate further, suppose there is a video and the video is tagged with "#EveryoneLearnsSingSongXXX". Then, information about music recognition can be obtained based on this tag, such as "SongXXX" mentioned above.

[0101] In an optional embodiment, if the video description information includes video source information, and the video source information includes the name of the video publisher, the time of video publication, the location of video publication, etc., the video source information is used to extract features to obtain source text features, and the source text features are used as part of the first feature.

[0102] To illustrate further, suppose there is a video whose publisher is named "Only Listen to YYY's Songs", and the video source information includes the publisher's name. In this case, the "YYY" in the publisher's name can help in identifying the music in the video.

[0103] It should be noted that by extracting features from at least one of the video title information, video tag information, and video source information, the text description dimension of the video data is enriched. Even when the audio quality in the video data is poor, the recognition accuracy can still be maintained through the features corresponding to the video description information, thereby improving the recognition accuracy of music in the video.

[0104] In this embodiment, feature extraction is performed on the video title information of the video data to obtain title text features, wherein the first feature includes title text features, and the video description information includes video title information; feature extraction is performed on the video tag information of the video data to obtain tag text features, wherein the first feature includes tag text features, and the video description information includes video tag information; feature extraction is performed on the video source information of the video data to obtain source text features, wherein the first feature includes source text features, and the video description information includes video source information. By performing feature extraction on at least one of the video title information, video tag information, and video source information, a high recognition accuracy can be maintained even when the audio quality in the video data is poor, thereby achieving the technical effect of improving the recognition accuracy of music in the video.

[0105] As an optional approach, the first text features corresponding to the video data are obtained, including:

[0106] Feature extraction is performed on the playback data in the video data to obtain the second feature, wherein the first text feature includes the second feature.

[0107] In optional embodiments, the playback data may be, but is not limited to, data obtained during video playback, and may include, but is not limited to, text appearing in the images of the video data during playback, and may also include, but is not limited to, text appearing in the audio of the video data during playback.

[0108] It should be noted that by recognizing playback data in video data, text information present in the video frame or in the video audio can be effectively obtained. The text information obtained from the playback data of the video data is then used to extract features, and these extracted features are used as one of the references for video music recognition. Thus, even when the music in the video to be recognized has been modulated or varied, the features extracted from the playback data can still be used for relatively accurate recognition, thereby improving the accuracy of video music recognition.

[0109] Through the embodiments of this application, feature extraction is performed on playback data in video data to obtain a second feature, wherein the first text feature includes the second feature. By identifying playback data in video data and extracting features, the technical objective of enabling more accurate identification through features extracted from playback data is achieved, thereby improving the technical effect of improving the accuracy of video music recognition.

[0110] As an alternative approach, feature extraction is performed on the playback data in the video data to obtain a second feature, including:

[0111] S2-1, Perform optical character recognition on the video playback data to obtain the first text corresponding to the video playback data;

[0112] S2-2, extract features from the first text to obtain the image text features, wherein the second feature includes the image text features.

[0113] In an optional embodiment, the video playback data may be, but is not limited to, text information contained in the images that appear during video playback.

[0114] In optional embodiments, Optical Character Recognition (OCR) may be, but is not limited to, a technology for converting text in video footage into editable and searchable text data, and may be, but is not limited to, for converting text appearing in video images into editable and searchable text data.

[0115] In optional embodiments, the first text can be, but is not limited to, converting text in the video footage into editable and searchable text data using optical character recognition technology.

[0116] In optional embodiments, the image text features may be, but are not limited to, features extracted from the first text, and may be, but are not limited to, features used to characterize the text content in the video frame during video playback.

[0117] It should be noted that optical character recognition technology identifies text information in video footage and converts it into processable first text. This text may contain information such as music titles, composers, or lyrics. Then, feature extraction is performed on the first text to obtain the text features of the video. This provides additional feature information for video music recognition based on the text features of the video, thereby improving the accuracy of video music recognition.

[0118] This application embodiment performs optical character recognition (OCR) on the playback data in video data to obtain the first text corresponding to the playback data; it then extracts features from the first text to obtain the video text features, wherein the second feature includes the video text features. By using OCR technology to identify text information in video footage and converting it into processable first text, and then extracting features from the first text to obtain the video text features, the technical objective of providing additional feature information for video music recognition based on the video text features is achieved, thereby improving the accuracy of video music recognition.

[0119] As an alternative approach, feature extraction is performed on the playback data in the video data to obtain a second feature, including:

[0120] S3-1, Automatic speech recognition is performed on the audio playback data in the video data to obtain the second text corresponding to the audio playback data;

[0121] S3-2, extract features from the second text to obtain audio text features, wherein the second features include audio text features.

[0122] In optional embodiments, the audio playback data may be, but is not limited to, data in the audio stream of the video that includes information such as music lyrics, dialogue text, and narration text.

[0123] In an optional embodiment, Automatic Speech Recognition (ASR) may be, but is not limited to, a technique for transcribing text information from audio signals containing text information in the audio stream of a video.

[0124] In an optional embodiment, the second text may be understood, but is not limited to, as a collection of text information transcribed by automatic speech recognition technology from audio signals containing text information in the audio stream of a video.

[0125] In optional embodiments, the audio text features may be, but are not limited to, features extracted from the second text, and may be, but are not limited to, used to characterize the text content in the video frame during playback.

[0126] It should be noted that by using automatic speech recognition technology, the audio signal containing text information in the audio stream of the video is transcribed into a set of text information as the second text. The second text may contain text information such as music title, music author or lyrics. Then, feature extraction is performed on the second text to obtain audio text features. This provides additional feature information for video music recognition based on the audio text features, thereby improving the accuracy of video music recognition.

[0127] This application embodiment performs automatic speech recognition on audio playback data in video data to obtain second text corresponding to the audio playback data; feature extraction is then performed on the second text to obtain audio text features, wherein the second features include audio text features. By using automatic speech recognition technology to transcribe audio signals containing text information in the video audio stream into a set of text information as the second text, and then performing feature extraction on the second text to obtain audio text features, the technical objective of providing additional feature information for video music recognition is achieved, thereby improving the technical effect of improving the accuracy of video music recognition.

[0128] As an optional approach, the music corresponding to the audio data can be used as the music identified in the video data, including:

[0129] S4-1, use the music segment corresponding to the audio data as the music identified in the video data; or,

[0130] S4-2 uses the music segment corresponding to the audio data as the music identified in the video data.

[0131] It should be noted that when the music segment corresponding to the audio data is a complete piece of music or song, or when it is set to use a music segment as the recognition result, the music segment can be directly used as the music recognized in the video data.

[0132] If the goal is to output a complete song as the result of music recognition, it is necessary to determine whether the music segment corresponding to the audio data is a complete song to determine the result of music recognition from the video data. If the music segment corresponding to the audio data is not a complete song, the recognition result is based on the complete song to which the music segment belongs, thereby improving the accuracy of video music recognition.

[0133] Through the embodiments of this application, the music corresponding to the audio data is used as the music identified in the video data; or, the music to which the music segment corresponding to the audio data belongs is used as the music identified in the video data. By using different music recognition results in different situations, the technical objective of using the music segment to which the music segment belongs as the recognition result is achieved even when the music segment corresponding to the audio data is not a complete piece of music, thereby improving the technical effect of increasing the flexibility of video music recognition.

[0134] As an alternative approach, when the video data contains identity information, after obtaining the video data of the music to be identified, the method further includes:

[0135] S5-1, Obtain the first text feature, the first audio feature, and the identity feature corresponding to the video data, wherein the identity feature is used to represent identity information;

[0136] S5-2, perform feature fusion on the first text feature, the first audio feature, and the identity feature to obtain the third multimodal feature corresponding to the video data;

[0137] S5-3, using the third multimodal feature, the target audio data is determined from the music database, and the music corresponding to the target audio data is used as the music identified in the video data;

[0138] Among them, the fourth multimodal feature corresponding to the target audio data matches the first multimodal feature. The fourth multimodal feature is obtained by feature fusion of the third text feature and the third audio feature. The third text feature is used to characterize the text information associated with the target audio data, and the third audio feature is used to characterize the audio information in the target audio data.

[0139] In optional embodiments, the identity features may be, but are not limited to, features that characterize identity information, and may be, but are not limited to, facial information appearing in video footage.

[0140] To illustrate further, suppose singer B is singing a song in video A. Singer B's identity information in the video can be used as one of the bases for identifying the music in the video.

[0141] In an optional embodiment, after obtaining the first text feature, the first audio feature, and the identity feature corresponding to the video data, the first text feature, the first audio feature, and the identity feature are fused to obtain the third multimodal feature corresponding to the video data.

[0142] Furthermore, based on the fused third multimodal feature, the audio multimodal feature corresponding to each audio data is retrieved in the music database. The audio multimodal feature is formed by fusing the audio text feature and audio feature of the audio data, and can be used to characterize the feature representation of the audio data. The third multimodal feature is then used to match multiple audio multimodal features.

[0143] Finally, if the audio multimodal feature corresponding to the audio data matches the third multimodal feature in the audio database, then the audio data is identified, and the audio multimodal feature of the audio data is the fourth multimodal feature. The music corresponding to the audio data is then used as the music identified in the video data.

[0144] It should be noted that by introducing identity features corresponding to video data that can be used to represent identity information, and by fusing features based on the first text feature, the first audio feature, and the identity feature to obtain the third multimodal feature, the image information appearing in the video screen can also be included as one of the bases for video music recognition, thereby improving the accuracy of video music recognition.

[0145] Next, through the embodiments of this application, a first text feature, a first audio feature, and an identity feature corresponding to the video data are obtained, wherein the identity feature is used to represent identity information; the first text feature, the first audio feature, and the identity feature are fused to obtain a third multimodal feature corresponding to the video data; using the third multimodal feature, target audio data is determined from a music database, and the music corresponding to the target audio data is used as the music identified in the video data; wherein, the fourth multimodal feature corresponding to the target audio data matches the first multimodal feature, and the fourth multimodal feature is obtained by fusing the third text feature and the third audio feature, wherein the third text feature is used to represent the text information associated with the target audio data, and the third audio feature is used to represent the audio information in the target audio data. By introducing an identity feature corresponding to the video data that can be used to represent identity information, and by fusing the first text feature, the first audio feature, and the identity feature to obtain the third multimodal feature, the technical objective of including the facial information appearing in the video frame as one of the bases for video music recognition is achieved, thereby improving the technical effect of improving the accuracy of video music recognition.

[0146] As an alternative approach, audio data is determined from a music database using first multimodal features, including:

[0147] S6-1, retrieve the music multimodal feature that has the closest spatial distance to the first multimodal feature from the music retrieval database, where the spatial distance is negatively correlated with the similarity between features;

[0148] S6-2, if the feature similarity between the music multimodal feature and the first multimodal feature is greater than a preset threshold, the data corresponding to the music multimodal feature is determined as audio data.

[0149] In optional embodiments, the music retrieval library may be, but is not limited to, a database storing music multimodal features corresponding to music data, and may be, but is not limited to, used to match music multimodal features based on the first multimodal feature.

[0150] In optional embodiments, spatial distance can be understood, but is not limited to, as an indicator for judging feature similarity. The closer the spatial distance between two features, the higher the similarity between the two features. Therefore, spatial distance and feature similarity are negatively correlated.

[0151] It should be noted that spatial distance is calculated between the first multimodal feature and all music multimodal features in the music retrieval database. Music multimodal features whose spatial distance to the first multimodal feature is less than a spatial distance threshold are then considered to have a similarity greater than a preset threshold. Finally, the data corresponding to these music multimodal features with a similarity greater than the preset threshold are identified as audio data. By calculating the spatial distance between the first multimodal feature and the music retrieval database, music multimodal features with a similarity greater than a specific threshold can be quickly identified, thereby improving the efficiency of obtaining music data through the first multimodal feature.

[0152] This application embodiment retrieves the music multimodal feature with the closest spatial distance to the first multimodal feature from a music retrieval database, where spatial distance and feature similarity are negatively correlated. If the feature similarity between the music multimodal feature and the first multimodal feature is greater than a preset threshold, the data corresponding to the music multimodal feature is identified as audio data. By calculating the spatial distance between the first multimodal feature and the music retrieval database, the technical objective of quickly identifying music multimodal features with a similarity greater than a specific threshold to the first multimodal feature is achieved, thereby improving the efficiency of acquiring music data through the first multimodal feature.

[0153] Optionally, according to another aspect of the embodiments of this application, another method for identifying video music is provided, such as... Figure 4 As shown, the specific steps include:

[0154] S402, using an audio multimodal model, processes multiple audio reference data to obtain the multimodal features corresponding to each audio reference data in the multiple audio reference data;

[0155] S404, add the multimodal features corresponding to each audio reference data to the music database;

[0156] S406, Obtain video data of the music to be identified, wherein multiple audio reference data include audio data;

[0157] S408 utilizes a video multimodal model to process video data and obtain the first multimodal feature corresponding to the video data. The audio multimodal model and the video multimodal model are jointly trained.

[0158] S410: Using the first multimodal feature, determine the audio data from the music database, and use the music corresponding to the audio data as the music identified in the video data;

[0159] The second multimodal feature corresponding to the audio data matches the first multimodal feature. The second multimodal feature is obtained by fusing the second text feature and the second audio feature. The second text feature is used to characterize the text information associated with the audio data, and the second audio feature is used to characterize the audio information in the audio data.

[0160] In an optional embodiment, the audio reference data may be, but is not limited to, music data existing in a music database used for matching with the first multimodal feature.

[0161] In optional embodiments, the music database may be, but is not limited to, a database that stores music reference data and music multimodal features corresponding to the music reference data, and may be, but is not limited to, used to match music multimodal features based on the first multimodal feature.

[0162] In optional embodiments, multimodal features can be understood, but are not limited to, as feature representations used to indicate specific music reference data in a music retrieval library.

[0163] In optional embodiments, the video multimodal model may be, but is not limited to, a model capable of processing video data and obtaining first multimodal features through feature extraction.

[0164] In optional embodiments, the audio multimodal model may be, but is not limited to, a model that can process multiple audio reference data in a music database and extract features to obtain multiple audio multimodal features corresponding to the multiple audio reference data.

[0165] In an optional embodiment, the video data may include, but is not limited to, data containing video information extracted from the video stream.

[0166] In optional embodiments, the video multimodal model may be, but is not limited to, a model capable of processing video data and obtaining first multimodal features through feature extraction.

[0167] In optional embodiments, the video multimodal model and the audio multimodal model may be, but are not limited to, obtained by joint training. Joint training may be, but is not limited to, training two or more models simultaneously during the model training process and utilizing the correlation between tasks to improve the generalization ability of the trained model.

[0168] In an optional embodiment, the first multimodal feature may be, but is not limited to, a feature obtained by feature fusion of the first text feature and the first audio feature.

[0169] In an optional embodiment, the first text feature may be, but is not limited to, text information used to characterize the association of video data.

[0170] To further illustrate, for a video on a short video application, the text information that can be used to characterize the video data association can include, but is not limited to, video source information, video title, video author name, video tags, video description, text information appearing in the video, text information corresponding to the audio information in the video, text information of bullet comments appearing in the video, text information in the comment section of the video, and other text information associated with the video.

[0171] In an optional embodiment, the first audio feature may be, but is not limited to, extracted from the audio information of the video data, used to characterize the audio data in the video, and may include, but is not limited to, information such as musical melody, musical rhythm, and musical spectrum.

[0172] In an optional embodiment, the second multimodal feature may be, but is not limited to, a feature obtained by fusing the second text feature and the second audio feature, which can be used to represent a certain audio data in a music database.

[0173] In an optional embodiment, the second text feature may be, but is not limited to, text information associated with a certain audio data in the music database, and the second audio feature may be, but is not limited to, audio information associated with a certain audio data in the aforementioned music database.

[0174] To further illustrate, suppose there is audio data A. The second text feature of audio data A can be, but is not limited to, text information associated with audio data A, such as the author of the audio, the lyrics of the audio, and the time when the audio was released. The audio data feature of audio data A can be, but is not limited to, audio information associated with audio data A, such as music melody, music rhythm, and music spectrum.

[0175] In an optional embodiment, after receiving a request to identify music in a video, an audio multimodal model is used to process multiple audio reference data in the music database to obtain multimodal features corresponding to each audio reference data.

[0176] Furthermore, the multimodal data corresponding to each audio reference data is added to the music database.

[0177] Then, acquire the video data of the music to be identified.

[0178] Next, the video data is processed using a video multimodal model to obtain first text features and first audio features. The first text features and first audio features are then fused to obtain the first multimodal feature. Then, the second text features and second audio features corresponding to each audio data point are retrieved from the music database. The second text features and second audio features are then fused to obtain the audio multimodal feature. Finally, the first multimodal feature is matched with multiple audio multimodal features.

[0179] Finally, if the audio multimodal feature corresponding to the audio data matches the first multimodal feature in the audio database, the audio data is identified, the audio multimodal feature of the audio data is identified as the second multimodal feature, and the music corresponding to the audio data is identified as the music in the video data.

[0180] It should be noted that accurate music recognition in videos is achieved through multimodal fusion of audio and text information. First, the system acquires video data from the video stream. Then, it extracts first text features and first audio features using text processing and audio processing models, respectively, representing the text and audio information of the video data. Next, the first text and first audio features are fused to obtain the first multimodal feature, which incorporates comprehensive information from both the video's audio and text. Then, the system uses the first multimodal feature to search a music database, finding the audio data with the highest matching degree. The music corresponding to this audio data is then used as the recognition result for the video's music. Each song or audio segment in the music database is pre-processed through feature extraction and fusion to obtain a second multimodal feature, enabling effective comparison between the music or audio in the database and the multimodal information in the video.

[0181] Furthermore, since the identified or matched audio data may be a musical fragment from a song rather than a complete piece of music, the music corresponding to the audio data is returned as the music identified in the video data when returning the identification result of the video music. In this embodiment, video data of the music to be identified is first obtained, and then the first text feature associated with the video data and the first audio feature in the video are extracted respectively. The first text feature is used to characterize the text information related to the video, while the first audio feature directly reflects the audio information in the video.

[0182] Further feature fusion of the first text features and the first audio features yields the first multimodal features corresponding to the video data. By fusing information from both text and audio dimensions, a more comprehensive and accurate music representation is formed, thereby enhancing the ability to recognize music.

[0183] Furthermore, the first multimodal feature is used to search for matching audio data in the music database. Each audio data in the music database is also fused with its associated second text feature and its own second audio feature in a similar manner to form a second multimodal feature. By comparing the similarity between the first and second multimodal features, audio data matching the video music can be identified.

[0184] In the process of identifying the music corresponding to the matched audio data as the music in the video data, since the multimodal features integrate information from both text and audio dimensions, even if there are problems such as variations, pitch shifts, or noise in the audio of the video, the text information can still provide useful clues to assist the recognition process. This reduces the possibility of deviations in the recognition results of video music, thereby achieving the technical effect of improving the recognition accuracy of video music.

[0185] As an optional approach, a video multimodal model is used to process the video data to obtain the first multimodal features corresponding to the video data, including:

[0186] S7-1 segments the video data to obtain multiple video data segments.

[0187] S7-2 uses the first segment model in the video multimodal model to extract features from multiple video segments, obtaining the text features and audio features corresponding to each video segment.

[0188] S7-3, through the first fusion model in the video multimodal model, performs feature fusion on the text features and audio features corresponding to each video data segment to obtain the multimodal features corresponding to each video data segment;

[0189] S7-4 uses the first sequence model in the video multimodal model to perform temporal fusion of the multimodal features corresponding to multiple video data segments to obtain the first multimodal feature.

[0190] In optional embodiments, video data segmentation may be, but is not limited to, the process of dividing video data into segments to obtain multiple video data segments.

[0191] In an optional embodiment, the first segment model may be, but is not limited to, a model that extracts features from the obtained multiple video segments and obtains the text features and audio features corresponding to each video segment in the multi-terminal video data.

[0192] In an optional embodiment, the first fusion model may be, but is not limited to, a model that fuses the text features and audio features corresponding to each segment of video data to obtain the multimodal features corresponding to each segment of time data.

[0193] In an optional embodiment, the first sequence model may be, but is not limited to, temporally fusing the multimodal features corresponding to multiple video segments to obtain a model of the first multimodal features.

[0194] In an optional embodiment, after acquiring the video data of the music to be identified and before acquiring the first multimodal feature, the video data is first segmented into multiple segments. Then, each segment is processed using a first segment model to extract audio and text features. Next, the audio and text features of the video segments are fused using a first fusion model to obtain the multimodal features of each video segment. Finally, the multimodal features of all video segments are temporally fused using a first sequence model to obtain the first multimodal feature of the entire video data.

[0195] It should be noted that by segmenting the video and using different models in the video multimodal model for feature extraction and fusion, the system can more effectively process multimodal information in video and audio data and perform feature extraction and fusion, thereby improving the processing efficiency of multimodal information in video and thus enhancing the efficiency of feature extraction and fusion of multimodal information in video data.

[0196] This application embodiment segments video data to obtain multiple video segments. Using a first segment model within the video multimodal model, features are extracted from these segments to obtain text and audio features for each segment. A first fusion model within the video multimodal model fuses these text and audio features to obtain multimodal features for each segment. Finally, a first sequence model within the video multimodal model performs temporal fusion of the multimodal features from the multiple video segments to obtain the first multimodal feature. By segmenting the video and using different models within the video multimodal model for feature extraction and fusion, the system achieves its technical objective of more effectively processing multimodal information in the video and extracting and fusing features, thereby improving the efficiency of feature extraction and fusion for multimodal information in video data.

[0197] As an optional approach, in the process of using an audio multimodal model to process multiple audio reference data and obtain the multimodal features corresponding to each audio reference data, the method further includes:

[0198] S8-1, segment the audio data to obtain multiple audio data segments;

[0199] S8-2 uses the second segment model in the audio multimodal model to extract features from multiple audio segments, obtaining the text features and audio features corresponding to each segment of the audio data;

[0200] S8-3 uses the second fusion model in the audio multimodal model to fuse the text features and audio features corresponding to each audio data segment, thereby obtaining the multimodal features corresponding to each audio data segment.

[0201] S8-4 uses the second sequence model in the audio multimodal model to perform temporal fusion of the multimodal features corresponding to multiple audio data segments to obtain the second multimodal features.

[0202] In optional embodiments, the segmentation of audio data may be, but is not limited to, the process of dividing audio data into segments to obtain multiple audio data segments.

[0203] In an optional embodiment, the second segment model may be, but is not limited to, a model that extracts features from the obtained multiple audio segments and obtains the text features and audio features corresponding to each audio segment in the multiple audio data.

[0204] In an optional embodiment, the second fusion model may be, but is not limited to, a model that fuses the text features and audio features corresponding to each segment of audio data to obtain the multimodal features corresponding to each segment of time data.

[0205] In an optional embodiment, the second sequence model may be, but is not limited to, a model that temporally fuses the multimodal features corresponding to multiple audio segments to obtain a second multimodal feature model.

[0206] In an optional embodiment, the audio data is segmented into multiple segments before obtaining the second multimodal features. Then, each segment is processed using a second segment model to extract corresponding audio and text features. Next, the audio and text features of each segment are fused using a second fusion model to obtain the multimodal features of each segment. Finally, the multimodal features of all audio segments are temporally fused using a second sequence model to obtain the second multimodal features of the entire audio data.

[0207] It should be noted that by segmenting the audio data and using different models in the audio multimodal model for feature extraction and fusion, the system can more effectively process multimodal information in video and audio data and perform feature extraction and fusion, thereby improving the processing efficiency of multimodal information in multiple video and audio data, and thus enhancing the efficiency of feature extraction and fusion of multimodal information.

[0208] Furthermore, extracting the first and second multimodal features based on the multimodal information of video and audio data enables the multimodal information of video and audio data to assist in the recognition process of video background music, thereby improving the accuracy of video background music recognition.

[0209] This application embodiment segments audio data to obtain multiple audio segments. Using a second segment model within the audio multimodal model, features are extracted from these segments to obtain text and audio features for each segment. A second fusion model within the audio multimodal model fuses these text and audio features to obtain multimodal features for each segment. Finally, a second sequence model within the audio multimodal model performs temporal fusion of the multimodal features to obtain second multimodal features. By segmenting the audio and using different models within the audio multimodal model for feature extraction and fusion, the system achieves its technical objective of more effectively processing multimodal information in audio data and extracting and fusing features, thereby improving the efficiency of feature extraction and fusion for multimodal information in audio data.

[0210] As an optional approach, before processing the video data using a video multimodal model to obtain the first multimodal feature corresponding to the video data, the method further includes:

[0211] S9-1, obtain multiple positive and negative sample pairs, wherein the positive sample pairs in the positive and negative sample pairs consist of the same video and the corresponding music, and the negative sample pairs in the positive and negative sample pairs consist of the same video and the non-corresponding music.

[0212] S9-2 uses multiple positive and negative sample pairs to jointly train the initial video multimodal model and the initial audio multimodal model until the joint convergence condition is met, resulting in a trained video multimodal model and a trained audio multimodal model.

[0213] In an optional embodiment, positive and negative sample pairs may be, but are not limited to, during the training process of the model, the dataset used to train the model is divided into positive and negative sample pairs.

[0214] In an optional embodiment, positive sample pairs may, but are not limited to, consist of the same video and corresponding music.

[0215] In an optional embodiment, negative sample pairs may, but are not limited to, consist of the same video and non-corresponding music.

[0216] In an optional embodiment, the parameters of the video multimodal model and the audio multimodal model are first initialized before training the video multimodal model and the audio multimodal model.

[0217] Furthermore, positive sample pairs consisting of the same video and corresponding music and negative sample pairs consisting of the same video and non-corresponding music were obtained.

[0218] Next, repeat the following steps until the parameters of the video multimodal model and the audio multimodal model meet the convergence criteria:

[0219] Positive and negative sample pairs are input into the video multimodal model and the audio multimodal model for multimodal feature extraction. During the extraction process, modality regularization is used to prevent overfitting in the video and audio multimodal models. Domain alignment loss and cover group loss are used to measure the similarity of positive and negative sample pairs. Based on the calculated loss function, the parameters of the video and audio multimodal models are updated using the backpropagation algorithm to maximize the similarity of the video and audio multimodal models on positive sample pairs and minimize the similarity on negative sample pairs. Specifically, the domain alignment loss is obtained by aligning the audio features of the video with the audio features of the music, the text features of the video with the text features of the music, and the multimodal features of the video with the multimodal features of the music. The cover group loss is obtained from the music multimodal features.

[0220] Finally, the video multimodal model and audio multimodal model that meet the convergence condition are used as the trained video multimodal model and audio multimodal model.

[0221] In optional embodiments, modal regularization can, but is not limited to, improve the model's generalization ability by reducing its complexity, enabling the model to make more accurate predictions when faced with new data.

[0222] In optional embodiments, the domain alignment loss function can be, but is not limited to, the loss function used for domain adaptation tasks, in order to reduce the distributional differences between the source and target domains, thereby enabling the model trained on the source domain to generalize better to the target domain. The core idea of ​​domain alignment loss is to align the feature distributions of the two domains so that the model performs better in the target domain.

[0223] In an optional embodiment, the cover group loss function is a loss function used to address the long-tailed distribution problem, specifically to resolve class imbalance. In long-tailed distribution data, the number of samples in the minority class (head class) far exceeds that of the majority class (tail class), causing the model to favor learning the head class while ignoring the tail class. The cover group loss function helps the model better learn the features of the tail class by grouping samples and optimizing the representation within and between groups.

[0224] It should be noted that by constructing and jointly training positive and negative sample pairs, the accuracy and robustness of recognition are improved, and rich training data is provided for the deep learning of the model, thereby improving the model's generalization ability and recognition efficiency.

[0225] This application's embodiments obtain multiple positive and negative sample pairs. Positive sample pairs consist of the same video and corresponding music, while negative sample pairs consist of the same video and non-corresponding music. Using these multiple positive and negative sample pairs, an initial video multimodal model and an initial audio multimodal model are jointly trained until the joint convergence condition is met, resulting in trained video and audio multimodal models. By constructing and jointly training positive and negative sample pairs, the technical objective of providing rich training data for the model is achieved, thereby improving the model's generalization ability and recognition efficiency.

[0226] As an alternative, the above-mentioned video music recognition method can be applied to short video application software scenarios.

[0227] In an optional embodiment, this embodiment proposes a music recognition method based on multimodal data for the task of background music recognition in videos. In addition to the audio data in the video, OCR data of video frames and text data of the title are also extracted. By combining the above multimodal data content, the background music in the video is identified.

[0228] Further examples, such as Figure 5 As shown, Figure 5 (a) is a screenshot from a video, with audio of Z performing a live version of "Song X". The video title clearly states the song title "Song X", and the video also displays the corresponding lyrics text "Lyrics A". These texts can be directly matched with the song title and specific lyrics to be identified, such as... Figure 5 In (b), there is also the song title "Song X", the singer Z, and lyrics such as "Lyrics A", "Lyrics B", "Lyrics C", "Lyrics D", and "Lyrics E" from "Song X". This text modality data can also serve as a supplement to the audio data, assisting this embodiment in better identifying background music. It should be noted that in... Figure 5 In scenarios used for identification Figure 5 The music library for video music information (a) in the text is a music library that stores a large amount of audio data and the corresponding multimodal features of the audio data before video music recognition is performed. Therefore, when performing video music recognition, it is only necessary to obtain the multimodal features corresponding to the video to perform the corresponding retrieval in the music library. At the same time, the results of music retrieval based on the video can further adjust the music in the music library and the corresponding multimodal features to improve the accuracy and efficiency of music recognition through video.

[0229] In this embodiment, after identifying the background music used in the video in the background, a music link can be added to the video interface. When the user clicks on this music link, they will be redirected to the music playback interface in the music player to play the corresponding song content in its entirety.

[0230] To further illustrate, continue based on Figure 5 The scenarios that can be selected include, for example Figure 6 As shown, in Figure 6 (a) in the example shows the identification of the music corresponding to the video and the display of identification information 602, where identification information 602 indicates that the music in the video has been successfully identified. Figure 6 (a) After successful video and music recognition, as shown in the example Figure 6 As shown in (b), a music control 604 is displayed in the original video interface. The music control 604 contains a music tag and a music link. The music tag is used to prompt the user to identify the music information. Clicking the music control 604 will redirect the user to the music player on the music platform to play the music. After clicking the music control 604, as shown... Figure 6 As shown in (c), the user is redirected to the music player on the music platform, where the song content corresponding to the music identified through the video is played in its entirety. The user can pause playback at any time and the playback information 606 is displayed, which indicates that the identified song content will be played in its entirety.

[0231] It should be noted that this embodiment effectively utilizes multimodal data from videos (video titles / OCR / ASR / text, audio) and text data from music (song titles, lyrics) to perform song recognition. It integrates multimodal data (audio, text) from these two domains to identify songs in videos, and at the same time solves several problems of the aforementioned prior art. It is more in line with business needs and upgrades the retrieval from audio to audio to multimodal cover song recognition from video audio and text to song audio and text.

[0232] To further illustrate, the overall model architecture of this embodiment is a typical dual-tower structure, divided into a video tower model and a music tower model according to the data domain or domain, corresponding to the video source domain (feed domain) and the music library song domain (domain), as detailed below:

[0233] (1) The training method of video to music is adopted, and the domain alignment is achieved by using a dual-tower model. The video source (feed) side and the music side are processed by models that do not share parameters.

[0234] (2) Use video-music pairs to construct training data, instead of just using music data;

[0235] (3) In addition to the audio modality, a text modality was added to assist in cover song recognition:

[0236] a. Add title text, text generated by optical character recognition, or text generated by automatic speech recognition to the video;

[0237] b. Add title text or lyrics text to the music;

[0238] (4) The Multi-segment Model is used to process audio of different durations by dividing the feed and music into multiple segments and adding a sequence model for merging multiple segments to generate a global embedding. It also supports segment-to-segment retrieval and whole-to-whole retrieval.

[0239] Further examples, such as Figure 7 As shown, in the video tower model, the video data of feed (video) segments #1, #2 (not shown in the figure), ..., #N (not shown in the figure) are first split into audio data and text data. The text data includes the title, text data recognized by optical character recognition, and text data recognized by automatic speech recognition. Then, the audio data and text data are input into the feed (video) segment model to divide the audio data and text data into multiple segments. Then, the segmented segments are input into the feed (video) sequence model, and the feed (video) audio embedding, feed (video) text embedding, and feed (video) multimodal embedding are obtained through modal regularization.

[0240] For the music tower model, the video data of music fragment #1, music fragment #2 (not shown in the figure), ..., music fragment #N (not shown in the figure) are first split into audio data and text data. The text data includes the title and lyrics. Then, the audio data, title and lyrics are input into the music fragment model to divide the audio data and text data into multiple segments. The segmented segments are then input into the music sequence model, and music-audio embedding, music-text embedding, and music-multimodal embedding are obtained through modal regularization.

[0241] Finally, similarity is calculated using domain alignment loss and coverage group loss. In the domain alignment loss, the data is obtained by aligning the feed (video) audio embedding with the music-audio embedding, the feed (video) text embedding with the music-text embedding, and the feed (video) multimodal embedding with the music-multimodal embedding. The coverage group loss is obtained by aligning the data with the music multimodal embedding.

[0242] In an optional embodiment, the model structure of the Feed model and the Music model is the same, only the model parameters are different. In this embodiment, the music side is used as an example to introduce this model architecture from bottom to top.

[0243] To further illustrate, consider the Segment Model: After dividing a video / music segment into multiple segments (15 seconds each) at a temporal granularity, the audio of each segment and its corresponding lyrics or ASR text are input into the Segment Model to obtain the segment's three embeddings (audio embedding, text embedding, and multimodal embedding). For example... Figure 8 As shown, a single music segment model consists of an audio model and a text model. Features are extracted from the audio data and text data respectively to obtain audio embeddings and text embeddings. The two are then fused through a model to output a music multimodal embedding that integrates text and audio information. The text data includes the song title of song X and the lyrics A of song X.

[0244] Sequence Model: This model fuses the embeddings of multiple segments in a temporal sequence to obtain three global embeddings for the video or music (audio embedding, text embedding, and multimodal embedding). Ultimately, this embodiment uses the global multimodal embedding as the final representation of the video / music for subsequent retrieval and recognition tasks.

[0245] The recognition process is as follows Figure 9 and Figure 10 As shown, specifically:

[0246] like Figure 9 As shown, this is the offline stage. All the music in the music library is input into the music tower model to extract the music-side multimodal embeddings and add them to the music retrieval library to build the index offline.

[0247] like Figure 10As shown, this is the online stage. The video is input into the video tower model to process the video content, obtain the video-side multimodal embedding, and retrieve the nearest music multimodal embedding from the music retrieval library. The similarity between the two is calculated. If it is higher than the similarity threshold, it is considered a match, the corresponding music recognition process is completed and recognition is returned as successful; otherwise, recognition failure is returned.

[0248] This application's embodiments identify music in videos based on multimodal information in the video data, significantly improving the recall and precision of the identification results in scenarios with background music. Particularly for videos with poor audio quality, severe speed changes, or pitch distortion, the method of this embodiment can greatly improve the accuracy of background music identification in such videos.

[0249] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0250] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0251] According to another aspect of the embodiments of this application, a video music recognition apparatus for implementing the above-described video music recognition method is also provided. For example... Figure 11 As shown, the device includes:

[0252] The first acquisition unit 1102 is used to acquire video data of the music to be identified;

[0253] The second acquisition unit 1104 is used to acquire the first text feature corresponding to the video data and the first audio feature corresponding to the video data, wherein the first text feature is used to characterize the text information associated with the video data and the first audio feature is used to characterize the audio information in the video data.

[0254] The fusion unit 1106 is used to fuse the first text features and the first audio features to obtain the first multimodal features corresponding to the video data;

[0255] The first determining unit 1108 is used to determine audio data from a music database using the first multimodal feature, and to use the music corresponding to the audio data as the music identified in the video data;

[0256] Among them, the second multimodal feature corresponding to the audio data is matched with the first multimodal feature. The second multimodal feature is obtained by feature fusion of the second text feature and the second audio feature. The second text feature is used to represent the text information associated with the audio data, and the second audio feature is used to represent the audio information in the audio data.

[0257] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0258] As an optional solution, the second acquisition unit 1104 includes: a first extraction module, used to extract features from the video description information associated with the video data to obtain a first feature, wherein the first text feature includes the first feature.

[0259] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0260] As an optional solution, the extraction module includes: a first extraction submodule, used to extract features from the video title information of the video data to obtain title text features, wherein the first feature includes title text features and the video description information includes video title information; a second extraction submodule, used to extract features from the video tag information of the video data to obtain tag text features, wherein the first feature includes tag text features and the video description information includes video tag information; and a third extraction submodule, used to extract features from the video source information of the video data to obtain source text features, wherein the first feature includes source text features and the video description information includes video source information.

[0261] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0262] As an optional solution, the second acquisition unit 1104 includes: a second extraction module, used to extract features from the playback data in the video data to obtain a second feature, wherein the first text feature includes the second feature.

[0263] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0264] As an optional solution, the second extraction module includes: a first recognition submodule, used to perform optical character recognition on the playback data in the video data to obtain the first text corresponding to the playback data; and a fourth extraction submodule, used to extract features from the first text to obtain the screen text features, wherein the second feature includes the screen text features.

[0265] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0266] As an optional solution, the second extraction module includes: a second recognition submodule, used to perform automatic speech recognition on the audio playback data in the video data to obtain the second text corresponding to the audio playback data; and a fifth extraction submodule, used to extract features from the second text to obtain audio text features, wherein the second features include audio text features.

[0267] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0268] As an optional solution, the first determining unit 1108 includes: a first determining module, used to determine the music segment corresponding to the audio data as the music identified in the video data; and a second determining module, used to determine the music to which the music segment corresponding to the audio data belongs as the music identified in the video data.

[0269] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0270] As an optional solution, the first acquisition unit 1102 includes: a first acquisition module, used to acquire a first text feature, a first audio feature, and an identity feature corresponding to the video data, wherein the identity feature is used to represent identity information; a first fusion module, used to perform feature fusion on the first text feature, the first audio feature, and the identity feature to obtain a third multimodal feature corresponding to the video data; and a third determination module, used to use the third multimodal feature to determine the target audio data from the music database, and to use the music corresponding to the target audio data as the music identified in the video data; wherein the fourth multimodal feature corresponding to the target audio data matches the first multimodal feature, the fourth multimodal feature is obtained by feature fusion of the third text feature and the third audio feature, the third text feature is used to represent the text information associated with the target audio data, and the third audio feature is used to represent the audio information in the target audio data.

[0271] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0272] As an optional solution, the first determining unit 1108 includes: a retrieval module, used to retrieve the music multimodal feature that has the closest spatial distance to the first multimodal feature from the music retrieval library, wherein the spatial distance is negatively correlated with the similarity between the features; and a fourth determining module, used to determine the data corresponding to the music multimodal feature as audio data when the feature similarity between the music multimodal feature and the first multimodal feature is greater than a preset threshold.

[0273] For specific implementation examples, please refer to the examples shown in the above video music recognition method, which will not be repeated here.

[0274] According to another aspect of the embodiments of this application, another video music recognition apparatus is also provided for implementing the other video music recognition method described above. For example... Figure 12 As shown, the device includes:

[0275] The first processing unit 1202 is used to process multiple audio reference data using an audio multimodal model to obtain the multimodal features corresponding to each audio reference data in the multiple audio reference data.

[0276] Add unit 1204 to add the multimodal features corresponding to each audio reference data to the music database;

[0277] The third acquisition unit 1206 is used to acquire video data of the music to be identified, wherein multiple audio reference data include audio data;

[0278] The second processing unit 1208 is used to process video data using a video multimodal model to obtain the first multimodal feature corresponding to the video data, wherein the audio multimodal model and the video multimodal model are jointly trained.

[0279] The second determining unit 1210 is used to determine audio data from the music database using the first multimodal features, and to use the music corresponding to the audio data as the music identified in the video data;

[0280] Among them, the second multimodal feature corresponding to the audio data is matched with the first multimodal feature. The second multimodal feature is obtained by feature fusion of the second text feature and the second audio feature. The second text feature is used to represent the text information associated with the audio data, and the second audio feature is used to represent the audio information in the audio data.

[0281] For specific implementation examples, please refer to the example shown in the other video music recognition method described above, which will not be repeated here.

[0282] As an optional solution, the second processing unit 1208 includes: a first processing module for segmenting video data to obtain multiple video segments; a third extraction module for extracting features from the multiple video segments using a first segment model in the video multimodal model to obtain text features and audio features corresponding to each video segment; a second fusion module for fusing the text features and audio features corresponding to each video segment using a first fusion model in the video multimodal model to obtain multimodal features corresponding to each video segment; and for temporally fusing the multimodal features corresponding to the multiple video segments using a first sequence model in the video multimodal model to obtain first multimodal features.

[0283] For specific implementation examples, please refer to the example shown in the other video music recognition method described above, which will not be repeated here.

[0284] As an optional solution, the first processing unit 1202 includes: a second processing module for segmenting audio data to obtain multiple audio segments; a fourth extraction module for extracting features from the multiple audio segments using a second segment model in the audio multimodal model to obtain text features and audio features corresponding to each audio segment; and a third fusion module for performing temporal fusion of the text features and audio features corresponding to each audio segment using a second fusion model in the audio multimodal model to obtain the multimodal features corresponding to each audio segment; and a fourth extraction module for performing temporal fusion of the multimodal features corresponding to the multiple audio segments using a second sequence model in the audio multimodal model to obtain the second multimodal features.

[0285] For specific implementation examples, please refer to the example shown in the other video music recognition method described above, which will not be repeated here.

[0286] As an optional solution, the second processing unit 1208 includes: a second acquisition module, used to acquire multiple positive and negative sample pairs, wherein the positive sample pairs in the positive and negative sample pairs consist of the same video and corresponding music, and the negative sample pairs in the positive and negative sample pairs consist of the same video and non-corresponding music; and a training module, used to jointly train the initial video multimodal model and the initial audio multimodal model using the multiple positive and negative sample pairs until the joint convergence condition is met, thereby obtaining the trained video multimodal model and the trained audio multimodal model.

[0287] For specific implementation examples, please refer to the example shown in the other video music recognition method described above, which will not be repeated here.

[0288] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described video music recognition method is also provided. This electronic device may, but is not limited to, […]. Figure 1 The user equipment 102 or server 112 shown in the figure, in this embodiment, is taken as an example of an electronic device, namely user equipment 102. Further, as shown in the figure... Figure 13 As shown, the electronic device includes a memory 1302 and a processor 1304. The memory 1002 stores a computer program, and the processor 1304 is configured to execute the steps of any of the above method embodiments via the computer program.

[0289] In an optional embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0290] In an optional embodiment, the processor described above may be configured to perform the following steps via a computer program:

[0291] S1, Obtain the video data of the music to be identified;

[0292] S2, obtain the first text feature corresponding to the video data and obtain the first audio feature corresponding to the video data, wherein the first text feature is used to characterize the text information associated with the video data and the first audio feature is used to characterize the audio information in the video data;

[0293] S3, perform feature fusion on the first text feature and the first audio feature to obtain the first multimodal feature corresponding to the video data;

[0294] S4. Using the first multimodal feature, determine the audio data from the music database, and use the music corresponding to the audio data as the music identified in the video data;

[0295] Among them, the second multimodal feature corresponding to the audio data is matched with the first multimodal feature. The second multimodal feature is obtained by feature fusion of the second text feature and the second audio feature. The second text feature is used to represent the text information associated with the audio data, and the second audio feature is used to represent the audio information in the audio data.

[0296] Alternatively, as those skilled in the art will understand, Figure 13 The structure shown is for illustrative purposes only. Figure 13 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 13 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 13 The different configurations shown.

[0297] The memory 1302 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video music recognition method and apparatus in this embodiment. The processor 1304 executes various functional applications and data processing by running the software programs and modules stored in the memory 1302, thereby realizing the aforementioned video music recognition method. The memory 1302 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1302 may further include memory remotely located relative to the processor 1304, and these remote memories can be connected to electronic devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1302 may be used, but is not limited to, to store information such as video data, first text features, first audio features, first multimodal features, audio data, and second multimodal features. As an example, such as... Figure 13 As shown, the memory 1302 may include, but is not limited to, the first acquisition unit 1102, the second acquisition unit 1104, the fusion unit 1106, and the first determination unit 1108 in the aforementioned video music recognition device. It may also include, but is not limited to, the first processing unit 1202, the adding unit 1204, the third acquisition unit 1206, the second processing unit 1208, and the second determination unit 1210 in the aforementioned other video music recognition device (not shown in the figure). Furthermore, it may include, but is not limited to, other module units in the aforementioned video music recognition device, which will not be elaborated upon in this example.

[0298] Optionally, the transmission device 1306 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1306 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1306 is a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0299] In addition, the aforementioned electronic device also includes: a display 1308 for displaying the aforementioned video data, first text features, first audio features, first multimodal features, audio data, second multimodal features, and other information; and a connection bus 1310 for connecting the various module components in the aforementioned electronic device.

[0300] In other embodiments, the aforementioned user equipment or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer network, and any form of computing device, such as a server, user equipment, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.

[0301] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.

[0302] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0303] It should be noted that the computer system of the electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0304] A computer system includes a Central Processing Unit (CPU), which performs various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) or loaded from RAM. ROM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output interfaces (I / O interfaces) are also connected to the bus.

[0305] The following components are connected to the input / output interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processing via a network such as the Internet. Drives are also connected to the input / output interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.

[0306] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions defined in the system of this application.

[0307] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of a computer device reads computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.

[0308] In an optional embodiment, the computer-readable storage medium described above may be configured to store a computer program for performing the following steps:

[0309] S1, Obtain the video data of the music to be identified;

[0310] S2, obtain the first text feature corresponding to the video data and obtain the first audio feature corresponding to the video data, wherein the first text feature is used to characterize the text information associated with the video data and the first audio feature is used to characterize the audio information in the video data;

[0311] S3, perform feature fusion on the first text feature and the first audio feature to obtain the first multimodal feature corresponding to the video data;

[0312] S4. Using the first multimodal feature, determine the audio data from the music database, and use the music corresponding to the audio data as the music identified in the video data;

[0313] Among them, the second multimodal feature corresponding to the audio data is matched with the first multimodal feature. The second multimodal feature is obtained by feature fusion of the second text feature and the second audio feature. The second text feature is used to represent the text information associated with the audio data, and the second audio feature is used to represent the audio information in the audio data.

[0314] Optionally, in embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0315] In optional embodiments, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware of an electronic device. The program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0316] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0317] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0318] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0319] In the several embodiments provided in this application, it should be understood that the disclosed user equipment can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0320] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0321] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0322] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for recognizing music in a video, characterized in that, include: Obtain the video data of the music to be identified; Obtain a first text feature corresponding to the video data and obtain a first audio feature corresponding to the video data, wherein the first text feature is used to characterize the text information associated with the video data and the first audio feature is used to characterize the audio information in the video data; The first text feature and the first audio feature are fused to obtain the first multimodal feature corresponding to the video data; Using the first multimodal feature, audio data is determined from the music database, and the music corresponding to the audio data is used as the music identified in the video data; The second multimodal feature corresponding to the audio data matches the first multimodal feature. The second multimodal feature is obtained by fusing the second text feature and the second audio feature. The second text feature is used to characterize the text information associated with the audio data, and the second audio feature is used to characterize the audio information in the audio data.

2. The method according to claim 1, characterized in that, The step of obtaining the first text feature corresponding to the video data includes: Feature extraction is performed on the video description information associated with the video data to obtain the first feature, wherein the first text feature includes the first feature.

3. The method according to claim 2, characterized in that, The first feature is obtained by extracting features from the video description information associated with the video data, including at least one of the following: Feature extraction is performed on the video title information of the video data to obtain title text features, wherein the first feature includes the title text features, and the video description information includes the video title information; Feature extraction is performed on the video tag information of the video data to obtain tag text features, wherein the first feature includes the tag text features, and the video description information includes the video tag information; Feature extraction is performed on the video source information of the video data to obtain source text features, wherein the first feature includes the source text features, and the video description information includes the video source information.

4. The method according to claim 1, characterized in that, The step of obtaining the first text feature corresponding to the video data includes: Feature extraction is performed on the playback data in the video data to obtain the second feature, wherein the first text feature includes the second feature.

5. The method according to claim 4, characterized in that, The step of extracting features from the playback data in the video data to obtain the second feature includes: Optical character recognition is performed on the playback data in the video data to obtain the first text corresponding to the playback data. Feature extraction is performed on the first text to obtain the image text features, wherein the second feature includes the image text features.

6. The method according to claim 4, characterized in that, The step of extracting features from the playback data in the video data to obtain the second feature includes: Automatic speech recognition is performed on the audio playback data in the video data to obtain the second text corresponding to the audio playback data; Feature extraction is performed on the second text to obtain audio text features, wherein the second features include the audio text features.

7. The method according to any one of claims 1 to 6, characterized in that, The step of using the music corresponding to the audio data as the music identified in the video data includes: The music segment corresponding to the audio data is used as the music identified in the video data; or... The music to which the music segment corresponding to the audio data belongs is used as the music identified in the video data.

8. The method according to any one of claims 1 to 6, characterized in that, If the video data contains identity information, after obtaining the video data of the music to be identified, the method further includes: Obtain the first text feature, the first audio feature, and the identity feature corresponding to the video data, wherein the identity feature is used to characterize the identity information; The first text feature, the first audio feature, and the identity feature are fused to obtain the third multimodal feature corresponding to the video data; Using the third multimodal feature, target audio data is determined from the music database, and the music corresponding to the target audio data is used as the music identified in the video data; The fourth multimodal feature corresponding to the target audio data matches the first multimodal feature. The fourth multimodal feature is obtained by fusing the third text feature and the third audio feature. The third text feature is used to characterize the text information associated with the target audio data, and the third audio feature is used to characterize the audio information in the target audio data.

9. The method according to any one of claims 1 to 6, characterized in that, The step of determining audio data from the music database using the first multimodal feature includes: The music multimodal feature that is spatially closest to the first multimodal feature is retrieved from the music retrieval database, wherein the spatial distance is negatively correlated with the similarity between the features; If the feature similarity between the music multimodal feature and the first multimodal feature is greater than a preset threshold, the data corresponding to the music multimodal feature is determined as the audio data.

10. A method for recognizing music in a video, characterized in that, include: By using an audio multimodal model, multiple audio reference data are processed to obtain the multimodal features corresponding to each audio reference data in the multiple audio reference data. The multimodal features corresponding to each of the audio reference data are added to the music database; Acquire video data of the music to be identified, wherein the plurality of audio reference data includes the audio data; The video data is processed using a video multimodal model to obtain the first multimodal feature corresponding to the video data, wherein the audio multimodal model and the video multimodal model are jointly trained. Using the first multimodal feature, audio data is determined from the music database, and the music corresponding to the audio data is used as the music identified in the video data; The second multimodal feature corresponding to the audio data matches the first multimodal feature. The second multimodal feature is obtained by fusing the second text feature and the second audio feature. The second text feature is used to characterize the text information associated with the audio data, and the second audio feature is used to characterize the audio information in the audio data.

11. The method according to claim 10, characterized in that, The step of processing the video data using a video multimodal model to obtain the first multimodal feature corresponding to the video data includes: The video data is segmented to obtain multiple video data segments; By using the first segment model in the video multimodal model, feature extraction is performed on the multiple video segments to obtain the text features and audio features corresponding to each video segment. The first fusion model in the video multimodal model is used to fuse the text features and audio features corresponding to each video data segment to obtain the multimodal features corresponding to each video data segment. The first multimodal feature is obtained by temporally fusing the multimodal features corresponding to the multiple video data segments using the first sequence model in the video multimodal model.

12. The method according to claim 10 or 11, characterized in that, In the process of processing multiple audio reference data using an audio multimodal model to obtain the multimodal features corresponding to each audio reference data in the multiple audio reference data, the method further includes: The audio data is segmented to obtain multiple audio data segments; By using the second segment model in the audio multimodal model, feature extraction is performed on the multiple audio segments to obtain the text features and audio features corresponding to each audio segment. The second fusion model in the audio multimodal model is used to fuse the text features and audio features corresponding to each audio data segment to obtain the multimodal features corresponding to each audio data segment. The second multimodal feature is obtained by temporally fusing the multimodal features corresponding to the multiple audio data segments using the second sequence model in the audio multimodal model.

13. The method according to claim 10 or 11, characterized in that, Before processing the video data using a video multimodal model to obtain the first multimodal feature corresponding to the video data, the method further includes: Multiple positive and negative sample pairs are obtained, wherein the positive sample pairs in the positive and negative sample pairs consist of the same video and the corresponding music, and the negative sample pairs in the positive and negative sample pairs consist of the same video and the non-corresponding music; Using the multiple positive and negative sample pairs, the initial video multimodal model and the initial audio multimodal model are jointly trained until the joint convergence condition is met, resulting in a trained video multimodal model and a trained audio multimodal model.

14. A video music recognition device, characterized in that, include: The first acquisition unit is used to acquire video data of the music to be identified; The second acquisition unit is used to acquire a first text feature corresponding to the video data and acquire a first audio feature corresponding to the video data, wherein the first text feature is used to characterize the text information associated with the video data and the first audio feature is used to characterize the audio information in the video data. The fusion unit is used to fuse the first text features and the first audio features to obtain the first multimodal features corresponding to the video data; The first determining unit is used to determine audio data from a music database using the first multimodal feature, and to use the music corresponding to the audio data as the music identified in the video data; The second multimodal feature corresponding to the audio data matches the first multimodal feature. The second multimodal feature is obtained by fusing the second text feature and the second audio feature. The second text feature is used to characterize the text information associated with the audio data, and the second audio feature is used to characterize the audio information in the audio data.

15. A device for recognizing video music, characterized in that, include: The first processing unit is used to process multiple audio reference data using an audio multimodal model to obtain the multimodal features corresponding to each audio reference data in the above-mentioned multiple audio reference data. Add a unit to add the multimodal features corresponding to the above audio reference data to the music database; The third acquisition unit is used to acquire video data of the music to be identified, wherein the aforementioned multiple audio reference data include the aforementioned audio data; The second processing unit is used to process the video data using a video multimodal model to obtain the first multimodal feature corresponding to the video data, wherein the audio multimodal model and the video multimodal model are jointly trained. The second determining unit is used to determine audio data from the music database using the first multimodal feature mentioned above, and to use the music corresponding to the audio data as the music identified in the video data. The second multimodal feature corresponding to the audio data matches the first multimodal feature. The second multimodal feature is obtained by fusing the second text feature and the second audio feature. The second text feature is used to characterize the text information associated with the audio data, and the second audio feature is used to characterize the audio information in the audio data.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program, wherein the computer program, when executed by an electronic device, performs the method according to any one of claims 1 to 9, or 10 to 13.

17. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1 to 9, or 10 to 13.

18. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 9, or 10 to 13, through the computer program.