Multimedia aggregation method and apparatus, electronic device, and storage medium
Patent Information
- Application Number
- PCT/CN2024/138296
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-04
- Filing Date
- 2024-12-11
- Publication Date
- 2025-10-02
AI Technical Summary
The song libraries of large music platforms contain a huge number of songs. Existing technologies make it difficult to directly use clustering methods for effective aggregation, and there are problems with the transitivity of dissimilar songs within a song group and incomplete recall.
By obtaining the correlation between songs and using multimodal methods such as audio fingerprint recognition, song title recognition and lyrics recognition, a sparse adjacency matrix and connectivity graph are constructed, preliminary partitioning and denoising are performed, and DBSCAN and graph segmentation algorithms are combined for precise clustering to optimize the correlation between songs.
It achieves accurate aggregation of songs in the song library, improves the recall rate and aggregation accuracy of the song group, effectively removes noise, and ensures the similarity of songs within the song group.
Smart Images

Figure CN2024138296_02102025_PF_FP_ABST
Abstract
Description
Multimedia aggregation method, device, electronic device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application number 202410244712.1, filed on March 4, 2024, entitled “Multimedia Aggregation Method, Device, Electronic Device and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to the field of data processing technology, and in particular to a multimedia aggregation method, device, electronic device, and storage medium. Background Art
[0004] Large music platforms usually have a song library with a massive song. When the songs in the song library need to be aggregated, related technologies usually cannot directly use clustering to aggregate the massive songs in the song library. Summary of the Invention
[0005] The present disclosure provides a multimedia aggregation method, device, electronic device and storage medium.
[0006] According to a first aspect of the present disclosure, a multimedia aggregation method is provided, the method comprising:
[0007] Acquire a multimedia set to be aggregated, where the multimedia set to be aggregated includes a plurality of multimedia to be aggregated;
[0008] Obtaining correlations between the plurality of multimedia to be aggregated, and dividing the multimedia set to be aggregated into a plurality of multimedia subsets to be aggregated based on the correlations;
[0009] Clustering processing is performed on the multimedia to be aggregated in each of the multimedia subsets to be aggregated to obtain an aggregation result of the multiple multimedia to be aggregated.
[0010] According to a second aspect of the present disclosure, a multimedia aggregation device is provided, the device comprising:
[0011] A multimedia set to be aggregated obtaining module, configured to obtain a multimedia set to be aggregated, wherein the multimedia set to be aggregated includes a plurality of multimedia to be aggregated;
[0012] a correlation obtaining module, configured to obtain correlations between the plurality of multimedia to be aggregated, and divide the multimedia set to be aggregated into a plurality of multimedia subsets to be aggregated based on the correlations;
[0013] The clustering processing module is used to perform clustering processing on the multimedia to be aggregated in each of the multimedia subsets to be aggregated, and obtain an aggregation result of the multiple multimedia to be aggregated.
[0014] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the program.
[0015] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above method of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0017] FIG1 is a flowchart of a multimedia aggregation method provided by an exemplary embodiment of the present disclosure;
[0018] FIG2 is a schematic block diagram of functional modules of a multimedia aggregation device provided by an exemplary embodiment of the present disclosure;
[0019] FIG3 is a structural block diagram of an electronic device provided by an exemplary embodiment of the present disclosure;
[0020] FIG4 is a structural block diagram of a computer system provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0022] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0023] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0024] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0026] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0027] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0028] As an optional but non-limiting implementation method, in response to receiving the user's active request, the method of sending a prompt message to the user can be, for example, a pop-up window, and the prompt message can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understandable that the above notification and the process of obtaining user authorization are only illustrative and do not constitute a limitation on the implementation method of the present disclosure. Other methods that meet relevant laws and regulations can also be applied to the implementation method of the present disclosure.
[0029] Since large music platforms have song libraries with massive data, users often need to find songs that are similar to their favorite songs or belong to the same category, or recommend songs in the same category to users based on their favorite songs. This requires aggregating the songs in the song library to meet user needs.
[0030] However, when related technologies aggregate songs in a song library, the song library is usually massive, with a large amount of data, making it impossible to directly use clustering to aggregate the songs in the song library. In addition, when related technologies aggregate songs in a song library, there may be errors in the songs aggregated into a group. For example, when using an audio similarity judgment model, it is impossible to output reliable judgment results for types of songs such as medleys, repeated melodies, and white noise. This leads to transitivity issues within the aggregated song groups. For example, a medley can be successfully grouped, resulting in multiple dissimilar songs in the same song group.
[0031] Furthermore, related technologies may not be able to fully retrieve all similar songs when recalling songs from a song library. The reason is as follows: For songs with numerous duplicate names, limited computing power means that only the top 100 songs can be retrieved for similarity analysis. Furthermore, songs with similar melodies but different titles cannot be aggregated. If a song group already contains low-quality audio, such as medleys or repeated melodies, the aggregation process is likely to fail, preventing successful aggregation.
[0032] Therefore, in order to improve the accuracy of song aggregation, the embodiments provided by the present disclosure can be used to consider the different quality and diversity of songs in the song library, thereby improving the recall rate of the song group while ensuring accuracy. The recall in the embodiments can be performed by fingerprint recognition retrieval or fuzzy search of song titles to recall all identical or similar songs.
[0033] It should be noted that the embodiments of the present disclosure are described using songs as an example. The embodiments may also aggregate multimedia such as audio or audio and video, but the embodiments are not limited thereto.
[0034] In an embodiment, for example, each song in a song library can be used as a song to be aggregated. For a song to be aggregated, the song can be used as a target song to be aggregated. Songs related to the target song to be aggregated can be retrieved from the song library by using cover recognition. Similarly, songs related to the remaining songs to be aggregated can be obtained in the song library. The song library in the embodiment can include multiple songs.
[0035] In this embodiment, songs related to the target song to be aggregated can also be obtained from the song library in a multimodal manner. For example, songs related to the target song to be aggregated can be retrieved from the song library through audio fingerprint recognition, song title recognition, or lyrics recognition. Similarly, songs related to the remaining songs to be aggregated can be retrieved from the song library separately in the multimodal manner.
[0036] In this embodiment, by searching the song library for songs related to each song to be aggregated, for example, based on correlation, songs that are relatively related to the top K songs to be aggregated can be obtained. Then, a connection relationship can be established between the related songs to be aggregated, for example, by connecting the related songs with edges in a connectivity graph. In this embodiment, a sparse adjacency matrix can also be constructed to represent the correlation between the songs to be aggregated in the song library, but the embodiment is not limited to this. K is a positive integer that can be set as needed, but the embodiment is not limited to this.
[0037] By obtaining songs related to each song to be aggregated from the song library as described above, it is possible that some related songs to be aggregated are not particularly related. To further optimize the correlation between the songs to be aggregated, optimization can be performed using similarity. For example, the similarity between two related songs to be aggregated can be obtained. If the similarity is below a threshold, such as 0.6 (this embodiment is not limited to this), the correlation between the two songs to be aggregated can be cancelled; otherwise, the correlation between the two songs to be aggregated can be retained.
[0038] When establishing the correlation between the songs to be aggregated by means of a connected graph, the edges between the songs to be aggregated whose similarity is lower than a threshold value can be removed to achieve the purpose of removing redundant edges and realize the optimization of the correlation between the songs to be aggregated.
[0039] Through the processing in the above manner, songs that are relevant to each song to be aggregated can be retrieved from the song library. In this way, the songs to be aggregated in the song library can be divided into multiple categories according to the correlation between the songs. When dividing through a connected graph, it is equivalent to dividing the connected graph into multiple connected blocks, each connected block is equivalent to a category, thereby realizing coarse clustering of the songs to be aggregated in the song library.
[0040] For each of the above-divided categories or connected blocks, if the number of songs to be aggregated contained in the category or connected block is not greater than M, the category or connected block can be directly used as the final classification result, that is, the songs to be aggregated contained in the category or connected block can be regarded as a category. Wherein, M is a positive integer and can be set as needed, and the embodiment is not limited to this.
[0041] In addition, if the number of songs to be aggregated contained in the above-mentioned category or connected block is greater than M, it means that the number of songs to be aggregated contained in the category or the connected block is large, and there may be different types of songs to be aggregated contained in the category or the connected block, so further processing is required.
[0042] Specifically, the above-mentioned category or connected block containing more than M songs to be aggregated can be called an initial category. The songs to be aggregated in the initial category can be denoised. For example, the songs to be aggregated in the initial category can be identified for musical components to determine the proportion of musical components in each song to be aggregated. For example, some songs to be aggregated may be audio of an interview program, where the background may be music but the main content is conversation. That is, the proportion of conversation is large and greater than a certain threshold. The interview program audio is mistakenly classified into the song library. Therefore, it is necessary to remove such songs to be aggregated to achieve the purpose of denoising the songs to be aggregated in the above-mentioned initial cluster.
[0043] After denoising the above-mentioned initial categories, each denoised initial category can also be accurately clustered. For example, a density-based spatial clustering algorithm (DBSCAN) can be used for accurate clustering, which will not be described in detail here. In this way, accurate clustering of the above-mentioned denoised initial categories can be achieved. The embodiment can further remove abnormal nodes in the clusters obtained by the above-mentioned accurate clustering through a graph segmentation algorithm (Spectral Clustering) to obtain the final aggregation result of the songs to be aggregated in the above-mentioned song library. The aggregation result may include multiple different song categories.
[0044] Based on the above embodiments, the present disclosure further provides a multimedia aggregation method, as shown in FIG1 , which may include the following steps:
[0045] In step S110 , a multimedia set to be aggregated is obtained, where the multimedia set to be aggregated includes a plurality of multimedia to be aggregated.
[0046] In the embodiment, the multimedia set to be aggregated may be part or all of the multimedia in the multimedia library, and the multimedia to be aggregated in the multimedia set may be used as the multimedia to be aggregated.
[0047] The multimedia to be aggregated may specifically be video, audio, or audio and video data. In the embodiment, the multimedia to be aggregated is music as an example. The multimedia set to be aggregated may be equivalent to the song library in the above embodiment, but the embodiment is not limited thereto.
[0048] In step S120 , correlations between a plurality of multimedia to be aggregated are obtained, and the multimedia set to be aggregated is divided into a plurality of multimedia sub-sets to be aggregated based on the correlations.
[0049] In the embodiment, the multimedia to be aggregated in the multimedia set to be aggregated may be preliminarily classified into categories to obtain multiple multimedia sets to be aggregated.
[0050] Specifically, the multimedia collection to be aggregated can be searched for other multimedia to be aggregated that is related to each multimedia to be aggregated. Feature information of each multimedia to be aggregated can be extracted and, based on this feature information, each multimedia to be aggregated can be searched for in the multimedia collection to be aggregated. The correlation between the multiple multimedia to be aggregated can be determined based on the search results. This feature information can be, for example, audio fingerprints, song titles, or lyrics. For example, multimedia to be aggregated that is related to each multimedia to be aggregated can be found based on audio fingerprints, song titles, or content.
[0051] Taking music as an example of multimedia to be aggregated, other music related to each music can be found in the multimedia library to be aggregated based on the correlation between the music. For example, for a certain target music, the correlation between the song title of the target music and the song titles of other music, specifically the similarity between the song titles, etc., can be used to find other music related to the target music, and establish connections between the related music. For example, connections can be established through the edges connecting the songs in a connectivity graph.
[0052] Similarly, correlations can be established based on the correlations between the lyrics of the music, or by fingerprinting the songs. The correlations obtained through song titles, lyrics, and fingerprinting can be combined to obtain correlations between each piece of music and other pieces of music.
[0053] In an embodiment, the correlation relationship between songs can be established in a multimodal manner through the above-mentioned method. Specifically, the weights between the multimodalities can also be considered. For example, different weights can be set between fingerprint recognition, song titles and lyrics to control the proportion of correlation between the music. For example, when the weights corresponding to the above-mentioned fingerprint recognition, song titles and lyrics are 0.5, 0.3 and 0.2 respectively, then among the other songs that are correlated with the target song, the proportion of songs that are correlated with the target song obtained through fingerprint recognition is 50%, the proportion of songs that are correlated with the target song obtained through song title matching is 30%, and the proportion of songs that are correlated with the target song obtained through lyrics matching is 20%. The size of the relevant weights, etc., can be set as needed, and the embodiment is not limited to this.
[0054] In addition, in the embodiment, multimedia to be aggregated that is relevant to each multimedia to be aggregated can be obtained through cover song recognition by means of a cover song model. For example, the top K multimedia to be aggregated that are respectively relevant to each multimedia to be aggregated can be obtained based on the correlation.
[0055] Through the above method, the correlation between each multimedia to be aggregated in the multimedia set can be obtained, and connections between the multimedia to be aggregated can be established based on the correlation. For example, a connected graph can be formed, with each multimedia to be aggregated as a vertex, and related multimedia to be aggregated can be connected through edges. In this way, the multimedia to be aggregated in the multimedia set to be aggregated can be divided into multiple categories, without any connection between the categories, that is, without edge connections.
[0056] In step S130 , clustering processing is performed on the multimedia to be aggregated in each of the multimedia subsets to be aggregated to obtain an aggregation result of the plurality of multimedia to be aggregated.
[0057] After performing the above-mentioned preliminary aggregation on the multimedia to be aggregated in the multimedia set to be aggregated, multiple categories can be obtained, each category is equivalent to a subset of multimedia to be aggregated, and then by further clustering the multimedia to be aggregated in each multimedia subset to be aggregated, the aggregation result of the above-mentioned multimedia to be aggregated can be obtained.
[0058] The multimedia aggregation method provided by the embodiments of the present disclosure obtains a multimedia set to be aggregated and obtains the correlation between the multimedia to be aggregated in the multimedia set to be aggregated, thereby dividing the multimedia set to be aggregated into multiple multimedia subsets to be aggregated. By clustering the multimedia to be aggregated in the multimedia subsets to be aggregated, an aggregation result of the multimedia to be aggregated in the multimedia set to be aggregated can be obtained. In this way, preliminary aggregation of the multimedia to be aggregated is achieved through the correlation between the multimedia to be aggregated, and by further clustering each of the obtained multimedia subsets to be aggregated, accurate aggregation of the multiple multimedia to be aggregated in the set to be aggregated can be achieved. In addition, accurate aggregation can be achieved for songs of types such as medleys, repeated melodies, and white noise in the song library.
[0059] Based on the above embodiment, in another embodiment provided by the present disclosure, the above step S120 may further include the following steps:
[0060] In step S121 , a target sparse adjacency matrix is constructed based on the correlation.
[0061] The target sparse adjacency matrix represents the correlation between the multimedia to be aggregated.
[0062] In an embodiment, the correlation between the multimedia to be clustered in the multimedia set to be clustered can be determined by the above-mentioned multimodal method or cover song recognition method. The correlation can be represented by a numerical value between 0 and 1, for example. The larger the numerical value, the higher the correlation between the multimedia to be clustered. In this way, when the multimedia set to be clustered contains N multimedia to be clustered, an N×N matrix can be constructed, which can specifically be a target sparse adjacency matrix. The target sparse adjacency matrix can represent the correlation between the multimedia to be clustered. The correlation between any two multimedia to be clustered can be represented by their correlation.
[0063] In step S122, the similarity between each multimedia to be aggregated in the target sparse adjacency matrix is obtained, and the target correlation between each multimedia to be aggregated is re-determined based on the similarity, and the multimedia set to be aggregated is divided into multiple multimedia subsets to be aggregated based on the target correlation.
[0064] When constructing the target sparse adjacency matrix through multimodal or cover song recognition methods and obtaining the correlation between the multimedia to be aggregated, these correlations can also be optimized by obtaining the similarities between the multimedia to be aggregated. Correlations with low similarity between the multimedia to be aggregated are removed, thereby optimizing the correlations between the multimedia to be aggregated and obtaining more accurate target correlations. In this way, the multimedia collection to be aggregated can be divided into multiple subsets of multimedia to be aggregated based on the target correlations. Each subset of multimedia to be clustered is equivalent to a category, achieving preliminary aggregation of the multimedia to be clustered.
[0065] In an embodiment, in the process of re-determining the target correlation between the multimedia to be aggregated based on the aforementioned similarity, a connectivity graph corresponding to the target sparse adjacency matrix can be obtained, and based on the similarity between the multimedia to be aggregated, edges in the connectivity graph can be removed or retained to obtain a target connectivity graph, and the target correlation between the multimedia to be aggregated can be determined based on the target connectivity graph. The connectivity graph includes a plurality of vertices and edges between the vertices, where the vertices represent the multimedia to be aggregated, and the edges represent the correlation between the multimedia to be aggregated.
[0066] In an embodiment, a connectivity graph can be constructed by the correlation between each multimedia to be clustered in the above-mentioned target sparse adjacency matrix. For example, each multimedia to be clustered is taken as a vertex, and the vertices corresponding to two multimedia to be clustered whose correlation is greater than a certain threshold are connected through edges. In this way, a connection relationship between multiple vertices can be obtained, and the connection relationship represents the correlation relationship between the multimedia to be clustered. By obtaining the similarity between the multimedia to be clustered, when the similarity between the two multimedia to be clustered is lower than a certain threshold, the edges between the vertices between the two can be removed, that is, the connection relationship between the multimedia to be clustered with lower similarity can be removed, while retaining the edges between the multimedia to be clustered with higher similarity, thereby achieving the purpose of re-determining the target correlation relationship between each multimedia to be clustered based on the similarity. It should be noted that the similarity between the multimedia to be clustered can be obtained by existing methods, which will not be repeated here.
[0067] Based on the above embodiments, in the embodiments provided by the present disclosure, when the multimedia set to be aggregated is divided into multiple multimedia subsets to be aggregated based on the target correlation relationship, the connection relationship between the vertices in the target connectivity graph can be obtained, and the target connectivity graph can be divided into multiple connectivity blocks based on the connection relationship. In this way, the multimedia set to be aggregated can be divided into multiple multimedia subsets to be aggregated based on the multiple connectivity blocks.
[0068] In the embodiment, when the vertices in the connectivity graph are connected, that is, when there are edges between the vertices, it means that the multimedia to be clustered corresponding to the vertices are highly correlated and belong to the same category. When there are no edges between the vertices, it means that the multimedia to be clustered corresponding to the vertices are less correlated and do not belong to the same category. Thus, in the connectivity graph obtained above, multiple connected blocks are obtained, where there are no connections between the connected blocks, that is, no edges; and where the vertices within the connected blocks are connected, that is, edges exist. In this way, each connected block can be treated as a category, and multiple categories can be obtained. Each category can be treated as a subset of multimedia to be clustered, thereby achieving the purpose of dividing the multimedia set to be clustered into multiple subsets of multimedia to be clustered based on multiple connected blocks.
[0069] In this embodiment, for example, the priority of the multimedia to be aggregated can be set based on factors such as its popularity or importance. Different priorities can correspond to different weights. High-priority multimedia to be aggregated can be recalled first, while low-priority multimedia to be aggregated, such as long-tail, low-quality songs, can be "cold-treated." For example, high-priority multimedia to be aggregated can be re-precisely clustered, while low-priority multimedia to be aggregated can be aggregated together, reducing or eliminating the need for re-precise clustering of low-priority multimedia to be aggregated, thereby improving aggregation efficiency.
[0070] In an embodiment provided herein, to cluster the multimedia to be aggregated in each subset of multimedia to be aggregated, a target connected block from a plurality of connected blocks is obtained, the multimedia to be aggregated in the target connected block is denoised, and clustering is performed on the multimedia to be aggregated in the denoised target connected block. The number of multimedia to be aggregated in the target connected block is greater than a threshold.
[0071] In an embodiment, the number of multimedia to be clustered contained in the above-mentioned connected block can be detected. If the number of multimedia to be clustered contained in the connected block is greater than a threshold, it indicates that the number of multimedia to be clustered contained in the connected block is large, there may be noise, and the connected block may also contain multiple categories. Therefore, the multimedia to be clustered in the connected block can be denoised, and the multimedia to be clustered in the denoised connected block can be clustered. For example, if the multimedia to be clustered is music, the multimedia to be clustered contained in the connected block may include interview audio, which can be removed as noise. For details, please refer to the description of the above embodiment and will not be repeated here.
[0072] In this embodiment, by detecting the number of multimedia to be aggregated contained in each connected block, connected blocks with a number of multimedia to be aggregated greater than a threshold can be designated as target connected blocks, while connected blocks with a number of multimedia to be aggregated less than the threshold can be designated as non-target connected blocks. Since the number of multimedia to be aggregated contained in a non-target connected block is relatively small, the multimedia to be aggregated contained in the non-target connected block can be directly classified as a category without further clustering, thereby improving the efficiency of the aggregation of the multimedia to be aggregated. Specifically, non-target connected blocks are obtained from multiple connected blocks, and the multimedia to be aggregated contained in the non-target connected blocks are designated as corresponding categories in the aggregation results. The number of multimedia to be aggregated in the non-target connected blocks is less than the threshold. The threshold in this embodiment can be set as needed, and the embodiment is not limited thereto.
[0073] In the case of dividing each functional module according to each function, an embodiment of the present disclosure provides a multimedia aggregation device, which can be a server, a terminal, or a chip applied to a server. Figure 2 is a schematic block diagram of the functional modules of the multimedia aggregation device provided by an exemplary embodiment of the present disclosure. As shown in Figure 2, the multimedia aggregation device includes:
[0074] The multimedia set to be aggregated obtaining module 10 is used to obtain the multimedia set to be aggregated, wherein the multimedia set to be aggregated includes a plurality of multimedia to be aggregated;
[0075] a correlation obtaining module 20, configured to obtain correlations between the plurality of multimedia to be aggregated, and to divide the multimedia set to be aggregated into a plurality of multimedia subsets to be aggregated based on the correlations;
[0076] The clustering processing module 30 is configured to perform clustering processing on the multimedia to be aggregated in each of the multimedia subsets to be aggregated, and obtain an aggregation result of the plurality of multimedia to be aggregated.
[0077] In another embodiment provided by the present disclosure, the correlation acquisition module is specifically configured to:
[0078] Based on the correlation, constructing a target sparse adjacency matrix; the target sparse adjacency matrix represents the correlation between the multimedia to be aggregated;
[0079] The similarities between the multimedia to be aggregated in the target sparse adjacency matrix are obtained, and target correlation relationships between the multimedia to be aggregated are re-determined based on the similarities, and the multimedia set to be aggregated is divided into multiple multimedia subsets to be aggregated based on the target correlation relationships.
[0080] In another embodiment provided by the present disclosure, the correlation acquisition module is further configured to:
[0081] Obtaining a connectivity graph corresponding to the target sparse adjacency matrix, the connectivity graph comprising a plurality of vertices and edges between vertices, the vertices representing the multimedia to be aggregated, and the edges representing correlations between the multimedia to be aggregated;
[0082] Based on the similarity, edges in the connectivity graph are removed or retained to obtain a target connectivity graph, and target correlation relationships between the multimedia to be aggregated are determined based on the target connectivity graph.
[0083] In another embodiment provided by the present disclosure, the correlation acquisition module is further configured to:
[0084] Obtaining the connection relationship between each vertex in the target connectivity graph;
[0085] The target connectivity graph is divided into a plurality of connectivity blocks based on the connectivity relationships, and the multimedia set to be aggregated is divided into a plurality of multimedia subsets to be aggregated based on the plurality of connectivity blocks.
[0086] In another embodiment provided by the present disclosure, the cluster processing module is specifically configured to:
[0087] Obtaining a target connected block from the plurality of connected blocks, wherein the amount of multimedia to be aggregated in the target connected block is greater than a threshold;
[0088] De-noising is performed on the multimedia to be aggregated in the target connected block, and clustering is performed on the multimedia to be aggregated in the target connected block after de-noising.
[0089] In another embodiment provided by the present disclosure, the apparatus further includes:
[0090] a non-target connected block acquisition module, configured to obtain a non-target connected block from the plurality of connected blocks, wherein the amount of multimedia to be aggregated in the non-target connected block is not greater than a threshold;
[0091] The category determination module is configured to use the multimedia to be aggregated contained in the non-target connected block as a corresponding category in the aggregation result.
[0092] In another embodiment provided by the present disclosure, the apparatus further includes:
[0093] A feature information extraction module, configured to extract feature information of each multimedia to be aggregated;
[0094] The retrieval module is configured to search the multimedia to be aggregated in the multimedia set to be aggregated based on the feature information, and determine the correlation between the plurality of multimedia to be aggregated based on the obtained retrieval results.
[0095] For the relevant device part, which corresponds to the method, please refer to the description of the above method for details and will not be repeated here.
[0096] The multimedia aggregation device provided by the embodiments of the present disclosure obtains a multimedia set to be aggregated and determines the correlations between the multimedia to be aggregated in the multimedia set to be aggregated, thereby dividing the multimedia set to be aggregated into multiple multimedia subsets to be aggregated. Clustering the multimedia to be aggregated in the multimedia subsets to be aggregated can yield aggregation results for the multimedia to be aggregated in the multimedia set to be aggregated. This allows for preliminary aggregation of the multimedia to be aggregated based on the correlations between the multimedia to be aggregated, and further clustering of the resulting multimedia subsets to be aggregated can achieve accurate aggregation of the multiple multimedia to be aggregated in the set to be aggregated.
[0097] An embodiment of the present disclosure further provides an electronic device, comprising: at least one processor; a memory for storing instructions executable by the at least one processor; wherein the at least one processor is configured to execute the instructions to implement the above method disclosed in the embodiment of the present disclosure.
[0098] Figure 3 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present disclosure. As shown in Figure 3, the electronic device 1800 includes at least one processor 1801 and a memory 1802 coupled to the processor 1801. The processor 1801 can execute the corresponding steps of the above method disclosed in the embodiment of the present disclosure.
[0099] The processor 1801 can also be referred to as a central processing unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in the embodiments of the present disclosure can be completed by hardware integrated logic circuits in the processor 1801 or by software instructions. The processor 1801 can be a general-purpose processor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in the memory 1802, such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media mature in the art. The processor 1801 reads the information in the memory 1802 and, in conjunction with its hardware, completes the steps of the method.
[0100] In addition, when various operations / processes according to the present disclosure are implemented via software and / or firmware, the programs constituting the software can be installed from a storage medium or a network to a computer system having a dedicated hardware structure, such as computer system 1900 shown in FIG4 . When the various programs are installed, the computer system can perform various functions, including those described above. FIG4 is a block diagram of the structure of a computer system provided by an exemplary embodiment of the present disclosure.
[0101] Computer system 1900 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0102] As shown in FIG4 , computer system 1900 includes a computing unit 1901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage unit 1908 into a random access memory (RAM) 1903. Various programs and data required for the operation of computer system 1900 may also be stored in RAM 1903. Computing unit 1901, ROM 1902, and RAM 1903 are connected to each other via a bus 1904. An input / output (I / O) interface 1905 is also connected to bus 1904.
[0103] Several components within computer system 1900 are connected to I / O interface 1905, including an input unit 1906, an output unit 1907, a storage unit 1908, and a communication unit 1909. Input unit 1906 can be any type of device capable of inputting information into computer system 1900. Input unit 1906 can receive input numeric or character information and generate key input signals related to user settings and / or function control of an electronic device. Output unit 1907 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 1908 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 1909 allows computer system 1900 to exchange information / data with other devices over a network, such as the Internet, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0104] The computing unit 1901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1901 performs the various methods and processes described above. For example, in some embodiments, the above-mentioned method disclosed in the embodiments of the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1908. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via ROM 1902 and / or communication unit 1909. In some embodiments, the computing unit 1901 may be configured to perform the above-mentioned method disclosed in the embodiments of the present disclosure by any other appropriate means (e.g., by means of firmware).
[0105] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above method disclosed in the embodiment of the present disclosure.
[0106] The computer-readable storage medium in the embodiments of the present disclosure can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. The above-mentioned computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the above. More specifically, the above-mentioned computer-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0107] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0108] The embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor, the method disclosed in the embodiments of the present disclosure is implemented.
[0109] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer.
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0111] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the modules, components, or units do not necessarily limit the modules, components, or units themselves.
[0112] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0113] The above descriptions are merely some embodiments of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.
[0114] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A multimedia aggregation method, comprising: Acquire a multimedia set to be aggregated, where the multimedia set to be aggregated includes a plurality of multimedia to be aggregated; Obtaining correlations between the plurality of multimedia to be aggregated, and dividing the multimedia set to be aggregated into a plurality of multimedia subsets to be aggregated based on the correlations; as well as Clustering processing is performed on the multimedia to be aggregated in each of the multimedia subsets to be aggregated to obtain an aggregation result of the multiple multimedia to be aggregated.
2. The method according to claim 1, wherein dividing the multimedia set to be aggregated into a plurality of multimedia subsets to be aggregated based on the correlation comprises: Based on the correlation, construct a target sparse adjacency matrix; The target sparse adjacency matrix represents the correlation between the multimedia to be aggregated; as well as The similarities between the multimedia to be aggregated in the target sparse adjacency matrix are obtained, and target correlation relationships between the multimedia to be aggregated are re-determined based on the similarities, and the multimedia set to be aggregated is divided into multiple multimedia subsets to be aggregated based on the target correlation relationships.
3. The method according to claim 2, wherein the re-determining the target correlation relationship between the multimedia to be aggregated based on the similarity comprises: Obtaining a connectivity graph corresponding to the target sparse adjacency matrix, the connectivity graph comprising a plurality of vertices and edges between vertices, the vertices representing the multimedia to be aggregated, and the edges representing correlations between the multimedia to be aggregated; as well as Based on the similarity, edges in the connectivity graph are removed or retained to obtain a target connectivity graph, and target correlation relationships between the multimedia to be aggregated are determined based on the target connectivity graph.
4. The method according to claim 3, wherein dividing the multimedia set to be aggregated into a plurality of multimedia subsets to be aggregated based on the target correlation relationship comprises: Obtaining the connection relationship between each vertex in the target connectivity graph; as well as The target connectivity graph is divided into a plurality of connectivity blocks based on the connectivity relationships, and the multimedia set to be aggregated is divided into a plurality of multimedia subsets to be aggregated based on the plurality of connectivity blocks.
5. The method according to claim 4, wherein the clustering of the multimedia to be aggregated in each of the multimedia subsets to be aggregated comprises: Obtaining a target connected block from the plurality of connected blocks, wherein the amount of multimedia to be aggregated in the target connected block is greater than a threshold; as well as De-noising is performed on the multimedia to be aggregated in the target connected block, and clustering is performed on the multimedia to be aggregated in the target connected block after de-noising.
6. The method according to claim 5, further comprising: Obtaining a non-target connected block from the plurality of connected blocks, wherein the amount of multimedia to be aggregated in the non-target connected block is not greater than a threshold; as well as The multimedia to be aggregated contained in the non-target connected block is used as a corresponding category in the aggregation result.
7. The method according to claim 1, further comprising: Extracting characteristic information of each multimedia to be aggregated; as well as Based on the feature information, each multimedia to be aggregated is searched in the multimedia set to be aggregated, and the correlation between the multiple multimedia to be aggregated is determined based on the obtained search results.
8. A multimedia aggregation device, comprising: A multimedia set to be aggregated obtaining module, configured to obtain a multimedia set to be aggregated, wherein the multimedia set to be aggregated includes a plurality of multimedia to be aggregated; a correlation obtaining module, configured to obtain correlations between the plurality of multimedia to be aggregated, and divide the multimedia set to be aggregated into a plurality of multimedia subsets to be aggregated based on the correlations; as well as The clustering processing module is used to perform clustering processing on the multimedia to be aggregated in each of the multimedia subsets to be aggregated, and obtain an aggregation result of the multiple multimedia to be aggregated.
9. An electronic device comprising: at least one processor; as well as a memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 7.