Audio distribution information acquisition method, computer device and computer program product
By constructing the target audio playback sequence and extracting audio feature vectors, the problem of insufficient accuracy of audio preference analysis in the prior art is solved, and more accurate user audio preference analysis and better audio content push are achieved.
Patent Information
- Application Number
- CN202210726287.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-24
AI Technical Summary
In the prior art, the accuracy of user audio preference analysis is insufficient, resulting in the pushed audio content not meeting user preferences.
By obtaining the played audio of users in multiple playback platforms, a target audio playback sequence is constructed, and audio feature vectors are extracted for each playedback audio. Based on the distance relationship of these feature vectors in the vector space, the distribution of played audio of multiple playback platforms is determined.
It improves the accuracy of obtaining audio distribution information, can more accurately reflect the user's audio preferences, thereby improving the matching degree of pushed audio content.
Smart Images

Figure CN115146104B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and in particular to a method, apparatus, computer equipment, storage medium and computer program product for obtaining audio distribution information. Background Art
[0002] With the development of Internet technology, users can now play and listen to audio through mobile phones and other terminals. In order to improve the user's listening experience, it is necessary to analyze the user's audio preference content so that audio content that meets their preferences can be pushed to the user. The current method of analyzing the user's audio preference is usually based on the number of times the user's audio is played. However, user preference analysis based on the number of audio plays will result in inaccurate user preference content obtained through analysis.
[0003] Therefore, the current method for obtaining audio distribution information of users has the defect of insufficient accuracy. Summary of the invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium and computer program product for obtaining audio distribution information that can improve accuracy in response to the above technical problems.
[0005] In a first aspect, the present application provides a method for obtaining audio distribution information, the method comprising:
[0006] Obtain a target audio playback sequence according to the played audios of the user on multiple playback platforms;
[0007] Acquire an audio feature vector of each of the played audio in the target audio playback sequence to obtain a plurality of audio feature vectors corresponding to the target audio playback sequence; wherein the audio feature vector has a plurality of vector dimensions in a vector space, and each vector dimension represents an audio feature information of the played audio;
[0008] The distribution of the played audios of the multiple playback platforms is determined according to the distances of the multiple audio feature vectors of the target audio playback sequence in at least one vector dimension.
[0009] In one embodiment, obtaining a target audio playback sequence according to the played audio of the user on multiple playback platforms includes:
[0010] For each playback platform, obtaining multiple initial audio playback sequences obtained from the played audios corresponding to multiple users in the playback platform;
[0011] According to the set ratio, a corresponding number of initial audio playback sequences are selected from multiple initial audio playback sequences corresponding to each playback platform;
[0012] The initial audio playback sequences of the selected multiple playback platforms are combined to obtain the target audio playback sequence.
[0013] In one embodiment, for each playback platform, obtaining multiple initial audio playback sequences obtained from the played audios corresponding to multiple users in the playback platform includes:
[0014] From multiple playback platforms, multiple played audios of multiple users of each playback platform in the same time period are obtained to generate multiple initial audio playback sequences corresponding to each of the playback platforms.
[0015] In one embodiment, the step of obtaining the audio feature vector of each played audio in the target audio playback sequence includes:
[0016] For each initial audio playback sequence, combining initial audio feature vectors of each played audio in the initial audio playback sequence into an audio feature vector set;
[0017] For each of the played audio, the initial audio feature vectors of the played audio in different audio feature vector sets are merged according to the number of times the played audio is played in different initial audio playback sequences to obtain the audio feature vector of the played audio in the target audio playback sequence.
[0018] In one of the embodiments, for each of the played audios, according to the number of times the played audios are played in different initial audio playback sequences, initial audio feature vectors of the played audios in different audio feature vector sets are merged to obtain an audio feature vector of the played audio in the target audio playback sequence, including:
[0019] Obtaining initial audio feature vectors corresponding to the same played audio in each two sets of audio feature vectors, and obtaining an average value of the number of times the same played audio is played in the initial audio play sequences corresponding to the each two sets of audio feature vectors;
[0020] Determine an average audio feature vector of the initial audio feature vectors corresponding to the same played audio in each two sets of audio feature vectors according to the average value, and replace the initial audio feature vectors corresponding to the same played audio in each two sets of audio feature vectors with the average audio feature vector to obtain a new audio feature vector set;
[0021] Detecting whether the number of the new audio feature vector sets is one, and if not, returning to the step of obtaining initial audio feature vectors corresponding to each of two sets of audio feature vector sets for the same played audio;
[0022] If so, the audio feature vector in the new audio feature vector set is used as the audio feature vector of the played audio in the target audio playback sequence.
[0023] In one of the embodiments, determining the distribution of played audio on the multiple playback platforms according to the distances of the multiple audio feature vectors of the target audio playback sequence in at least one vector dimension includes:
[0024] Determining at least one target vector dimension among the multiple vector dimensions, wherein the vector discreteness of the multiple audio feature vectors in the dimensional space corresponding to the at least one target vector dimension is the largest;
[0025] According to the at least one target vector dimension, reducing the dimensions of the multiple audio feature vectors to obtain audio feature vectors after dimension reduction;
[0026] The multiple audio feature vectors after dimensionality reduction are distributed in a vector space corresponding to the at least one target vector dimension to obtain distribution of the played audio on the multiple playback platforms.
[0027] In one of the embodiments, determining the distribution of played audio on the multiple playback platforms according to the distances of the multiple audio feature vectors of the target audio playback sequence in at least one vector dimension includes:
[0028] Determining similarities between the multiple audio feature vectors according to the Euclidean distances between the multiple audio feature vectors in the vector space;
[0029] According to the similarities between the multiple audio feature vectors, distribution results of the played audios of the multiple playback platforms are determined.
[0030] In one embodiment, the step of obtaining the audio feature vector of each played audio in the target audio playback sequence includes:
[0031] Obtaining multiple audio feature information of each of the played audio in the target audio playback sequence;
[0032] Vectorized processing is performed on multiple audio feature information of the played audio through a preset natural language processing model to obtain an audio feature vector of the played audio.
[0033] In one embodiment, the target audio playback sequence includes target audio playback sequences corresponding to multiple different time periods; after determining the distribution of the played audios on the multiple playback platforms, the method further includes:
[0034] Based on the orthogonal Prucker method, the distribution of the played audio in the target audio playback sequence corresponding to the multiple different time periods is mapped to the same vector space to obtain the distribution position of each played audio;
[0035] In the same vector space, change information of the distribution position of the same played audio in the multiple different time periods is determined.
[0036] In a second aspect, the present application provides a method for obtaining audio distribution information, the method comprising:
[0037] In response to a multi-platform audio distribution information viewing instruction, platform information and viewing dimension information of multiple playback platforms are obtained, and distribution of multiple played audios of the multiple playback platforms under the viewing dimension information is obtained; the distribution of the played audios of the multiple playback platforms is determined based on the above method;
[0038] In the dimensional space corresponding to the viewing dimensional information, the distribution of multiple played audios of the multiple playback platforms is displayed.
[0039] In a third aspect, the present application provides a device for acquiring audio distribution information, the device comprising:
[0040] A sequence acquisition module, used to obtain a target audio playback sequence based on the user's played audio on multiple playback platforms;
[0041] A feature acquisition module, used to acquire an audio feature vector of each of the played audio in the target audio playback sequence, so as to obtain a plurality of audio feature vectors corresponding to the target audio playback sequence; wherein the audio feature vector has a plurality of vector dimensions in a vector space, and each vector dimension represents an audio feature information of the played audio;
[0042] The distribution determination module is used to determine the distribution of the played audio on the multiple playback platforms according to the distance between the multiple audio feature vectors of the target audio playback sequence in at least one vector dimension.
[0043] In a fourth aspect, the present application provides a device for acquiring audio distribution information, the device comprising:
[0044] A response module, configured to respond to a multi-platform audio distribution information viewing instruction, obtain platform information of multiple playback platforms and viewing dimension information, and obtain distribution of multiple played audios of the multiple playback platforms under the viewing dimension information; the distribution of the played audios of the multiple playback platforms is determined based on the above method;
[0045] A display module is used to display the distribution of multiple played audios of the multiple playback platforms in the dimensional space corresponding to the viewing dimensional information.
[0046] In a fifth aspect, the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0047] In a sixth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0048] In a seventh aspect, the present application provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0049] The above-mentioned audio distribution information acquisition method, apparatus, computer equipment, storage medium and computer program product obtain a target audio playback sequence containing the played audio corresponding to multiple users on multiple playback platforms based on the audio playback behavior of users on multiple playback platforms, and obtain multiple audio feature information corresponding to each played audio in the target audio playback sequence. For each played audio in the audio playback sequence, an audio feature vector with multiple dimensions corresponding to the played audio is obtained, and multiple audio feature vectors corresponding to the audio playback sequence are obtained. Based on the distance of the multiple audio feature vectors in at least one dimension in the vector space, the distribution results of the multiple played audios are determined. Compared with the traditional method of determining the user's audio distribution information only based on the number of times the user's audio is played, this scheme determines the audio playback sequence through the user's playback behavior on multiple playback platforms, and determines the audio feature vectors of multiple dimensions based on the sequence and the audio feature information, thereby determining the distribution result of the audio information based on the distance between the multiple vectors observed from the perspective of at least one dimension in the vector space, thereby improving the accuracy of obtaining the audio playback information. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A diagram showing an application environment of a method for obtaining audio distribution information in an embodiment;
[0051] Figure 2 A schematic diagram of a flow chart of a method for obtaining audio distribution information in one embodiment;
[0052] Figure 3 is a schematic flow chart of a vector merging step in one embodiment;
[0053] Figure 4 A schematic diagram of an interface for displaying audio distribution information in an embodiment;
[0054] Figure 5 It is a schematic diagram of an interface for displaying audio distribution information in another embodiment;
[0055] Figure 6 A schematic diagram of an interface for a step of changing audio distribution information in an embodiment;
[0056] Figure 7 A schematic diagram of an interface of a step of changing audio distribution information in another embodiment;
[0057] Figure 8 A schematic diagram of a flow chart of a method for obtaining audio distribution information in another embodiment;
[0058] Fig. 9 is a structural block diagram of an audio distribution information acquisition device in one embodiment;
[0059] Fig.10 is a structural block diagram of an audio distribution information acquisition device in another embodiment;
[0060] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0062] The audio distribution information acquisition method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can obtain the played audio corresponding to multiple users on multiple playback platforms from the server 104, and obtain the audio playback sequence based on the information, so that the terminal 102 can determine the distribution results of multiple played audios based on the multiple audio feature vectors in the audio playback sequence. Among them, the terminal 102 can be but not limited to various personal computers, laptops, smart phones and tablet computers. The server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.
[0063] In one embodiment, Figure 2 As shown, a method for obtaining audio distribution information is provided, and the method is applied to Figure 1 The terminal in is used as an example to illustrate, including the following steps:
[0064] Step S202, obtaining a target audio playback sequence according to the played audios of the users on multiple playback platforms.
[0065] Among them, the playback platform can be a platform used by the user to play audio, such as major music playback software, etc., and the target audio playback sequence includes the played audio corresponding to multiple users on multiple playback platforms. The user may use more than one audio playback software, that is, a user can perform audio playback on multiple playback platforms. Terminal 102 can determine the audio playback sequence based on the played audio of multiple users on multiple playback platforms with multiple users as the object. Among them, the audio playback sequence includes the played audio corresponding to multiple users on multiple playback platforms, and the played audio in the audio playback sequence can be sorted according to the number of plays, or can be sorted according to the audio playback time. Among them, the audio playback sequence obtained by terminal 102 can be a sequence obtained by merging the audio playback sequences corresponding to multiple playback platforms, so it can include the played audio corresponding to multiple playback platforms. In addition, the audio playback sequence corresponding to the above-mentioned multiple playback platforms can be an audio playback sequence of a non-specific user, that is, the audio playback sequence obtained by terminal 102 from multiple playback platforms is not for a specific user, and the user is anonymous in this scheme. Terminal 102 can treat the audio playback sequences of the above-mentioned multiple playback platforms as individual user behaviors, that is, one sequence corresponds to one user. Moreover, multiple playback platforms may be playback platforms of different sites. A site may represent the manager of a playback platform, and a manager may have multiple playback platforms. When a manager wants to know the distribution of audio playback behaviors of users of multiple playback platforms, since a manager can usually only obtain the user audio playback data of the playback platform managed by itself, the audio playback data of other managers' playback platforms can be obtained by obtaining their public data.
[0066] Step S204, obtaining the audio feature vector of each played audio in the target audio playback sequence to obtain multiple audio feature vectors corresponding to the target audio playback sequence; wherein the audio feature vector has multiple vector dimensions in the vector space, and each vector dimension represents an audio feature information of the played audio.
[0067] The target audio play sequence may be obtained by the terminal 102 after merging multiple audio play sequences obtained from different play platforms. The target audio play sequence is a play sequence that integrates the audio play conditions of users on multiple play platforms. The played audio may be the audio that the user has played, and the audio feature information may be various audio feature information of each played audio in the target audio play sequence. Taking the audio as a song as an example, the audio feature information may include but is not limited to the name of the song, the composition information of the song, the style of the song, the popularity index of the song, the literary index of the song user, etc.
[0068] Among them, there may be multiple played audios in the target audio playback sequence, and these played audios may come from different playback platforms, and each played audio may represent an audio. For each played audio in the target audio playback sequence, the terminal 102 may obtain an audio feature vector corresponding to each played audio, and the audio feature vector may be a vector in a high-dimensional vector space, then each audio feature vector may have corresponding multiple dimensions in the vector space, and these dimensions may be determined according to the above-mentioned audio feature information. The terminal 102 may vectorize each played audio in the audio playback sequence, thereby obtaining multiple audio feature vectors corresponding to the audio playback sequence.
[0069] Step S206: determining the distribution of played audio on multiple playback platforms according to the distances of multiple audio feature vectors of the target audio playback sequence in at least one vector dimension.
[0070] Among them, multiple audio feature vectors can be multiple vectors obtained after vectorization of each played audio in the above-mentioned target audio playback sequence. Terminal 102 can determine the distribution results of multiple played audio based on the distance of multiple vectors in at least one dimension in the vector space. For example, the above-mentioned audio feature vectors can be vectors in a high-dimensional space. When observing the distribution of each audio feature vector in the high-dimensional vector space from the angles corresponding to different dimensions, the distribution of these audio feature vectors will be different. For example, when observing the distribution of audio feature vectors from the angles corresponding to certain dimensions, some of the vectors will have overlapping positions, so that the specific distance relationship of these audio feature vectors cannot be known. Therefore, terminal 102 needs to determine which dimensions can maximize the distance relationship between multiple audio feature vectors based on the distance between the multiple audio feature vectors obtained at the observation angle corresponding to each dimension in the vector space, so as to determine the distribution results of multiple played audio.
[0071] In the above-mentioned method for obtaining audio distribution information, an audio playback sequence including the played audio corresponding to multiple users on multiple playback platforms is obtained based on the audio playback behavior of users on multiple playback platforms, and multiple audio feature information corresponding to each played audio in the target audio playback sequence is obtained. For each played audio in the target audio playback sequence, an audio feature vector with multiple dimensions corresponding to the played audio is obtained, and multiple audio feature vectors corresponding to the target audio playback sequence are obtained. Based on the distance of the multiple audio feature vectors in at least one dimension in the vector space, the distribution results of the multiple played audios are determined. Compared with the traditional method of determining the audio distribution information of a user only based on the number of times the user plays the audio, this scheme determines the target audio playback sequence through the user's playback behavior on multiple playback platforms, and determines the audio feature vectors of multiple dimensions based on the sequence and the audio feature information, thereby determining the distribution result of the audio information based on the distance between the multiple vectors observed from the perspective of at least one dimension in the vector space, thereby improving the accuracy of obtaining audio playback information.
[0072] In one embodiment, a target audio playback sequence is obtained based on the played audio of users on multiple playback platforms, including, for each playback platform, obtaining multiple initial audio playback sequences obtained from the played audio corresponding to multiple users on the playback platform; selecting a corresponding number of initial audio playback sequences from the multiple initial audio playback sequences corresponding to each playback platform according to a set ratio; and merging the selected initial audio playback sequences of the multiple playback platforms to obtain a target audio playback sequence.
[0073] In this embodiment, the terminal 102 can determine the target audio playback sequence from the played audio of users of multiple playback platforms. Among them, since multiple playback platforms can correspond to different management parties, for example, one of the multiple playback platforms belongs to the management party to which the terminal 102 belongs, and another playback platform belongs to other management parties other than the management party to which the terminal 102 belongs, then for the playback platform of the management party to which the terminal 102 belongs, the terminal 102 can obtain the played audio of all users of each playback platform under its management party to form the initial audio playback sequence corresponding to each playback platform; for the playback platform other than the management party to which the terminal 102 belongs, the terminal 102 can obtain the relevant data of the publicly played audio, which can be a non-full-user data, thereby obtaining the initial audio playback sequence of the playback platform of other management parties. The terminal 102 can merge the audio playback sequences obtained from each playback platform to obtain the merged target audio playback sequence.
[0074] The terminal 102 can obtain multiple initial audio playback sequences obtained from the played audio corresponding to multiple users in multiple playback platforms through the above-mentioned data acquisition method for different playback platforms, and select a corresponding number of initial audio playback sequences from the multiple initial audio playback sequences corresponding to each playback platform according to the set ratio corresponding to the multiple playback platforms. Thus, the terminal 102 can obtain the selected multiple corresponding numbers of initial audio playback sequences corresponding to each playback platform, and the terminal 102 can merge the selected multiple corresponding numbers of initial audio playback sequences corresponding to the playback platforms to obtain the fused target audio playback sequence.
[0075] Among them, the set ratios corresponding to the above-mentioned multiple playback platforms can be determined according to the number of users corresponding to each playback platform. For example, in one embodiment, according to the set ratios corresponding to the multiple playback platforms, before selecting a corresponding number of initial audio playback sequences from the multiple initial audio playback sequences corresponding to each playback platform, it also includes: obtaining the number of users corresponding to each playback platform; determining the user ratio between the multiple playback platforms based on the number of users; and determining the set ratios corresponding to the multiple playback platforms based on the user ratio. In this embodiment, before the terminal 102 merges the audio playback sequences of multiple playback platforms, it is necessary to first determine the set ratios between the multiple playback platforms, so that the terminal 102 can merge the initial audio playback sequences of the multiple playback platforms according to the set ratio.
[0076] Terminal 102 can obtain the number of users corresponding to each playback platform, so that terminal 102 can determine the user ratio between multiple playback platforms based on the number of users, and determine the set ratio corresponding to multiple playback platforms based on the user ratio. Specifically, assuming that the user's listening preference behavior will not have a huge jump in each platform, for example, a user is unlikely to listen to traditional opera songs on playback platform A, but listen to European and American electronic dance music on platform B, then terminal 102 can obtain the user's preferred audio playback sequence on two different playback platforms, and determine the set ratio when performing sequence fusion based on the number of users on different playback platforms, as two sub-corpora, merged into an overall corpus. Specifically, since audio information has playback copyright, not all playback platforms have the playback rights of all audios. For example, some audio can be played on one playback platform, but other audio needs to be played on another playback platform. Therefore, terminal 102 needs to integrate the audio on multiple playback platforms to achieve a comprehensive user played audio.
[0077] The terminal 102 can adjust the weights x and y between the two sub-corpora according to the number of audio playback sequences actually obtained from multiple playback platforms, and approximate the proportion of the number of listening users on different platforms in the actual market, so as to realize vectorization in the case of cross-platform content. Based on the above assumptions, the current method can also be applied to public user preference sequence data. That is, the terminal 102 can collect and merge the initial audio sequences of multiple playback platforms under one management party, and can also collect and merge the initial audio sequences of multiple playback platforms under different management parties. For example, let the first playback platform and the second playback platform be different playback platforms belonging to the same management party. If the terminal 102 extracts the initial audio playback sequence of the user in the first playback platform at a ratio of x as sub-corpora one; then extracts the initial audio playback sequence of the user in the second playback platform at a ratio of y as sub-corpora two in the same way. Where the ratio x and the ratio y can make the overall ratio of the two sub-corpora close to the number of users of the playback platform in the actual market, then the terminal 102 can merge sub-corpora one and sub-corpora two into an audio playback sequence preferred by the overall users of the management party.
[0078] After the terminal 102 determines the set ratio between the above-mentioned various playback platforms, it can select a corresponding number of initial audio playback sequences from the initial audio sequences of each playback platform based on the ratio, and merge the initial audio playback sequences of the selected multiple playback platforms to obtain a merged audio playback sequence. For example, for multiple playback platforms in the same management party, the terminal 102 can obtain the completed audio sequences of each user in the same time period, such as within a week, and perform sampling fusion according to actual needs, for example, sampling fusion is performed according to the set ratio between the above-mentioned various playback platforms, so that the terminal 102 can construct a high-dimensional vector of content within the same management party, which is used for analyzing the content distribution within the station in a specified time window. Specifically, the terminal 102 can use the following formula to obtain the fusion sequence of the initial audio playback sequences between various playback platforms under the same management party: 1*the audio playback sequence of the top 100 preferred users of the first playback platform in one week + 0.8*the audio playback sequence of the top 100 preferred users of the second playback platform in one week. Among them, the first playback platform and the second playback platform can both belong to the same management party. 1 and 0.8 can respectively represent the ratio of the number of users between the two playback platforms.
[0079] For playback platforms of different managers, terminal 102 can collect non-full public data of other playback platforms that do not belong to its manager based on a specific time period, and can also use the audio playback sequence of the playback platform in the manager of terminal 102 in the same time period for multiple random sampling, and merge it with the public audio playback sequence of the playback platform of the above-mentioned other managers to construct a high-dimensional content vector with ductility, so as to realize the analysis of the content distribution inside and outside the station that integrates multiple managers. For example, terminal 102 can use the following formula to obtain the fusion sequence of the initial audio sequences between various playback platforms under different managers: (0.2*the first playback platform's full user Top100 preference sequence in one week + the crawled third playback platform's non-full user Top100 preference sequence) + (0.1*the second playback platform's full user Top100 preference sequence in one week + the crawled third playback platform's non-full user Top100 preference sequence). Among them, the third playback platform can be other playback platforms that do not belong to the management party of the above-mentioned terminal 102. Terminal 102 can only obtain its public audio playback data from these playback platforms, such as the user's weekly listening ranking and other information; 0.2 and 0.1 can respectively represent the ratio of the number of users between different playback platforms. After terminal 102 obtains the above-mentioned fused audio playback sequence, it can use the audio playback sequence as the data source of Word2Vec, that is, terminal 102 can input the above-mentioned audio playback sequence into Word2Vec and vectorize the data therein. Among them, the Word2Vec algorithm is one of the most common tools in natural language processing scenarios. It converts words into vectors in semantic space through batch learning of data and predicting the most likely content in the windows before and after words.
[0080] Through this embodiment, the terminal 102 can determine the set fusion ratio based on the number of users between different playback platforms, and based on the ratio, fuse the initial audio playback sequences of multiple playback platforms belonging to the same manager or different managers to obtain the user's vectorized audio playback sequence, so that the terminal 102 can analyze the audio preference distribution based on the vectorized played audio, thereby improving the accuracy of obtaining audio distribution information.
[0081] In one embodiment, obtaining an audio feature vector of each played audio in a target audio playback sequence includes: obtaining multiple audio feature information of each played audio in the target audio playback sequence; vectorizing multiple audio feature information of the played audio by setting a natural language processing model to obtain an audio feature vector of the played audio.
[0082] In this embodiment, the target audio playback sequence may be an audio playback sequence obtained by the terminal 102 after fusing the initial audio playback sequences collected from multiple playback platforms. Based on the target audio playback sequence, the terminal 102 may vectorize the multiple audio feature information of each played audio therein to obtain the audio feature vector corresponding to each played audio. Wherein, the above-mentioned vectorization of the played audio may be performed by setting a natural language processing model. The terminal 102 may vectorize the multiple audio feature information of each played audio contained in the target audio playback sequence by using the set natural language processing model and the multiple audio feature information corresponding to the played audio, and obtain the audio feature vector of each played audio in the target audio playback sequence. Wherein, the above-mentioned audio feature vector may be a multi-dimensional vector, and the multi-dimensional audio feature vector may be a vector in a high-dimensional space, and each dimension in the high-dimensional space may be obtained by processing the multiple audio feature information of the played audio by setting a natural language processing model, that is, the above-mentioned vectorized audio feature vector may be displayed in a high-dimensional vector space.
[0083] Among them, the above-mentioned natural language processing model can be Word2Vec. The Word2Vec algorithm is one of the most common tools in natural language processing scenarios. It converts words into vectors in semantic space by batch learning and predicting the most likely content in the windows before and after words through data. Terminal 102 can use each played audio in the audio playback sequence obtained after the above fusion as the input corpus of the natural language processing model Word2Vec, and output each played audio in the above audio playback sequence to Word2Vec, so as to obtain audio feature vectors of multiple dimensions corresponding to each played audio in the audio playback sequence through Word2Vec. Among them, the Word2Vec algorithm can generate audio feature vectors of multiple dimensions based on each played audio in the above-mentioned audio playback sequence, for example, generate an audio feature vector of fifty dimensions, that is, the terminal 102 can generate a vector space of multiple dimensions, for example, a vector space of fifty dimensions, then the terminal 102 can generate audio feature vectors of multiple dimensions corresponding to each played audio in the audio playback sequence through Word2Vec, and place the audio feature vectors of multiple dimensions in the corresponding vector space of multiple dimensions, for example, the terminal 102 generates a fifty-dimensional audio feature vector through Word2Vec, and places it at the corresponding position in the fifty-dimensional vector space; after the terminal 102 vectorizes each played audio in the above-mentioned audio playback sequence, the distribution of the audio feature vectors corresponding to each played audio in the high-dimensional vector space can be obtained, and the terminal 102 can determine the distance relationship between each audio feature vector based on the distribution of the audio feature vectors corresponding to each played audio in the high-dimensional space, thereby analyzing the audio distribution information.
[0084] Through this embodiment, the terminal 102 can vectorize each played audio in the audio playback sequence through the set natural language processing model, and obtain the audio feature vector corresponding to each played audio in the high-dimensional space, so that the terminal 102 can analyze the audio distribution information based on the audio feature vector in the high-dimensional space, thereby improving the accuracy of obtaining the audio distribution information.
[0085] In one embodiment, for each playback platform, multiple initial audio playback sequences are obtained from the played audios corresponding to multiple users in the playback platform, including: from multiple playback platforms, multiple played audios of multiple users on each playback platform in the same time period are obtained to generate multiple initial audio playback sequences corresponding to each of the playback platforms.
[0086] In this embodiment, for other playback platforms that do not belong to the management party to which the terminal 102 belongs, the terminal 102 can obtain the audio playback data disclosed by these other playback platforms, such as the user's song ranking information within a week, which can be a non-full amount of public data. In order to ensure that the audio playback sequence used for analysis has a relatively complete amount of information, the terminal 102 can collect multiple played audios for multiple playback platforms under the management party to which the terminal 102 belongs in the same time window, and collect multiple non-full amount of public data for playback platforms of other management parties other than the management party to which the terminal 102 belongs, and the terminal 102 can merge based on the collected data to obtain a fused audio playback sequence. The terminal 102 can obtain multiple initial audio playback sequences corresponding to each playback platform collected in the same time period. Among them, each playback platform can include at least one of the playback platform of the management party to which the terminal 102 belongs and the playback platform other than the management party to which the terminal 102 belongs.
[0087] After the terminal 102 obtains multiple initial audio playback sequences collected in the same time period, it can determine the audio feature vector corresponding to the audio playback sequence obtained by merging multiple initial audio playback sequences based on the audio feature vectors in each of the multiple initial audio playback sequences. For example, in one embodiment, the audio feature vector of each played audio in the target audio playback sequence is obtained, including: for each initial audio playback sequence, combining the initial audio feature vectors of each played audio in the initial audio playback sequence into an audio feature vector set; for each played audio, according to the number of times the played audio is played in different initial audio playback sequences, the initial audio feature vectors of the played audio in different audio feature vector sets are merged to obtain the audio feature vector of the played audio in the target audio playback sequence.
[0088] In this embodiment, for multiple initial audio playback sequences collected by the terminal 102 in the same time period, the terminal 102 can obtain multiple initial audio feature vectors corresponding to each initial audio playback sequence, and each initial audio playback sequence can correspond to a set of initial audio feature vector sets, so that the terminal 102 can obtain multiple sets of audio feature vector sets corresponding to the above multiple initial audio playback sequences. For each played audio in the above multiple initial audio playback sequences, the terminal 102 can merge the initial audio feature vectors in the above multiple sets of audio feature vector sets according to the number of times each played audio is played in different sequences, so as to obtain the audio feature vector corresponding to the played audio.
[0089] Among them, when merging the audio feature vectors in the above-mentioned audio feature vector sets, for the same audio feature vectors in each audio feature vector set, such as the played audio corresponding to the same audio, since when using the word2vec algorithm, there will be a random starting point, which points to the overall general direction, so the directions of the audio feature vectors in each audio feature vector set may be different, so the terminal 102 needs to convert each audio feature vector in each audio feature vector set to the same direction, so as to merge each audio feature vector in each audio feature vector set from the same angle. Specifically, for two vector spaces with the same vocabulary, the terminal 102 can use the method of solving the orthogonal Prucker problem to find the transformation matrix on one side, and then use the converted vector space to take the mean, so as to merge the audio feature vectors. The conversion formula is as follows: M ij =1 / 2(V i A ij +V j ), where A ij In order to make the vector V i and vector V j The transformation matrix in the same direction, M ij is the mean vector, AA T =I indicates that A is an orthogonal matrix. After the terminal 102 obtains the above conversion matrix, it can convert one side vector into a vector with the same direction as the other vector based on the conversion matrix. For example, the vector V i Convert to vector V j The vectors of the same measurement scale are obtained by taking the vector mean to obtain the mean vector M. ij Taking the audio content as a song as an example, the terminal 102 can obtain a spatial transformation matrix A through the above method: ijThe terminal 102 can use the matrix to scale and rotate one group of song content vectors into a vector with the same measurement scale (same vector direction) as the other group of song content vectors, so that the terminal 102 can obtain the influence of the two groups of corpora collected twice on this group of songs and improve the accuracy of the vector.
[0090] Among them, since the number of the above-mentioned audio feature vector sets can be multiple, the number of the above-mentioned merging can be multiple times, and the merging method of the multiple merging can be an iterative merging, that is, the terminal 102 can perform the next merging based on the new audio feature vector set obtained after the first merging, so that the final merging results in a unique audio feature vector set. For example, in one embodiment, for each played audio, the initial audio feature vectors of the played audio in different audio feature vector sets are merged according to the number of times the played audio is played in different initial audio playback sequences to obtain the audio feature vector of the played audio in the target audio playback sequence, including: obtaining the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets, and obtaining the average number of times the same played audio is played in the initial audio playback sequences corresponding to each two groups of audio feature vector sets; determining the average audio feature vector of the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets according to the average value, replacing the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets with the average audio feature vector, to obtain a new audio feature vector set; detecting whether the number of new audio feature vector sets is one, and if not, returning to the step of obtaining the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets; if so, using the audio feature vectors in the new audio feature vector set as the audio feature vector of the played audio in the target audio playback sequence.
[0091] In this embodiment, the terminal 102 can merge multiple groups of initial audio feature vector sets multiple times to obtain a fused group of audio feature vector sets as the audio feature vector set of the target audio playback sequence. The above initial audio feature vector sets can have multiple groups, and each group of initial audio feature vector sets can have multiple initial audio feature vectors. For each played audio in each of the above-mentioned initial audio playback sequences, the terminal 102 can obtain the initial audio feature vectors corresponding to the same played audio in each of the two sets of audio feature vectors, and obtain the average of the number of times the same played audio feature vector information is played in each of the two sets of audio feature vectors, so that the terminal 102 can determine the average audio feature vector corresponding to each initial audio feature vector corresponding to the same played audio in each of the two sets of audio feature vectors according to the average; wherein the terminal 102 can also determine the average of each corresponding initial audio feature vector based on other audio feature information of the played audio feature vector information, and the above-mentioned determination of the average of the initial audio feature vectors of the same played audio in different audio feature vector sets by calculating the average of the number of times played can be one of the calculation methods, that is, the terminal 102 can obtain the audio feature vectors corresponding to the same played audio in different audio feature vector sets, and obtain the position information of these audio feature vectors in the multidimensional vector space, and determine the average audio feature vector of these audio feature vectors. wherein the above-mentioned position information is determined based on each audio feature information corresponding to the audio feature vector and the dimension of the above-mentioned multidimensional vector space. For the played audio that exists in only one of the two audio feature vector sets, the corresponding average audio feature vector may be itself.
[0092] After the terminal 102 determines the average audio feature vector corresponding to each initial audio feature vector in the two sets of audio feature vectors, each initial audio feature vector corresponding to the same played audio in each of the two sets of audio feature vectors can be replaced with its corresponding average audio feature vector. After the terminal 102 replaces each initial audio feature vector in each of the two sets of audio feature vectors with its corresponding average audio feature vector, the audio feature vector set corresponding to the replaced initial audio feature vector no longer exists. The terminal 102 can combine multiple average audio feature vectors into a new audio feature vector set. After each merger to obtain a new audio feature vector set, the terminal 102 can detect the number of audio feature vector sets. If it is detected that the number of new audio feature vector sets is not one, the terminal 102 determines that the audio feature vector sets need to be merged again. The terminal 102 can return to the step of obtaining the initial audio feature vectors corresponding to the same played audio in each of the two sets of audio feature vectors, and perform the next merger based on the new audio feature vector set until the number of audio feature vector sets is one. When the terminal 102 detects that the combined new audio feature vector set is one, the terminal 102 may use the audio feature vector in the current new audio feature vector set as the audio feature vector of the played audio in the target audio playback sequence.
[0093] Specifically, the schematic diagram of merging the above multiple groups of audio feature vector sets can be as follows: Figure 3 As shown, Figure 3 The terminal 102 collects user audio playback sequences of various playback platforms of the management party to which it belongs, and the terminal 102 also collects public audio playback sequences in other playback platforms other than the management party to which it belongs. The terminal 102 can set multiple fusion experiments, for example Figure 3 The eight fusions in the above are realized, and V k ,For example Figure 3 V1-V8 in V k It can be an audio feature vector set, which can be an audio feature vector set corresponding to an audio playback sequence obtained by terminal 102 after fusing audio playback sequences of multiple playback platforms. Specifically, since the above audio playback sequence contains audio playback information of other playback platforms, terminal 102 can set a ratio that can be simulated in real life, sample the data of the playback platform of the management party to which terminal 102 belongs, and then fuse it with the audio playback information of other playback platforms. For example, terminal 102 can determine the above fusion ratio based on the number of users of different playback platforms, thereby fusing the audio playback of multiple playback platforms, and obtain V based on the fused audio playback sequence. k ,like Figure 3As shown, the terminal 102 can perform eight fusions to obtain eight audio feature vector sets corresponding to the fused audio playback sequences. Thus, the terminal 102 can obtain the average value based on the eight audio feature vector sets. Specifically, the terminal 102 can convert the eight audio feature vector sets into two groups by using the conversion matrix A. ij The initial audio feature vectors in one set of audio feature vector sets are converted into the initial audio feature vectors in the same direction as the other set, and the converted audio feature vector set is averaged with the other set of audio feature vector sets to obtain the averaged M k ,For example Figure 3 Terminal 102 can use M1-M7. k Replace the corresponding fused audio feature vector set, so that the merged audio feature vector set M can be obtained k The terminal 102 may perform the process of averaging and merging multiple times until finally merging an audio feature vector set corresponding to an audio playback sequence, for example Figure 3 M7 in , so that terminal 102 can obtain a set of audio feature vectors with more complete information.
[0094] Through the above embodiments, the terminal 102 can collect audio playback sequences of multiple playback platforms multiple times to obtain multiple initial audio playback sequences corresponding to the multiple playback platforms, and the terminal 102 can also average and merge the multiple initial audio playback sequences multiple times to obtain an audio feature vector set corresponding to a final audio playback sequence, so that the terminal 102 can analyze the user's audio distribution information based on the audio feature vector set to improve the accuracy of audio distribution information acquisition.
[0095] In one embodiment, the distribution of played audio on multiple playback platforms is determined based on the distance between multiple audio feature vectors of a target audio playback sequence in at least one vector dimension, including: determining multiple dimensions in a vector space based on multiple audio feature information corresponding to multiple audio feature vectors; determining at least one target vector dimension among the multiple dimensions, and the vector discreteness of multiple audio feature vectors in the dimensional space corresponding to at least one target vector dimension is the largest; according to at least one target vector dimension, multiple audio feature vectors are reduced in dimension to obtain audio feature vectors after dimension reduction; and multiple audio feature vectors after dimension reduction are distributed in a vector space corresponding to at least one target vector dimension to obtain the distribution of played audio on multiple playback platforms.
[0096] In this embodiment, after obtaining the final user audio playback sequence and its corresponding audio feature vector set, the terminal 102 can analyze the user's audio distribution information based on each audio feature vector in the audio feature vector set. The terminal 102 can determine multiple vector dimensions in the vector space based on multiple audio feature information corresponding to multiple audio feature vectors, so that the terminal 102 can obtain at least one target vector dimension based on multiple vector dimensions, wherein, under the observation angle corresponding to the target vector dimension, the multiple audio feature vectors corresponding to the above audio playback sequence have the largest vector discreteness in the dimensional space corresponding to the above at least one target vector dimension. The number of the above at least one target vector dimension can be two-dimensional or three-dimensional; the above dimensional space can be composed of at least one target hyperplane. After the terminal 102 determines the angles corresponding to which target vector dimensions are used to observe the above multiple audio feature vectors, it can reduce the dimensions of multiple high-dimensional audio feature vectors based on the target hyperplane corresponding to the determined target vector dimension to obtain the audio feature vectors in the dimensional space of the selected number of dimensions. Among them, the above-mentioned dimensionality reduction method can be a dimensionality reduction method of PCA (Principal Component Analysis). PCA converts a group of variables that may be correlated into a group of linearly unrelated variables through orthogonal transformation. The converted group of variables is called principal component. The terminal 102 can reduce the dimension of high-dimensional vectors through PCA to achieve vector space visualization. For example, the terminal 102 can first select a target vector dimension in the high-dimensional vector space. Under the observation angle corresponding to the target vector dimension, the discrete degree of multiple audio feature vectors is the largest. The terminal 102 can generate a corresponding target hyperplane based on the target vector dimension, so that the multiple audio feature vectors in the high-dimensional space can be reduced in dimension; if the number of target vector dimensions is greater than one, the terminal 102 can determine a target vector dimension that can maximize the discrete degree between multiple audio feature vectors under the current number of dimensions from the vector space after the dimensionality reduction through the target hyperplane corresponding to a target vector dimension, and determine a second target hyperplane based on the target vector dimension. Based on the first determined target hyperplane and the second target hyperplane, the multiple audio feature vectors are reduced in dimension. The terminal 102 may repeat the above process until the number of dimensions is reduced to the number of dimensions consistent with the selected number of dimensions. Thus, the terminal 102 may distribute the multiple audio feature vectors after dimensionality reduction in the vector space corresponding to the number of dimensions of at least one target vector, and obtain the distribution information results of the played audio on multiple playback platforms.
[0097] Specifically, Figure 4 As shown, Figure 4FIG. 1 is a schematic diagram of an interface for displaying audio distribution information in an embodiment. The terminal 102 may reduce the dimensions of the above-mentioned multiple high-dimensional audio feature vectors to two or three dimensions, so that the terminal 102 may Figure 4 As shown, the content distribution after the fusion of audio playback information of multiple playback platforms is displayed to the user. Among them, the terminal 102 can also receive the user's selection of a certain audio feature vector, and determine other played audios similar to the played audio corresponding to the audio feature vector, and mark them with corresponding identifiers, wherein the similarity of two audio feature vectors can be determined based on the number of times the two audio feature vectors appear together in the same audio playback sequence. For example, if the number of times two audio feature vectors appear in the same audio playback sequence at the same time is greater, the more similar the two audio feature vectors are, and the closer the positions of the two audio feature vectors in the vector space are, the similarity between each audio feature vector is judged.
[0098] Through this embodiment, the terminal 102 can determine the similarity between each audio feature vector based on the number of times each audio feature vector appears in the same audio playback sequence, and the terminal 102 can observe multiple audio feature vectors based on the dimension that maximizes the degree of discreteness between the multiple audio feature vectors, so that the distance relationship between each audio feature vector can be determined more accurately, thereby improving the accuracy of obtaining audio distribution information.
[0099] In one embodiment, the distribution of played audio on multiple playback platforms is determined based on the distance between multiple audio feature vectors of the target audio playback sequence in at least one vector dimension, including: obtaining the Euclidean distance between the multiple audio feature vectors in the vector space, and determining the similarity between the multiple audio feature vectors; and determining the distribution results of played audio for users on multiple playback platforms based on the similarity between the multiple audio feature vectors.
[0100] In this embodiment, in addition to reducing the dimensionality of the high-dimensional audio feature vector, the terminal 102 can also determine the audio distribution information between each audio feature vector by calculating the Euclidean distance. The terminal 102 can obtain the Euclidean distance between multiple audio feature vectors in the above vector space, so that the terminal 102 can determine the similarity between the multiple audio feature vectors based on the Euclidean distance. The terminal 102 can determine the distribution results of the played audio of users on multiple playback platforms based on the similarity between the multiple audio feature vectors. Specifically, Figure 5 As shown, Figure 5 The terminal 102 can obtain the user's selection of a certain audio, calculate the Euclidean distance between the audio feature vectors corresponding to other audios and the audio feature vector corresponding to the audio, and display the similarity between the other audios and the audio in a table.
[0101] Through this embodiment, the terminal 102 can determine the similarity between multiple audio feature vectors by means of Euclidean distance, and obtain the distribution result of the played audio based on the similarity, thereby improving the accuracy of obtaining audio distribution information.
[0102] In one embodiment, after determining the distribution of played audio on multiple playback platforms, it also includes: based on the orthogonal Prucker method, mapping the distribution of played audio in the target audio playback sequence corresponding to multiple different time periods to the same vector space to obtain the distribution position of each played audio; in the same vector space, determining the distribution position change information corresponding to the distribution result of the same played audio in the multiple time periods.
[0103] In this embodiment, the terminal 102 can also collect audio playback sequences in different time periods, so that the terminal 102 can determine the changes of the same played audio over time. For the above-mentioned multiple playback platforms, the terminal 102 can obtain the audio playback sequences collected based on multiple time periods, and based on the audio playback sequence of each time period, determine the distribution of the played audio corresponding to each time period, and then use the orthogonal Prucker method to map the distribution of the played audio in the target audio playback sequence corresponding to multiple different time periods to the same vector space. In the same vector space, the terminal 102 can determine the change information of the distribution results corresponding to the same played audio in the distribution positions corresponding to multiple time periods.
[0104] Specifically, Figure 6 As shown, Figure 6The interface diagram of the step of changing the audio distribution information in one embodiment. The semantics of words will change with the change of context in different time periods, and the semantic changes of these words may be stable or drastic. To study the semantic changes of words, the orthogonal Prucker method can be used to map one of the semantic spaces and approximate another semantic space, and then compare the degree of change of the vector of this word. In the analysis of audio distribution information, the terminal 102 can take the content vector space U in time period A, transform and map it to the content vector space V in time period B, and then U is transformed into U*, which is similar to V. Then the vector U(S) of the content S in U is transformed into U*(S), and at this time, the terminal 102 can compare the difference between U(S) and U*(S), and then it can be known that the user preference nature of the current content S in the two time periods from A to B has changed. For example, the terminal 102 can study the distribution changes of the played audio in the time dimension through the semantic change method in natural language processing. The comparison formula can be as follows: the first audio feature vector content vector of time window A vs. the first audio feature vector content vector of time window B; the second audio feature vector content vector of time window A vs. the second audio feature vector content vector of time window B. Among them, the first audio feature vector content may include the audio feature vectors corresponding to the played audio of each playback platform of the management party to which the terminal 102 belongs, and the second audio feature vector content may include the audio feature vectors corresponding to the played audio feature vector information of the playback platform of the management party to which the terminal 102 belongs and other playback platforms other than the management party to which the terminal 102 belongs. Among them, as Figure 7 As shown, Figure 7 FIG. 1 is a schematic diagram of an interface for the step of changing the audio distribution information in another embodiment. The terminal 102 can display the change of the audio feature vector in the time dimension through a two-dimensional platform or a three-dimensional surface. Figure 6 The lines in represent the position changes of the same audio feature vector in the vector space corresponding to different time periods.
[0105] Through this embodiment, the terminal 102 can obtain the distribution of audio feature vectors for multiple playback platforms in different time periods, so as to analyze the distribution change information of the audio feature vectors, thereby improving the accuracy of obtaining the audio distribution information.
[0106] In one embodiment, Figure 8 As shown, a method for obtaining audio distribution information is provided, and the method is applied to Figure 1 The terminal in is used as an example to illustrate, including the following steps:
[0107] Step S302, in response to the multi-platform audio distribution information viewing instruction, obtain the platform information of multiple playback platforms and the viewing dimension information, and obtain the distribution of multiple played audios of the multiple playback platforms under the viewing dimension information; the distribution of the played audios of the multiple playback platforms is determined based on the above method.
[0108] Among them, the user can trigger the multi-platform audio distribution information viewing instruction. After the terminal 102 obtains the instruction, it can obtain the multiple playback platform information and viewing dimension quantity information input by the user, so that the terminal 102 can obtain the distribution results of multiple played audios based on multiple playback platforms under the viewing dimension quantity information, so that the user can obtain the distribution of each played audio in the vector space.
[0109] Step S304: Display the distribution of multiple played audios on multiple playback platforms in the dimensional space corresponding to the viewing dimensional information.
[0110] The number of viewing dimensions may be a number selected by the user, which may be two-dimensional or three-dimensional. After the terminal 102 obtains the distribution result of the above-mentioned audio distribution information, it may display the distribution results of multiple played audios based on multiple playback platforms in the dimensional space corresponding to the number of viewing dimensions. The user may determine the distance relationship between each played audio based on the distribution result, thereby determining the user's preference distribution, such as the user's preference for different types of audio, which audios are similar audios, etc.
[0111] In the above-mentioned audio distribution information acquisition method, a target audio playback sequence including played audios corresponding to multiple users on multiple playback platforms is obtained based on the audio playback behavior of users on multiple playback platforms, and multiple audio feature information corresponding to each played audio in the audio playback sequence is obtained. For each played audio in the target audio playback sequence, an audio feature vector with multiple dimensions corresponding to the played audio is obtained, and multiple audio feature vectors corresponding to the target audio playback sequence are obtained. Based on the distance between the multiple audio feature vectors in at least one dimension in the vector space, the distribution results of the multiple played audios are determined. Compared with the traditional method of determining the user's audio distribution information only based on the number of times the user's audio is played, this solution determines the audio playback sequence through the user's playback behavior in multiple playback platforms, and determines audio feature vectors of multiple dimensions based on the sequence and audio feature information, thereby determining the distribution result of the audio information based on the distance between the multiple vectors observed from the perspective of at least one dimension in the vector space, thereby improving the accuracy of obtaining the audio playback information. In addition, the terminal 102 can also display the distribution results of the played audio in the dimensional space of the corresponding number of dimensions based on the viewing dimension quantity information and playback platform information selected by the user, which can improve the readability of the distribution results of the played audio, and can improve the accuracy of the relationship analysis between audios by visually displaying the audio distribution information.
[0112] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0113] Based on the same inventive concept, the embodiment of the present application also provides an audio distribution information acquisition device for implementing the audio distribution information acquisition method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more audio distribution information acquisition device embodiments provided below can refer to the limitations of the audio distribution information acquisition method above, and will not be repeated here.
[0114] In one embodiment, Fig. 9As shown, an audio distribution information acquisition device is provided, comprising: a sequence acquisition module 500, a vector acquisition module 502 and a distribution module 504, wherein:
[0115] The sequence acquisition module 500 is used to obtain a target audio playback sequence based on the user's played audio on multiple playback platforms.
[0116] The vector acquisition module 502 is used to obtain the audio feature vector of each played audio in the target audio playback sequence to obtain multiple audio feature vectors corresponding to the target audio playback sequence; wherein the audio feature vector has multiple vector dimensions in the vector space, and each vector dimension represents an audio feature information of the played audio.
[0117] The distribution module 504 is used to determine the distribution results of multiple played audios according to the distances of multiple audio feature vectors in at least one dimension in the vector space.
[0118] In one embodiment, the above-mentioned sequence acquisition module 500 is specifically used to obtain, for each playback platform, multiple initial audio playback sequences obtained from the played audios corresponding to multiple users in the playback platform; according to a set ratio, a corresponding number of initial audio playback sequences are selected from the multiple initial audio playback sequences corresponding to each playback platform; and the initial audio playback sequences of the selected multiple playback platforms are merged to obtain a target audio playback sequence.
[0119] In one embodiment, the above-mentioned device also includes: a ratio acquisition module, which is specifically used to obtain the number of users using each playback platform; determine the user ratio between multiple playback platforms based on the number of users using; and determine the set ratio corresponding to multiple playback platforms based on the user ratio.
[0120] In one embodiment, the above-mentioned vector acquisition module 502 is specifically used to obtain multiple audio feature information of each played audio in the target audio playback sequence; vectorize the multiple audio feature information of the played audio through a preset natural language processing model to obtain the audio feature vector of the played audio.
[0121] In one embodiment, the sequence acquisition module 500 is specifically used to acquire multiple played audios of multiple users of each playback platform in the same time period from multiple playback platforms to generate multiple initial audio playback sequences corresponding to each playback platform.
[0122] In one embodiment, the above-mentioned vector acquisition module 502 is specifically used to obtain multiple initial audio feature vectors of multiple played audios of each initial audio playback sequence corresponding to the target audio playback sequence, and obtain multiple groups of audio feature vector sets corresponding to the multiple initial audio playback sequences based on the multiple initial audio feature vectors; wherein each group of audio feature vector sets includes multiple initial audio feature vectors corresponding to each initial audio playback sequence; for each played audio in the multiple initial audio playback sequences, according to the number of times each played audio is played in different sequences, the initial audio feature vectors in the multiple groups of audio feature vector sets are merged to obtain the audio feature vector of the played audio in the target audio playback sequence.
[0123] In one embodiment, the above-mentioned vector acquisition module 502 is specifically used to obtain the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets, and obtain the average number of times the same played audio is played in the initial audio playback sequence corresponding to each two groups of audio feature vector sets; determine the average audio feature vector of the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets based on the average value, replace the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets with the average audio feature vector, and obtain a new audio feature vector set; detect whether the number of new audio feature vector sets is one, if not, return to the step of obtaining the initial audio feature vectors corresponding to the same played audio in each two groups of audio feature vector sets; if so, use the audio feature vectors in the new audio feature vector set as the audio feature vectors of the played audio in the target audio playback sequence.
[0124] In one embodiment, the distribution module 504 is specifically used to determine at least one target vector dimension among multiple vector dimensions, wherein the vector discreteness of multiple audio feature vectors in the dimensional space corresponding to the at least one target vector dimension is the largest; according to the at least one target vector dimension, multiple audio feature vectors are reduced in dimension to obtain reduced-dimensional audio feature vectors; multiple reduced-dimensional audio feature vectors are distributed in the vector space corresponding to the at least one target vector dimension to obtain the distribution of played audio on multiple playback platforms.
[0125] In one embodiment, the distribution module 504 is specifically used to determine the similarity between multiple audio feature vectors based on the Euclidean distance between the multiple audio feature vectors in the vector space; and determine the distribution results of the played audio on multiple playback platforms based on the similarity between the multiple audio feature vectors.
[0126] In one embodiment, the above-mentioned device also includes: a change module, which is used to map the distribution of played audio in the target audio playback sequence corresponding to multiple different time periods to the same vector space based on the orthogonal Prucker method, so as to obtain the distribution position of each played audio; in the same vector space, determine the change information of the distribution position of the same played audio in multiple different time periods.
[0127] In one embodiment, Fig.10 As shown, a device for obtaining audio distribution information is provided, including: a response module 600 and a display module 602, wherein:
[0128] Response module 600 is used to respond to a multi-platform audio distribution information viewing instruction, obtain platform information of multiple playback platforms and viewing dimension information, and obtain the distribution of multiple played audios on multiple playback platforms under the viewing dimension information; the distribution of played audios on multiple playback platforms is determined based on the above method.
[0129] The display module 602 is used to display the distribution of multiple played audios of multiple playback platforms in the dimensional space corresponding to the viewing dimensional information.
[0130] Each module in the above-mentioned audio distribution information acquisition device can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above modules.
[0131] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.11 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for obtaining audio distribution information is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0132] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0133] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above-mentioned audio distribution information acquisition method when executing the computer program.
[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for obtaining audio distribution information is implemented.
[0135] In one embodiment, a computer program product is provided, including a computer program, which implements the above-mentioned audio distribution information acquisition method when executed by a processor.
[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0137] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0138] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0139] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for obtaining audio distribution information, characterized in that: The method comprises: Obtain a target audio playback sequence according to the played audios of the user on multiple playback platforms; Acquire the audio feature vector of each played audio in the target audio playback sequence to obtain multiple audio feature vectors corresponding to the target audio playback sequence, including: combining the initial audio feature vectors of each played audio in the initial audio playback sequence into an audio feature vector set; merging the initial audio feature vectors of the played audio in different audio feature vector sets according to the number of times the played audio is played in different initial audio playback sequences, including: acquiring the initial audio feature vectors corresponding to the same played audio in each of two groups of audio feature vector sets, and the average of the number of times the played audio is played in the initial audio playback sequences corresponding to each of the two groups of audio feature vector sets; determining the average audio feature vector of the initial audio feature vectors corresponding to the same played audio in each of the two groups of audio feature vector sets according to the average, and determining the audio feature vector of the played audio in the target audio playback sequence according to the average audio feature vector; wherein the audio feature vector has multiple vector dimensions in the vector space, and each vector dimension represents an audio feature information of the played audio; The distribution of the played audios of the multiple playback platforms is determined according to the distances of the multiple audio feature vectors of the target audio playback sequence in at least one vector dimension.
2. The method according to claim 1, characterized in that The step of obtaining a target audio playback sequence according to the played audios of the user on the multiple playback platforms includes: For each playback platform, obtaining multiple initial audio playback sequences obtained from the played audios corresponding to multiple users in the playback platform; According to the set ratio, a corresponding number of initial audio playback sequences are selected from multiple initial audio playback sequences corresponding to each playback platform; The initial audio playback sequences of the selected multiple playback platforms are combined to obtain the target audio playback sequence.
3. The method according to claim 2, characterized in that The step of obtaining, for each playback platform, a plurality of initial audio playback sequences obtained from the played audios respectively corresponding to a plurality of users on the playback platform comprises: From multiple playback platforms, multiple played audios of multiple users of each playback platform in the same time period are obtained to generate multiple initial audio playback sequences corresponding to each of the playback platforms.
4. The method according to claim 1, characterized in that The step of determining the audio feature vector of the played audio in the target audio playback sequence according to the average audio feature vector includes: Replacing the initial audio feature vectors corresponding to the same played audio in each of the two groups of audio feature vector sets with the average audio feature vector to obtain a new audio feature vector set; Detecting whether the number of the new audio feature vector sets is one, and if not, returning to the step of obtaining initial audio feature vectors corresponding to each of two sets of audio feature vector sets for the same played audio; If so, the audio feature vector in the new audio feature vector set is used as the audio feature vector of the played audio in the target audio playback sequence.
5. The method according to claim 1, characterized in that The determining, according to the distances of the multiple audio feature vectors of the target audio playback sequence in at least one vector dimension, the distribution of the played audios of the multiple playback platforms comprises: Determining at least one target vector dimension among the multiple vector dimensions, wherein the vector discreteness of the multiple audio feature vectors in the dimensional space corresponding to the at least one target vector dimension is the largest; According to the at least one target vector dimension, reducing the dimensions of the multiple audio feature vectors to obtain audio feature vectors after dimension reduction; A plurality of audio feature vectors after dimensionality reduction are distributed in a vector space corresponding to the dimension of the at least one target vector to obtain distribution of the played audios of the plurality of playback platforms.
6. The method according to claim 1, characterized in that The determining, according to the distances of the multiple audio feature vectors of the target audio playback sequence in at least one vector dimension, the distribution of the played audios of the multiple playback platforms comprises: Determining similarities between the multiple audio feature vectors according to the Euclidean distances between the multiple audio feature vectors in the vector space; According to the similarities between the multiple audio feature vectors, distribution results of the played audios of the multiple playback platforms are determined.
7. The method according to claim 1, characterized in that The step of obtaining an audio feature vector of each played audio in the target audio playback sequence includes: Obtaining multiple audio feature information of each of the played audio in the target audio playback sequence; Vectorized processing is performed on multiple audio feature information of the played audio through a preset natural language processing model to obtain an audio feature vector of the played audio.
8. The method according to claim 1, characterized in that The target audio playback sequence includes target audio playback sequences corresponding to multiple different time periods; After determining the distribution of the played audio on the multiple playback platforms, the method further includes: Based on the orthogonal Prucker method, the distribution of the played audio in the target audio playback sequence corresponding to the multiple different time periods is mapped to the same vector space to obtain the distribution position of each played audio; In the same vector space, change information of the distribution position of the same played audio in the multiple different time periods is determined.
9. A method for obtaining audio distribution information, characterized in that: The method comprises: In response to a multi-platform audio distribution information viewing instruction, platform information and viewing dimension information of multiple playback platforms are obtained, and distribution of multiple played audios of the multiple playback platforms under the viewing dimension information is obtained; the distribution of the played audios of the multiple playback platforms is determined based on the method described in any one of claims 1 to 8; In the dimensional space corresponding to the viewing dimensional information, the distribution of multiple played audios of the multiple playback platforms is displayed.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Audio resource processing method and device, equipment and storage medium
CN111046225A
Audio recognition method and device, computer equipment and storage medium
CN113823320A