Live-based virtual resource configuration method, computer device and storage medium
By using voiceprint feature matching technology, the problem of identifying non-real-faced anchors on live streaming platforms has been solved, enabling efficient and low-cost anchor recruitment and management, and improving the quality of live streaming.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2023-03-31
- Publication Date
- 2026-07-21
Smart Images

Figure CN116506650B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of live streaming technology, and in particular to a method for configuring virtual resources based on live streaming, computer equipment, and storage media. Background Technology
[0002] With the development of internet and communication technologies, society has entered an era of intelligent interconnection. Interacting, entertaining, and working on the internet is becoming increasingly common. Among these, live streaming technology is particularly prevalent. People can watch or participate in live streams anytime, anywhere through smart devices, greatly enriching their lives and broadening their horizons.
[0003] With the rapid development of the live streaming industry, the forms of live streaming interaction have also become diverse. During a live stream, the streamer can perform, such as singing or dancing. Viewers can interact with the streamer while watching their performance, such as sending virtual gifts, posting comments, or giving likes. However, currently, the recruitment costs for streamers to acquire virtual resources from the platform, such as anime / manga streamers, are very high, requiring extensive manual verification to ensure that streamers are properly vetted for acquiring virtual resources. Summary of the Invention
[0004] The main technical problem addressed in this application is to provide a method for configuring virtual resources based on live streaming, computer equipment, and storage media, which facilitates the management of virtual resources acquired by broadcasters.
[0005] To address the aforementioned technical problems, this application provides a technical solution: a method for configuring virtual resources based on live streaming. This method includes: a broadcaster terminal uploading an audio file to a server; the server extracting the broadcaster's voiceprint features from the audio file; the server performing feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library; if the broadcaster's voiceprint features do not exist in the preset voiceprint library, the server sends virtual resources to the broadcaster terminal based on the broadcaster's broadcaster identity identifier; and the broadcaster terminal responding to the broadcaster's selection of virtual resources generates a live video combining virtual resources and broadcaster features.
[0006] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a method for configuring virtual resources based on live streaming. This method includes: extracting the voiceprint features of the broadcaster from the audio file uploaded by the broadcaster's terminal; performing feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library; if it is determined that the broadcaster's voiceprint features do not exist in the preset voiceprint library, then issuing virtual resources to the broadcaster's terminal based on the broadcaster's identity identifier, so that the broadcaster's terminal can respond to the broadcaster's selection of virtual resources and generate a live video combining virtual resources and broadcaster features.
[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a computer device, which includes a processor, a memory, and a communication circuit; the communication circuit and the memory are coupled to the processor; the memory stores a computer program, and the processor is used to execute the computer program to implement the live streaming-based virtual resource configuration method provided in this application as described above.
[0008] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium storing a computer program, which is executed by a processor to implement the live streaming-based virtual resource configuration method provided in this application as described above.
[0009] The beneficial effects of this application are as follows: Unlike existing technologies, the server extracts the broadcaster's voiceprint features from the audio file uploaded by the broadcaster's terminal and performs feature matching between these features and those in a preset voiceprint database to determine if the broadcaster's voiceprint features exist in the database. If the broadcaster's voiceprint features are not found in the database, the server can send virtual resources to the broadcaster's terminal based on the broadcaster's identity identifier. The broadcaster's terminal can respond to the broadcaster's selection of virtual resources and generate a live video combining the virtual resources and the broadcaster's features. Thus, identifying and recruiting broadcasters through machine recognition reduces the difficulty of recruiting broadcasters during virtual resource broadcasting, allowing for broadcaster identity verification without relying on real faces. This improves the efficiency and cost of broadcaster recruitment, reduces manual intervention, and enhances the accuracy of recruitment, thereby improving the quality of live broadcasts and streamer management. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of the system composition of an embodiment of the live streaming system of this application;
[0011] Figure 2 This is a flowchart illustrating the first embodiment of the virtual resource configuration method based on live streaming in this application;
[0012] Figure 3 This is a flowchart illustrating the second embodiment of the virtual resource configuration method based on live streaming in this application.
[0013] Figure 4 This is a timing diagram of the second embodiment of the virtual resource configuration method based on live streaming in this application;
[0014] Figure 5 This is a schematic diagram of the server interface of the second embodiment of the virtual resource configuration method based on live streaming in this application;
[0015] Figure 6 This is a flowchart illustrating the preset deep learning model in the second embodiment of the live streaming-based virtual resource configuration method of this application;
[0016] Figure 7 This is a schematic diagram of the interface of the broadcaster terminal in the second embodiment of the virtual resource configuration method based on live streaming in this application;
[0017] Figure 8 This is a schematic diagram of virtual resources in the second embodiment of the virtual resource configuration method based on live streaming in this application;
[0018] Figure 9 This is a schematic diagram of the circuit structure of an embodiment of the computer device of this application;
[0019] Figure 10 This is a schematic diagram of the circuit structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0021] With the rapid development of the live streaming industry, the forms of live streaming interaction have also become diverse. During a live stream, the streamer performs on their own device, and viewers watch the performance on their own devices and interact with the streamer. Viewers can also send virtual gifts to the streamer to interact with them.
[0022] Through long-term research, the inventors discovered that in existing live streaming processes, most streamers use their real faces to start broadcasts, resulting in a relatively simple interactive format. Some live streams utilize virtual resources, such as anime / manga avatars. However, because streamers don't use their real faces during the recruitment process, streamer identification is difficult, making accurate identification challenging and hindering the management of streamers broadcasting using virtual resources. To address these technical problems, this application proposes the following embodiments.
[0023] like Figure 1As shown in the embodiment of the live streaming system described in this application, the live streaming system 1 may include a server 10, a broadcaster terminal 20, and a viewer terminal 30. The broadcaster terminal 20 and the viewer terminal 30 can be electronic devices; specifically, they are electronic devices with corresponding client programs installed, i.e., client terminals. The electronic devices can be mobile terminals, computers, servers, or other terminals, etc. Mobile terminals can be mobile phones, laptops, tablets, smart wearable devices, etc., and computers can be desktop computers, etc.
[0024] Server 10 can pull live data streams from broadcaster terminal 20, process the acquired live data streams accordingly, and then push them to viewer terminals 30. Viewer terminals 30 can then watch the live stream from the broadcaster or guest. Live data stream mixing can occur at least among server 10, broadcaster terminal 20, and viewer terminals 30. Video and audio communication is possible between broadcaster terminals 20 and between broadcaster terminals 20 and viewer terminals 30. During the live stream, broadcaster terminal 20 can push the live data stream, including video, to server 10, which in turn pushes the corresponding live data to each viewer terminal 30 in the live stream room corresponding to broadcaster terminal 20. Broadcaster terminals 20 and viewer terminals 30 can then display the corresponding live stream content in their respective live stream rooms. Specifically, server 10 can be, for example, a server cluster, which can not only be used to collect and push live data streams, but also to process business requests and related matters, such as storing and processing business-related data generated during the live stream, such as handling virtual gift giving, virtual currency recharge and consumption, public screen information sending and receiving, authentication, live chat, and automatic identification of sensitive words / images.
[0025] Of course, the terms "anchor terminal 20" and "viewer terminal 30" are relative. The terminal that is in the process of live streaming is the anchor terminal 20, and the terminal that is watching the live stream is the viewer terminal 30.
[0026] like Figure 2 As shown, the first embodiment of the virtual resource configuration method based on live streaming in this application may include the following steps: M100: The broadcaster terminal uploads an audio file to the server. M200: The server extracts the broadcaster's voiceprint features from the audio file. M300: The server performs feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library. M400: If it is determined that the broadcaster's voiceprint features do not exist in the preset voiceprint library, the server issues virtual resources to the broadcaster terminal based on the broadcaster's broadcaster identity identification identifier. M500: The broadcaster terminal responds to the broadcaster's selection operation for virtual resources and generates a live video combining virtual resources and broadcaster features.
[0027] During the configuration of virtual resources based on live streaming, the server 10 can interact with the broadcaster terminal 20 to complete the configuration. Specifically, the broadcaster terminal 20 can upload audio files to the server 10. After receiving the audio files, the server 10 can extract the broadcaster's voiceprint features from the audio files and perform feature matching with voiceprint features in a preset voiceprint database to determine whether the broadcaster's voiceprint features exist in the database. If the server determines that the broadcaster's voiceprint features do not exist in the preset voiceprint database, the server 10 can issue virtual resources to the broadcaster terminal 20 based on the broadcaster's identity identification identifier. The broadcaster terminal 20 can respond to the broadcaster's selection of virtual resources and generate a live video combining virtual resources and broadcaster features. In this way, identifying and recruiting broadcasters through machine recognition helps reduce the difficulty of identification during broadcaster recruitment in the process of virtual resource broadcasting, allowing for broadcaster identification without relying on real faces. On the one hand, it can improve the efficiency of recruiting anchors, reduce the cost of recruiting anchors, and reduce the manual involvement in the process. On the other hand, it can improve the accuracy of recruiting anchors, which is conducive to improving the quality of anchor live broadcasts and managing anchors.
[0028] like Figure 3 As shown, the second embodiment of the virtual resource configuration method based on live streaming in this application can use server 10 as the execution subject. This embodiment may include the following steps: S100: Extract the voiceprint features of the broadcaster from the audio file uploaded by the broadcaster terminal. S200: Perform feature matching between the broadcaster's voiceprint features and the voiceprint features in the preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library. S300: If it is determined that the broadcaster's voiceprint features do not exist in the preset voiceprint library, then send virtual resources to the broadcaster terminal according to the broadcaster's broadcaster identity identification mark, so that the broadcaster terminal can respond to the broadcaster's selection operation of virtual resources and generate a live video combining virtual resources and broadcaster features.
[0029] During the live stream, the voiceprint features of the broadcaster are extracted from the audio file uploaded by the broadcaster terminal 20. These voiceprint features are then matched against those in a pre-set voiceprint database to determine if they exist. If the voiceprint features are not found in the database, virtual resources are sent to the broadcaster terminal 20 based on the broadcaster's identity identifier. This allows the broadcaster terminal 20 to respond to the broadcaster's selection of virtual resources and generate a live video combining the virtual resources and the broadcaster's features. This machine-based identification and recruitment of broadcasters reduces the difficulty of recruitment during virtual resource broadcasts, allowing for identification without relying on real faces. This improves recruitment efficiency, reduces costs, and minimizes manual intervention. Furthermore, it enhances recruitment accuracy, leading to higher live stream quality and better broadcaster management.
[0030] The method described in this embodiment can be applied to scenarios where the broadcaster uses virtual resources to start a live broadcast, such as... Figure 4 As shown, the following is a detailed description of this embodiment, with server 10 as the execution subject.
[0031] S100: Extract the voiceprint features of the broadcaster from the audio file uploaded by the broadcaster's terminal.
[0032] The audio file can be obtained by the broadcaster terminal 20 through a recording device that is communicatively connected to the broadcaster terminal 20. During the live broadcast, the broadcaster terminal 20 can upload the audio file to the server 10, which then sends the audio file to all viewer terminals 30 corresponding to all viewers currently in the live broadcast room, so that the viewer terminals 30 can play the audio file while playing the live broadcast.
[0033] Voiceprint features are the speech features contained in speech that can represent and identify the speaker, as well as the speech models built based on these features. Since voiceprint recognition is the process of identifying the speaker corresponding to a speech segment based on the voiceprint features of the speech to be recognized, it is necessary to extract the voiceprint features of the broadcaster in order to identify the broadcaster's identity based on voiceprint features.
[0034] In one implementation, before extracting the broadcaster's voiceprint features from the audio file uploaded by the broadcaster's terminal, the following steps may be included:
[0035] S110: Obtain the audio file obtained by the anchor reading a preset text during the trial broadcast on the anchor terminal.
[0036] A trial broadcast can be a process by which a broadcaster obtains audio files to verify their identity before officially starting the broadcast.
[0037] The preset text can be pre-set in server 10 and used to acquire audio files to identify the broadcaster. By setting the preset text, it is easy to limit the content and duration of the broadcaster's reading during the trial broadcast, thereby improving the efficiency of broadcaster voiceprint feature extraction.
[0038] During the trial broadcast, the anchor can read a preset text aloud. The anchor terminal 20 acquires the audio of the anchor reading the preset text, forms an audio file, and sends the audio file to the server 10. The server 10 acquires the audio file obtained by the anchor terminal 20 from the anchor reading the preset text during the trial broadcast, and can then obtain the anchor's voiceprint characteristics based on the audio file.
[0039] In one implementation, the following steps in S110 can be used to obtain the audio file obtained by the broadcaster reading a preset text during the trial broadcast:
[0040] S111: Retrieves the playback link generated after the trial playback ends, which is used to link to the audio file.
[0041] S112: Obtain audio files via playback links.
[0042] Playback links can be generated based on the audio from the trial broadcast and are used to link to audio files. Specifically, after the broadcaster finishes the trial broadcast on broadcaster terminal 20, broadcaster terminal 20 can automatically generate the corresponding playback link. Server 10 can obtain the playback link and access the audio file through it. The audio file can then be streamed into a pre-configured process on server 10. Figure 5 As shown, Figure 5 This is the interface of server 10 when running the configuration process. Through this configuration process, the voiceprint features of the broadcaster can be extracted from the audio file, and the virtual resources can be configured for the broadcaster based on the voiceprint features.
[0043] In one implementation, the steps for extracting the broadcaster's voiceprint features from the audio file uploaded by the broadcaster's terminal can be referenced in S100 as follows:
[0044] S120: Perform noise reduction processing on the audio file and extract the audio segments containing the anchor's voice.
[0045] Since audio files may contain background sounds other than the broadcaster's voice, such as background music and ambient sounds, these background sounds can interfere with the process of extracting the broadcaster's voiceprint features. Therefore, it is necessary to perform noise reduction processing on the audio files and extract the audio segments that carry the broadcaster's voice and can be used to extract the broadcaster's voiceprint features. This will help improve the accuracy of broadcaster voiceprint feature extraction and, consequently, the accuracy of broadcaster identification.
[0046] In one implementation, the steps for denoising the audio file and extracting the speech segment carrying the broadcaster's voice can be referenced in S120 as follows:
[0047] S121: Remove the background sound from the audio file and perform noise reduction on the audio file to obtain the foreground sound.
[0048] Since background noise can interfere with the broadcaster's voice, it can be removed from the audio file. Specifically, a deep learning-based sound event algorithm can be used to remove background noise from the audio file.
[0049] To eliminate noise interference in audio files, background noise can be removed before denoising the audio file to obtain the foreground noise. Specifically, deep learning-based denoising algorithms can be used to eliminate noise interference in the audio file during the denoising process.
[0050] For example, if an audio file contains both the broadcaster's voice and background music, it will affect the extraction of the broadcaster's voiceprint features during the voiceprint extraction process, thus affecting the accuracy of voiceprint recognition. By using a deep learning-based sound event algorithm to remove the background music from the audio file, and then using a deep learning-based noise reduction algorithm to eliminate noise interference, the foreground sound portion carrying the broadcaster's voice is finally obtained.
[0051] S122: Remove non-speech segments that do not carry the anchor's voice from the foreground audio portion to obtain the speech segments.
[0052] Non-audio segments can include audio segments that are not read from the pre-set text by the anchor during the test broadcast, such as breathing sounds, coughing sounds, etc. made by the anchor during the test broadcast.
[0053] Since non-speech segments do not carry the voice information used for extraction and recognition, after removing the background sound portion of the audio file and performing noise reduction to obtain the foreground sound portion, non-speech segments that do not carry the announcer's voice can be removed from the foreground sound portion to obtain the speech segments. Specifically, in the process of removing non-speech segments that do not carry the announcer's voice from the foreground sound portion, a deep learning-based speech activity detection algorithm can be used to remove non-speech segments that do not carry the announcer's voice from the foreground sound portion, thereby improving the accuracy of announcer voiceprint features and voiceprint recognition.
[0054] S130: Extract the broadcaster's voiceprint features from the speech segment.
[0055] After removing and denoising audio segments to obtain speech segments containing only the broadcaster's voice, the broadcaster's voiceprint features can be extracted from the speech segments, thus enabling feature matching based on the broadcaster's voiceprint features.
[0056] In one implementation, the following steps can be referenced to extract the broadcaster's voiceprint features from a speech segment:
[0057] S131: Divide the speech segment into several speech blocks.
[0058] To facilitate the processing of audio segments, they can be divided into several audio blocks. Specifically, the division of audio segments into audio blocks can be carried out according to pre-set division criteria in server 10. For example, the audio segments can be divided according to preset audio block size or duration. If the total duration of the audio segment is 30 seconds and the preset audio block duration is 5 seconds, then the audio segment can be divided into 6 audio blocks.
[0059] S132: Receive the result of determining whether each voice block is the voice of the broadcaster.
[0060] After dividing a speech segment into several speech blocks, it is possible to determine whether each speech block is the voice of a broadcaster and receive the determination result. Specifically, in the process of determining whether each speech block is the voice of a broadcaster, it can be done manually to confirm whether each speech block is the voice of the broadcaster being tested, thereby obtaining a determination result. Alternatively, a deep learning recognition model can be used to recognize each speech block to determine whether each speech block is the voice of a broadcaster and output the determination result.
[0061] S133: If the determination result is yes, then extract the voiceprint features of the speech block.
[0062] If the determination result is that the voice block is the voice of the anchor, then the voiceprint features of the voice block can be extracted as the voiceprint features of the anchor.
[0063] If the result indicates that the voice block is not the broadcaster's voice, then the voice block can be discarded and its voiceprint features will not be extracted.
[0064] In one implementation, the following steps included in S100 can be used as a reference for how to extract the broadcaster's voiceprint features from the audio file uploaded by the broadcaster's terminal:
[0065] S140: The temporal signal of the audio file is processed by a preset deep learning model to obtain a feature vector as the voiceprint feature of the anchor.
[0066] After obtaining several audio blocks containing the broadcaster's voice, voiceprint features can be extracted using a pre-defined deep learning model. Specifically, the temporal signal of the audio file can be processed using a pre-defined deep learning model to obtain a feature vector that serves as the broadcaster's voiceprint feature.
[0067] In one implementation, the steps in S140 for extracting features from the temporal signal of an audio file using a pre-defined deep learning model to obtain a feature vector representing the broadcaster's voiceprint can be referenced:
[0068] S141: Input the audio time-domain signal of the audio file into the first convolutional layer.
[0069] S142: Input the convolution result output from the first convolutional layer into the first residual neural network.
[0070] S143: Input the processing results of the first residual neural network into the second residual neural network and the second convolutional layer respectively.
[0071] S144: Input the processing results of the second residual neural network into the third residual neural network and the second convolutional layer respectively.
[0072] S145: Input the convolution result of the second convolutional layer into the max pooling layer.
[0073] S146: Input the pooling result of the max pooling layer into the fully connected layer, and output the feature vector through the fully connected layer.
[0074] like Figure 6 As shown, the preset deep learning model may include a first convolutional layer, a first residual neural network, a second residual neural network, a third residual neural network, a second convolutional layer, a max pooling layer, and a fully connected layer, connected in sequence. Specifically, the first and second convolutional layers are both one-dimensional convolutional layers.
[0075] In the process of feature extraction from the temporal signal of an audio file using a pre-defined deep learning model, the audio temporal signal is input into a first convolutional layer. The convolution result from the first convolutional layer is then input into a first residual neural network. The processing result from the first residual neural network is then input into a second residual neural network and a second convolutional layer. The processing result from the second residual neural network is then input into a third residual neural network and a second convolutional layer. The convolution result from the second convolutional layer is then input into a max pooling layer. Finally, the pooling result from the max pooling layer is input into a fully connected layer, which outputs a one-dimensional feature vector. This one-dimensional feature vector is used as the broadcaster's voiceprint feature. By combining a temporal delay neural network and a residual network, an effective neural network for extracting voiceprint features is constructed—the pre-defined deep learning model. This pre-defined deep learning model has a small size and strong learning ability, thereby improving the accuracy of the broadcaster's voiceprint feature extraction process.
[0076] After extracting the voiceprint features of the broadcaster, the following steps can be performed:
[0077] S200: Perform feature matching between the anchor's voiceprint features and the voiceprint features in the preset voiceprint database to determine whether the anchor's voiceprint features exist in the preset voiceprint database.
[0078] The preset voiceprint library may include a voiceprint library that stores all voiceprint features corresponding to all streamers currently broadcasting using virtual resources on the live streaming platform.
[0079] When recruiting broadcasters during the virtual resource broadcasting process, to avoid granting the same broadcaster virtual resource broadcasting privileges repeatedly, it is necessary to identify the broadcaster to determine whether they have already activated virtual resource broadcasting privileges. Specifically, since the preset voiceprint database stores all voiceprint features corresponding to all broadcasters currently using virtual resources on the live streaming platform, feature matching can be performed between the broadcaster's voiceprint features and those in the preset voiceprint database to determine whether the broadcaster has activated virtual resource broadcasting privileges.
[0080] In one implementation, the following steps in S200 can be used to perform feature matching between the broadcaster's voiceprint features and those in a preset voiceprint database to determine whether the broadcaster's voiceprint features exist in the preset voiceprint database:
[0081] S210: Calculate the cosine similarity between the anchor's voiceprint features and each voiceprint feature in the preset voiceprint library.
[0082] S220: Determine whether the cosine similarity is greater than the preset threshold.
[0083] S230: If it is less than, then it is determined that the anchor's voiceprint features do not exist in the preset voiceprint library.
[0084] In the process of matching the broadcaster's voiceprint features with those in a pre-defined voiceprint database, the cosine similarity between each feature and the pre-defined voiceprint database can be calculated. Each cosine similarity is then compared to a pre-defined threshold to determine if it is greater than the threshold. If the cosine similarity is less than the threshold, it can be determined that the two voiceprint features do not originate from the same speaker, thus indicating a mismatch between the broadcaster's voiceprint features and those in the pre-defined voiceprint database, and consequently, that the broadcaster's voiceprint features do not exist in the database.
[0085] If the cosine similarity is greater than the preset threshold, it can be determined that the two voiceprint features come from the same speaker. Therefore, it can be determined that the anchor's voiceprint features match the voiceprint features in the preset voiceprint library, and thus it can be determined that the anchor's voiceprint features exist in the preset voiceprint library. Virtual resource live streaming permissions can be denied for the duplicate anchor, that is, the anchor's virtual resource application is rejected.
[0086] In one implementation, the following steps in S200 can be used to determine whether the anchor's voiceprint features exist in the preset voiceprint database by performing feature matching between the anchor's voiceprint features and the voiceprint features in the preset voiceprint database:
[0087] S240: Determine whether the anchor's voiceprint features match the voiceprint features in the preset voiceprint library.
[0088] S250: If there is no match, determine whether the anchor identity identifier corresponding to the anchor's voiceprint feature is the same as the anchor identity identifier corresponding to the voiceprint feature in the preset voiceprint library.
[0089] S260: If they are different, it is determined that the anchor's voiceprint features do not exist in the preset voiceprint database.
[0090] The streamer identification identifier can be a marker used to distinguish streamers within virtual resource gameplay. Specifically, the streamer identification identifier can be stored in a preset voiceprint database, corresponding to the streamer's voiceprint features, thus facilitating the identification of the corresponding streamer based on the streamer identification identifier corresponding to the streamer's voiceprint features. The streamer identification identifier can be carried by the streamer's voiceprint features, i.e., it already exists for the streamer before voiceprint matching, or it can be assigned to the streamer by server 10.
[0091] In determining whether a broadcaster's voiceprint features match those in a pre-defined voiceprint database, the method based on cosine similarity described above can be used, and will not be repeated here. If the cosine similarity is less than a pre-defined threshold, it can be determined that the broadcaster's voiceprint features do not match those in the pre-defined voiceprint database. Furthermore, it can be determined whether the broadcaster's identity identifier corresponding to the broadcaster's voiceprint features is the same as the broadcaster's identity identifier corresponding to the voiceprint features in the pre-defined voiceprint database.
[0092] When extracting the voiceprint features of a broadcaster, if the broadcaster possesses a broadcaster identification identifier, this identifier can be obtained simultaneously. Furthermore, if a mismatch is found between the broadcaster's voiceprint and those in a pre-defined voiceprint database, it can be determined whether the broadcaster's voiceprint feature matches the corresponding identifier in the pre-defined database. If they differ, it can be concluded that the broadcaster's voiceprint feature does not exist in the pre-defined database.
[0093] In one implementation, after determining whether the broadcaster's voiceprint features match those in a preset voiceprint database, the following steps may be included:
[0094] S251: If there is no match, determine whether the anchor's voiceprint features have a corresponding anchor identity identification mark.
[0095] S252: If it is determined that there is no corresponding anchor identity identification mark, then an anchor identity identification mark is assigned to the anchor, and the anchor identity identification mark is bound to the anchor voiceprint feature. Then, it is determined whether the anchor identity identification mark corresponding to the anchor voiceprint feature is the same as the anchor identity identification mark corresponding to the voiceprint feature in the preset voiceprint library.
[0096] If the broadcaster's voiceprint features do not match those in the preset voiceprint database, it can be determined whether the broadcaster's voiceprint features have a corresponding broadcaster identification identifier. If there is no corresponding broadcaster identification identifier, the server 10 can assign a broadcaster identification identifier to the broadcaster and bind the broadcaster identification identifier to the broadcaster's voiceprint features, and then perform a judgment to determine whether the broadcaster identification identifier corresponding to the broadcaster's voiceprint features is the same as the broadcaster identification identifier corresponding to the voiceprint features in the preset voiceprint database.
[0097] In one implementation, after determining whether the anchor identity identifier corresponding to the anchor's voiceprint features is the same as the anchor identity identifier corresponding to the voiceprint features in the preset voiceprint database, the following steps may be included:
[0098] S270: If they are the same, it is determined that the anchor's voiceprint features exist in the preset voiceprint library.
[0099] If the anchor's voiceprint feature corresponds to the same anchor identity identifier as the anchor identity identifier in the preset voiceprint database, that is, if the anchor identity identifier already has other corresponding voiceprint features stored in the preset voiceprint database, then it can be determined that the anchor's voiceprint feature exists in the preset voiceprint database.
[0100] In one implementation, after determining that the broadcaster's voiceprint features do not exist in the preset voiceprint database, the following steps can be performed:
[0101] S280: If it is determined that the anchor's voiceprint features do not exist in the preset voiceprint library, then add the anchor's voiceprint features to the preset voiceprint library to update the preset voiceprint library.
[0102] Since the broadcaster's voiceprint characteristics are not found in the preset voiceprint database, meaning the broadcaster has no prior broadcasting history using virtual resources, they meet the recruitment criteria for broadcasters using virtual resources. Therefore, virtual resource live streaming permissions can be granted to the broadcaster, allowing them to broadcast using virtual resources. Thus, the broadcaster's voiceprint characteristics can be added to the preset voiceprint database to update it.
[0103] In one implementation, the steps for adding the broadcaster's voiceprint features to the preset voiceprint library can be referenced in S280 as follows:
[0104] S281: Obtain the index code by performing feature encoding on the anchor's voiceprint features.
[0105] S282: Add the anchor's voiceprint features and their index encoding to the preset voiceprint library.
[0106] Index coding can include encodings used to identify and distinguish the broadcaster's voiceprint features. For example, the index coding can be a unique code. Specifically, after extracting the broadcaster's voiceprint features, feature coding can be performed on those features to obtain the index coding.
[0107] In the process of adding the anchor's voiceprint features to the preset voiceprint library, the anchor's voiceprint features can be feature-encoded to obtain index codes, and then the anchor's voiceprint features and their index codes can be associated and added to the preset voiceprint library.
[0108] Furthermore, during the storage of the anchor's voiceprint features, the anchor's identity identification identifier and index code can be stored together. This allows the corresponding anchor or anchor's voiceprint features to be determined based on the anchor's identity identification identifier or index code, thereby improving the efficiency of managing voiceprint features in the preset voiceprint library.
[0109] In one implementation, after determining that the broadcaster's voiceprint features exist in a preset voiceprint database, the following steps can be performed:
[0110] S290: If it is determined that the anchor's voiceprint features exist in the preset voiceprint library, then the anchor's voiceprint features will not be added to the preset voiceprint library.
[0111] If it is determined that the broadcaster's voiceprint features exist in the preset voiceprint database, that is, if the broadcaster's voiceprint features correspond to the same broadcaster identity identifier as the voiceprint features in the preset voiceprint database, then the broadcaster's voiceprint features do not need to be added to the preset voiceprint database. However, it is still possible to grant the broadcaster virtual resource live streaming permissions and send virtual resources to the broadcaster terminal 20, so that the broadcaster can use the virtual resources to conduct live streaming on the broadcaster terminal 20.
[0112] For example, during virtual resource gameplay, the same guild might assign the same UID to different streamers. If streamer B obtains a UID that streamer A has used, and streamer A uses this UID to configure virtual resources (meaning streamer A's voiceprint features are already stored in the preset voiceprint database), then when streamer B uses this UID to match the corresponding UID in the preset voiceprint database, they will match streamer A's voiceprint features. In this case, it can be determined that the streamer's voiceprint features exist in the preset voiceprint database, and streamer B's voiceprint features do not need to be added to the database. However, it is still possible to grant streamer B virtual resource live streaming permissions and distribute virtual resources to streamer B's corresponding streamer terminal 20.
[0113] In one implementation, after performing feature matching between the broadcaster's voiceprint features and those in a preset voiceprint database to determine whether the broadcaster's voiceprint features exist in the preset voiceprint database, the following steps can also be performed:
[0114] S300: If it is determined that the anchor's voiceprint features do not exist in the preset voiceprint database, then virtual resources are sent to the anchor terminal according to the anchor's anchor identity identification mark, so that the anchor terminal can respond to the anchor's selection operation of virtual resources and generate a live video that combines virtual resources and anchor features.
[0115] The virtual resources can be pre-configured resources in server 10 for configuring virtual avatars. By sending the virtual resources to the broadcaster terminal 20, the broadcaster can select from the virtual resources, thereby generating a live video on the broadcaster terminal 20 that combines the virtual resources and the broadcaster's characteristics.
[0116] If it is determined that the broadcaster's voiceprint features do not exist in the preset voiceprint database, virtual resources can be sent to the broadcaster terminal 20 based on the broadcaster's identity recognition features. After receiving the virtual resources, the broadcaster terminal can respond to the broadcaster's selection operation of the virtual resources and generate a live video that combines the virtual resources and the broadcaster's features.
[0117] For example, such as Figure 7 and Figure 8As shown, after the virtual resources are sent to the broadcaster terminal 20, the broadcaster terminal 20 can display a virtual avatar configuration interface based on the virtual resources. The broadcaster can select the background, expression, actions and other configurations of the virtual avatar. The broadcaster's features can be the broadcaster's voiceprint features. Thus, the broadcaster terminal 20 can generate a live video that combines the virtual avatar and the broadcaster's voiceprint features, and send the live video to each viewer terminal 30 through the server 10 so that the viewer can watch it on the viewer terminal 30.
[0118] like Figure 9 As shown in the embodiments of the computer device described in this application, the computer device 100 can be the server 10 described above. The computer device 100 may include a processor 110, a memory 120, and a communication circuit 130.
[0119] The memory 120 is used to store computer programs and may be ROM (Read-Only Memory), RAM (Random Access Memory), or other types of storage devices. Specifically, the memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory is used to store at least one line of program code.
[0120] Processor 110 is used to control the operation of computer device 100. Processor 110 may also be referred to as CPU (Central Processing Unit). Processor 110 may be an integrated circuit chip with signal processing capabilities. Processor 110 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), off-the-shelf programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor may be a microprocessor, or processor 110 may be any conventional processor.
[0121] The processor 110 is used to execute the computer program stored in the memory 120 to implement the live-stream-based virtual resource configuration method described in the first embodiment and the second embodiment of the live-stream-based virtual resource configuration method of this application.
[0122] The computer device 100 may also include a communication circuit 130, which is a communication connection device or circuit used by the computer device 100 to communicate with external devices, so that the processor 110 can interact with external devices via the communication circuit 130.
[0123] For a detailed description of the functions and execution processes of each functional module or component in the computer device embodiments of this application, please refer to the descriptions in the first embodiment and the second embodiment of the live streaming-based virtual resource configuration method of this application, which will not be repeated here.
[0124] In the several embodiments provided in this application, it should be understood that the disclosed computer device 100 and the live-stream-based virtual resource configuration method can be implemented in other ways. For example, the embodiments of the computer device 100 described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0126] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0127] See Figure 10 If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in computer-readable storage medium 200. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions / computer programs to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks, as well as electronic terminals such as computers, mobile phones, laptops, tablets, and cameras that have the aforementioned storage media.
[0128] The description of the execution process of program data in computer-readable storage media can be found in the first embodiment and the second embodiment of the live-stream-based virtual resource configuration method of this application, and will not be repeated here.
[0129] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for configuring virtual resources based on live streaming, characterized in that, include: The broadcaster's terminal uploads audio files to the server; The server extracts the broadcaster's voiceprint features from the audio file; The server performs feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint database to determine whether the broadcaster's voiceprint features exist in the preset voiceprint database. If it is determined that the anchor's voiceprint features do not exist in the preset voiceprint database, the server sends virtual resources to the anchor's terminal based on the anchor's anchor identity identification identifier. The broadcaster terminal responds to the broadcaster's selection operation of the virtual resource and generates a live video that combines the virtual resource and the broadcaster's characteristics. The server extracts the voiceprint features of the broadcaster from the audio file uploaded by the broadcaster terminal, including: performing noise reduction processing on the audio file and extracting the speech segments carrying the broadcaster's voice from the audio file, and extracting the voiceprint features of the broadcaster from the speech segments; The step of performing feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library includes: determining whether the broadcaster's voiceprint features match the voiceprint features in the preset voiceprint library. If they do not match, then determine whether the anchor identity identification identifier corresponding to the anchor voiceprint feature is the same as the anchor identity identification identifier corresponding to the voiceprint feature in the preset voiceprint library. If they are different, it is determined that the anchor's voiceprint feature does not exist in the preset voiceprint library; Extracting the broadcaster's voiceprint features from the audio segment includes: The audio segment is divided into several audio blocks; Receive the determination result of whether each of the audio blocks is the voice of the broadcaster; If the determination result is yes, then the voiceprint features of the speech block are extracted.
2. A method for configuring virtual resources based on live streaming, characterized in that, include: Extract the voiceprint features of the broadcaster from the audio files uploaded by the broadcaster's terminal; The anchor's voiceprint features are matched with voiceprint features in a preset voiceprint database to determine whether the anchor's voiceprint features exist in the preset voiceprint database. If it is determined that the anchor's voiceprint features do not exist in the preset voiceprint database, then virtual resources are sent to the anchor terminal according to the anchor's anchor identity identification mark, so that the anchor terminal can respond to the anchor's selection operation of the virtual resources and generate a live video that combines the virtual resources and the anchor features. The step of extracting the voiceprint features of the broadcaster from the audio file uploaded by the broadcaster terminal includes: performing noise reduction processing on the audio file and extracting the speech segments carrying the broadcaster's voice from the audio file, and extracting the voiceprint features of the broadcaster from the speech segments; The step of performing feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library includes: determining whether the broadcaster's voiceprint features match the voiceprint features in the preset voiceprint library. If they do not match, then determine whether the anchor identity identification identifier corresponding to the anchor voiceprint feature is the same as the anchor identity identification identifier corresponding to the voiceprint feature in the preset voiceprint library. If they are different, it is determined that the anchor's voiceprint feature does not exist in the preset voiceprint library; Extracting the broadcaster's voiceprint features from the audio segment includes: The audio segment is divided into several audio blocks; Receive the determination result of whether each of the audio blocks is the voice of the broadcaster; If the determination result is yes, then the voiceprint features of the speech block are extracted.
3. The method according to claim 2, characterized in that: The step of extracting the voiceprint features of the broadcaster from the audio file uploaded by the broadcaster's terminal includes: performing feature extraction processing on the temporal signal of the audio file through a preset deep learning model to obtain a feature vector as the voiceprint features of the broadcaster.
4. The method according to claim 3, characterized in that: The preset deep learning model includes a first convolutional layer, a first residual neural network, a second residual neural network, a third residual neural network, a second convolutional layer, a max pooling layer, and a fully connected layer connected in sequence. The step of extracting features from the temporal signal of the audio file using a preset deep learning model to obtain a feature vector as the voiceprint feature of the broadcaster includes: inputting the audio temporal signal of the audio file into a first convolutional layer; The convolution result output from the first convolutional layer is input into the first residual neural network; The processing results of the first residual neural network are respectively input into the second residual neural network and the second convolutional layer; The processing results of the second residual neural network are respectively input into the third residual neural network and the second convolutional layer; The convolution result of the second convolutional layer is input into the max pooling layer; The pooling result of the max pooling layer is input into the fully connected layer, and the feature vector is output through the fully connected layer.
5. The method according to claim 2, characterized in that, Before extracting the voiceprint features of the broadcaster from the audio file uploaded by the broadcaster's terminal, the method includes: obtaining the playback link generated by the broadcaster at the end of the trial broadcast and used to link to the audio file; The audio file is obtained through the playback link.
6. The method according to claim 5, characterized in that: After performing feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library, the method includes: if it is determined that the broadcaster's voiceprint features do not exist in the preset voiceprint library, then adding the broadcaster's voiceprint features to the preset voiceprint library to update the preset voiceprint library.
7. The method according to claim 6, characterized in that, The step of adding the anchor's voiceprint features to the preset voiceprint library includes: performing feature encoding on the anchor's voiceprint features to obtain index encoding; The anchor's voiceprint features and their index encoding are associated and added to the preset voiceprint library.
8. The method according to claim 5, characterized in that, After determining whether the anchor identity identifier corresponding to the anchor voiceprint feature is the same as the anchor identity identifier corresponding to the voiceprint feature in the preset voiceprint library, the method includes: if they are the same, then determining that the anchor voiceprint feature exists in the preset voiceprint library; After performing feature matching between the broadcaster's voiceprint features and voiceprint features in a preset voiceprint library to determine whether the broadcaster's voiceprint features exist in the preset voiceprint library, the method includes: if it is determined that the broadcaster's voiceprint features exist in the preset voiceprint library, then the broadcaster's voiceprint features are not added to the preset voiceprint library.
9. The method according to claim 5, characterized in that, After determining whether the anchor's voiceprint features match the voiceprint features in the preset voiceprint database, the process includes: if they do not match, determining whether the anchor's voiceprint features have a corresponding anchor identity identification identifier. If it is determined that there is no corresponding anchor identity identification identifier, then an anchor identity identification identifier is assigned to the anchor, and the anchor identity identification identifier is bound to the anchor voiceprint feature. Then, the step of determining whether the anchor identity identification identifier corresponding to the anchor voiceprint feature is the same as the anchor identity identification identifier corresponding to the voiceprint feature in the preset voiceprint library is executed.
10. The method according to claim 2, characterized in that, The step of denoising the audio file and extracting the speech segment carrying the anchor's voice from the audio file includes: removing the background sound portion from the audio file and denoising the audio file to obtain the foreground sound portion; The non-speech segments that do not carry the anchor's voice are removed from the foreground audio portion to obtain the speech segment.
11. A computer device, characterized in that, Includes processor, memory, and communication circuitry; The communication circuit and the memory are coupled to the processor; The memory stores a computer program, and the processor executes the computer program to implement the live-stream-based virtual resource configuration method as described in any one of claims 1-10.
12. A computer-readable storage medium, characterized in that, The system contains a computer program that is executed by a processor to implement the live-stream-based virtual resource configuration method as described in any one of claims 1-10.