Voice packet sending method and device, voice chat system, electronic equipment and medium
Through real-time data analysis of the chat room, the voice package recommendation queue and custom creation function are generated, which solves the problem of limited number of voice packages in the chat room, improves user interaction effect and creative freedom, and realizes the interactive experience of intelligent recommendation and multimodal fusion.
Patent Information
- Application Number
- CN202510839971.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-08-15
AI Technical Summary
In the voice chat function of the existing chat room, the number of voice packets is limited and the style is fixed, which leads to a decrease in user freshness, making it difficult to determine the most suitable voice packet, affecting the chat interaction effect.
By conducting scenario analysis of the real-time chat data in the chat room, a user-specific voice package recommendation queue is generated, and a custom voice package creation function is provided. Combined with the reinforcement learning model optimization recommendation algorithm, the voice package content of other users is analyzed in real time for multimodal fusion recommendation.
It improves the user's voice package usage rate, increases the chat interaction effect, provides freedom of independent creation, and promotes the platform's content creation efficiency and interaction.
Smart Images

Figure CN120499141A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network live broadcast technology, and in particular to a method and device for sending a voice packet, a voice chat system, an electronic device, and a computer-readable storage medium. Background Art
[0002] Chat room is a very popular Internet chat tool that provides real-time chat function. In the chat room, users can chat and interact with other online users through real-time text, image and voice communication over the Internet.
[0003] Currently, when chat rooms provide voice chats, they generally provide configured pre-made voice packages. Users send voice packages in the chat room to make the chat interaction interesting. However, due to the limited number of voice packages and fixed styles, it is easy to cause users to lose their sense of freshness. As the chat deepens, it is difficult for users to determine the most appropriate voice package to send, which affects the chat interaction effect. Summary of the Invention
[0004] Based on this, it is necessary to provide a voice packet sending method, device, voice chat system, electronic device and computer-readable storage medium to enhance the chat interaction effect.
[0005] A method for sending a voice packet, comprising:
[0006] When a user enters a chat room, context analysis is performed on the real-time chat data in the chat room to obtain context-aware information;
[0007] Generating a voice package recommendation queue based on the context awareness information and the user's voice package resource library; wherein the voice package resource library stores voice packages exclusive to the user;
[0008] A first voice package selected by a user from a voice package recommendation queue is obtained, and the first voice package is sent to the chat room.
[0009] In one embodiment, the voice packet sending method further includes:
[0010] In response to the user entering the voice creation system page, requesting and loading the voice package editing page from the server;
[0011] Create a customized voice pack on the voice pack editing page;
[0012] Update the created exclusive voice pack to the server's voice pack resource library and download the audition link;
[0013] The number of voice packages in the voice package resource library is updated on the panel of the chat room, and a download and trial listening link is displayed for the user to audition.
[0014] In one embodiment, creating a customized voice package on the voice package editing page includes:
[0015] Selecting a voice package to be edited, and performing corpus mixing and timbre reorganization on the voice package to be edited;
[0016] Deconstruct the newly added text entered by the user into independent semantic units and dynamically render them in the editable area based on the part-of-speech tagging results;
[0017] After selecting an insertion position in the editable area, the timestamp of the insertion position and the newly added text are recorded, and the newly added text is converted into a fundamental frequency feature aligned with the original voice packet spectrum;
[0018] The fundamental frequency characteristics and the time domain stretching rate are dynamically adjusted, audio data of a target waveform is generated by a vocoder, and naturalness optimization is performed to obtain a customized voice package.
[0019] In one embodiment, performing context analysis on the real-time chat data in the chat room to obtain context awareness information includes:
[0020] Analyze the user's historical voice packet usage preference information and social style information to obtain user profile information, analyze the chat room theme and the number of online users to obtain scene feature information, and perform emotion recognition on the tone and speed of the chat room voice information to obtain emotional atmosphere information;
[0021] Situational perception information is obtained by performing multi-dimensional fusion based on the user portrait information, scene feature information and emotional atmosphere information.
[0022] In one embodiment, generating a voice package recommendation queue based on the context awareness information and the user's voice package resource library includes:
[0023] A recommendation algorithm model is used to calculate a recommendation score for each voice package in the user's voice package resource library according to the context awareness information, and a voice package recommendation queue is generated according to the recommendation score.
[0024] In one embodiment, the voice packet sending method further includes:
[0025] Periodically using a reinforcement learning model to dynamically optimize the recommendation algorithm model;
[0026] Adjusting the recommendation weight of the recommendation algorithm model according to the likes data or dislikes data;
[0027] The matching algorithm of the recommendation algorithm model is optimized according to the playback completion rate and the secondary usage rate.
[0028] In one embodiment, the voice packet sending method further includes: obtaining a second voice packet sent by other users in the chat room, and performing multimodal fusion based on the second voice packet to generate a matching quick reply voice packet.
[0029] In one embodiment, performing multimodal fusion based on the second voice package to generate a matching quick reply voice package includes:
[0030] Extracting core keywords from the second voice packet and analyzing the emotional tendency corresponding to the second voice packet through natural language processing technology, and constructing a semantic-emotional association model based on this;
[0031] Based on word vectors, the system calculates voice packets with semantic relevance greater than a set threshold, identifies matching content with emotional relevance greater than a set threshold, and filters scenes based on the current social scene characteristics to obtain multimodal matching results.
[0032] A quick reply voice package that matches the second voice package is obtained based on the semantic-emotional association model and the multimodal matching results, and is presented in the user chat box.
[0033] A voice packet sending device, comprising:
[0034] A context awareness module, configured to perform context analysis on real-time chat data in a chat room to obtain context awareness information when a user enters the chat room;
[0035] A voice recommendation module, configured to generate a voice package recommendation queue based on the context awareness information and the user's voice package resource library; wherein the voice package resource library stores voice packages exclusive to the user;
[0036] The voice sending module is used to obtain a first voice package selected by a user from a voice package recommendation queue and send the first voice package to the chat room.
[0037] A voice chat system comprises: a client and a chat server; wherein the client is connected to the chat server via a communication network;
[0038] The client is used for each user to access the chat room;
[0039] The chat server is used to send the voice packet in the chat room using the voice packet sending method.
[0040] An electronic device, comprising:
[0041] one or more processors;
[0042] Memory;
[0043] One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the steps of the voice packet sending method.
[0044] A computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded by the processor and executes the steps of the voice packet sending method.
[0045] The technical solution of the present application is to perform scenario analysis on the real-time chat data of the chat room to obtain scenario perception information when a user enters the chat room; generate a voice package recommendation queue based on the scenario perception information and the user's voice package resource library; the user selects a voice package from the voice package recommendation queue and sends it to the chat room; this technical solution can make intelligent scenario-based recommendations based on the user's actual usage data and combined with the scenario perception information, automatically recommend the most suitable voice package, increase the user's voice package usage rate, and improve the user's voice chat interaction effect in the chat room.
[0046] Furthermore, when other users send voice packets, the voice packet content is analyzed in real time based on AI, and multimodal fusion is performed using the second voice packet sent by other users in the chat room to obtain the most matching quick reply voice packet, helping users participate in the interaction more naturally.
[0047] Furthermore, it provides users with an independent creation system, allowing them to perform AI secondary creation based on the existing voice package timbre and corpus, thereby promoting the efficiency of platform content creation and interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a diagram of an example voice chat service application scenario;
[0049] Figure 2 is a flow chart of a method for sending a voice packet according to an embodiment;
[0050] Figure 3 This is an example flow chart of a user sending a voice packet during a chat;
[0051] Figure 4 This is a flow chart of a user independently creating a voice package according to an embodiment;
[0052] Figure 5 This is a diagram of a user-generated voice package panel;
[0053] Figure 6 This is a diagram of the voice panel of an example chat room;
[0054] Figure 7 is a structural diagram of a voice packet sending device according to an embodiment;
[0055] Figure 8 is a block diagram of an example electronic device. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0057] In the embodiments of the present application, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same effects and functions, "at least one" means one or more, and "a plurality of" means two or more, for example, a plurality of objects refers to two or more objects. Words such as "include" or "comprise" mean that the information appearing before "include" or "comprises" covers the information listed after "include" or "comprises" and its equivalents, and does not exclude other information. The "and / or" mentioned in the embodiments of the present application indicates that three relationships may exist, and the character " / " generally indicates that the objects before and after are in an "or" relationship.
[0058] The technical solution provided in the embodiments of this application can be applied to Figure 1 In the application scenario of the related method of this application shown in FIG. Figure 1 This is a diagram of an example voice chat application scenario. The voice chat system may include a chat server and clients. Each client communicates with the chat server via a communications network, enabling real-time chat interaction. The clients may include, but are not limited to, various personal computers, laptops, smartphones, and tablets. The chat server may be implemented as a standalone server or a server cluster consisting of multiple servers.
[0059] refer to Figure 2 As shown, Figure 2 The flowchart of the method for sending a voice packet according to an embodiment includes the following steps:
[0060] S10: When a user enters a chat room, context analysis is performed on the real-time chat data in the chat room to obtain context awareness information.
[0061] Specifically, when a user enters a chat room, the server can perform a scenario analysis on various chat data in the chat room to obtain scenario awareness information, thereby providing a reference for selecting a voice package.
[0062] refer to Figure 3 As shown, Figure 3This is an example flow chart of a user sending a voice packet during a chat, which provides several preferred embodiment processes of the voice packet sending method of the present application.
[0063] In one embodiment, step S10 of performing context analysis on the real-time chat data in the chat room to obtain context awareness information includes:
[0064] S101: Analyze the user's historical voice package usage preference information and social style information to obtain user portrait information, analyze the chat room theme and the number of online users to obtain scene feature information, and perform emotion recognition on the tone and speaking speed of the chat room voice information to obtain emotional atmosphere information.
[0065] S102, performing multi-dimensional fusion of the user portrait information, scene feature information, and emotional atmosphere information to obtain situational awareness information.
[0066] For example, after a user enters a chat room, three-dimensional real-time analysis is initiated to analyze historical voice package usage preferences and social styles to obtain a user portrait; the room theme and the number of online users are analyzed to obtain scene feature information; and the emotional atmosphere is obtained through voice emotion recognition (intonation / speed).
[0067] S20 , generating a voice package recommendation queue according to the context awareness information and a voice package resource library of the user; wherein the voice package resource library stores voice packages exclusive to the user.
[0068] In this step, the server can intelligently generate a voice package recommendation queue from the user's voice package resource library according to the context awareness information; wherein the voice package resource library stores the user-specific voice packages.
[0069] In one embodiment, Figure 3 As shown, step S20 generates a voice package recommendation queue based on the context awareness information and the user's voice package resource library, including:
[0070] A recommendation algorithm model is used to calculate a recommendation score for each voice package in the user's voice package resource library according to the context awareness information, and a voice package recommendation queue is generated according to the recommendation score.
[0071] Specifically, the user portrait information, scene feature information, and emotional atmosphere information in the above-mentioned situational awareness information can be input into a recommendation algorithm to form a voice package recommendation queue, and the voice package queue can be recommended to the user.
[0072] In one embodiment, further, in order to improve the performance of the recommendation algorithm, the method of the present application further includes:
[0073] The recommendation algorithm model is dynamically optimized using a reinforcement learning model periodically; the recommendation weight of the recommendation algorithm model is adjusted according to the like data or dislike data; and the matching algorithm of the recommendation algorithm model is optimized according to the playback completion rate and the secondary usage rate.
[0074] Exemplarily, the recommendation algorithm uses a reinforcement learning model for dynamic optimization. On the one hand, explicit feedback (active likes or dislikes, etc.) directly adjusts the recommendation weight, and implicit feedback (playback completion rate / secondary usage rate) optimizes the matching algorithm through offline analysis. Incremental training is performed every 5 minutes, and at the same time, it is monitored whether there is user usage feedback. If so, it can be dynamically optimized based on user feedback, thereby dynamically improving the recommendation accuracy.
[0075] S30: Obtain a first voice package selected by the user from a voice package recommendation queue, and send the first voice package to the chat room.
[0076] In this step, the user can select a target voice package from the voice package recommendation queue and send it to the chat room according to the chat needs, so that the most suitable voice package can be automatically recommended based on the user's actual usage data and combined with the current chat topic, atmosphere and user preferences, thereby increasing the user's voice package usage rate.
[0077] In order to make the technical solution of this application clearer, more embodiments are described below.
[0078] In one embodiment, in order to continue to provide users with intelligent recommended voice package references after sending the voice package, the voice package sending method of the present application, step S30 may also include: obtaining a second voice package sent by other users in the chat room, and performing multimodal fusion based on the second voice package to generate a matching quick reply voice package.
[0079] Specifically, when other users send voice packets, the voice packet content is analyzed in real time based on AI, and the second voice packet sent by other users in the chat room is used for multimodal fusion to obtain the most matching quick reply voice packet.
[0080] The solution of the above embodiment uses AI to analyze the content of the voice package in real time to obtain the most matching quick reply voice package, which can help users participate in the interaction more naturally.
[0081] In one embodiment, Figure 3 As shown, step S30 of performing multimodal fusion according to the second voice package to generate a matching quick reply voice package includes:
[0082] S301, extracting core keywords from the second voice package and analyzing the emotional tendency corresponding to the second voice package through natural language processing technology, and constructing a semantic-emotional association model based on this.
[0083] S302, based on word vectors, calculates voice packets whose semantic relevance is greater than a set threshold, identifies adaptive content whose emotional consistency is greater than a set threshold, and filters scenes based on the characteristics of the current social scene to obtain a multimodal matching result.
[0084] S303: Obtain a quick reply voice package that matches the second voice package based on the semantic-emotion association model and the multimodal matching result, and present it on the user chat box.
[0085] For example, in a gaming scenario, when a voice packet is received from another user, a dual-channel processing of semantic parsing and multimodal matching is initiated. The semantic parsing channel uses NLP (Natural Language Processing) technology to extract core keywords in the voice content, such as "tactics" and "celebration," and analyzes the emotional tendencies in the voice, such as excitement and seriousness, to build a semantic-emotional association model. A triple screening of multimodal matching is performed simultaneously; for example, based on word vectors, voice packets with a semantic relevance of ≥70% are calculated, and matching content with an emotional fit of ≥65% is identified. Combined with the current social scene characteristics, such as filtering out violations in game battles, the dual-channel processing fusion algorithm determines the final quick reply voice packet, which is then displayed above the user's chat box, allowing the user to quickly send it by clicking on it.
[0086] As in the above-mentioned embodiment, real-time intelligent matching of quick reply voice packages is realized. When other users send voice packages, AI analyzes the content semantics and emotions, and intelligently recommends quick reply voice packages that are consistent with the context, emotions, and etiquette, thereby helping users participate in interactions more naturally.
[0087] In order to enable users to create their own voice packages, this application also provides a technical solution for users to create customized voice packages.
[0088] In one embodiment, reference Figure 4 As shown, Figure 4 This is a flowchart of a user independently creating a voice package according to an embodiment. The voice package sending method of the present application may further include:
[0089] (a) In response to the user entering the voice creation system page, request and load the voice package editing page from the server.
[0090] Specifically, such as Figure 5 As shown, Figure 5This is an example of a user-created voice package panel diagram. The user enters the voice creation system page, and the client requests creation permission and initialization parameter data from the server. The server returns the creation mode options, available voice material library, and user historical work data. The client renders the creation interface and loads the timbre adjustment control, corpus mixing matrix, and timbre upload component.
[0091] (b) Create a customized voice package on the voice package editing page.
[0092] For example, this embodiment provides the following two creation modes:
[0093] First and second creation mode
[0094] In this mode, users perform AI secondary creation based on the existing voice package timbre and corpus. The main process includes:
[0095] ① Select a voice package to be edited, and perform corpus mixing and timbre reorganization on the voice package to be edited.
[0096] ② Deconstruct the newly added text entered by the user into independent semantic units and dynamically render them in the editable area based on the part-of-speech tagging results.
[0097] ③ After selecting the insertion position in the editable area, record the timestamp of the insertion position and the newly added text, and convert the newly added text into a fundamental frequency feature aligned with the original voice packet spectrum.
[0098] ④ Dynamically adjust the fundamental frequency characteristics and time domain stretching rate, generate audio data of the target waveform through a vocoder, and optimize the naturalness to obtain a customized voice package.
[0099] For example, the user selects an existing voice package to enter the editing page. The user can mix the corpus and reorganize the timbre of the voice package. The server deconstructs the original text into independent semantic units based on the CRF (Conditional Random Field) word segmentation model. The client dynamically renders the editable area as a light gray box with a "+" mark according to the part-of-speech tagging results. After the user clicks the "+" sign in the interface to select the insertion position, the timestamp and newly added text of the position are recorded. The newly added content is converted into fundamental frequency features through the TTS (Text To Speech) engine and aligned with the original voice package spectrum. When the user adjusts the pitch / speech rate (±30% corresponds to the slider scale), the PSOLA (Pitch Synchronous Overlap Add) algorithm is used to dynamically adjust the fundamental frequency and time domain expansion rate, and then the target waveform is generated by the WaveGlow vocoder. The edited audio segment is optimized for naturalness through the LSTM-PPG model, and a trial audio is output for the user to audition and confirm. After the user confirms, the voice package content is uploaded to the server for review. After the review is passed, a secondary creation voice package is generated based on the timbre of the voice package;
[0100] 2. Customize the sound:
[0101] For example, users can upload a personal voice recording of more than 20 seconds. After the server automatically detects the sound quality, it generates a dedicated voice package within 5 minutes through intelligent modeling.
[0102] (c) Update the created exclusive voice package to the voice package resource library on the server and download the trial listening link.
[0103] Specifically, when the server completes the creation process, it digitally signs the generated work and stores it in the blockchain evidence library, updates the voice package holding data in the user's voice package resource library, and sends a creation completion notification and a work audition link to the client.
[0104] (d) updating the number of voice packages in the voice package resource library on the panel of the chat room and displaying a download link for the user to audition.
[0105] Specifically, after receiving the notification from the server, the client updates the number of voice packages in the user asset panel, records the operation timestamp and parameter snapshot in the creation log, and the client can send customized voice packages in the chat channel.
[0106] As in the above-mentioned embodiment, by providing users with the function of independent creation, users can perform AI secondary creation based on the existing voice package timbre and corpus, freely adjust the tone, speaking speed or corpus mixing, and open a user-defined timbre upload channel. By collecting the user's voiceprint sample, a personalized voice package can be generated, which promotes the efficiency and interaction of platform content creation.
[0107] Based on the embodiments of the audio packet sending method in the above embodiments, an application example based on the audio packet sending method is described below.
[0108] (1) After the user opens the voice package panel in the chat room, the user enters the voice creation system page, where the user can choose secondary creation and custom uploaded sound. If the user selects the secondary upload button, the user selects the existing voice package and enters the editing page.
[0109] (2) When the user is on the editing page, the client automatically marks the editable area. After the user selects the text insertion area, he can customize the input content to complete the corpus mixing.
[0110] (3) After mixing the corpus, the user can click Modify Parameters to enter the timbre reorganization panel.
[0111] In this editing panel, you can adjust controls in three dimensions, such as pitch (vertical slider ±50Hz), speech rate (horizontal drag bar ±40%), and emotional parameters (circular coordinate selector to locate the "lively / calm" quadrant) in real time. Each adjustment can be compared and auditioned instantly.
[0112] After the user confirms the audition, click Save. The self-created voice result is saved to the user's voice package library. The user can send the voice package in the chat channel.
[0113] (4) Users can click on the intelligently recommended quick reply voice package to interact. If they click on the trial listening button, they can try it out. If they click on the text part, it will be quickly sent to the current chat room.
[0114] (5) When another user sends a voice packet, the server automatically analyzes the voice packet according to the fusion algorithm and recommends the most matching quick reply voice packet to the current user, such as Figure 6 As shown, Figure 6 This is a sample chat room voice panel diagram. The current user can click on the voice package to listen to it and send it quickly.
[0115] As in the example above, users can remix existing voice packs and re-arrange their voices, eliminating the rigidity of pre-made voice packs and increasing creative freedom and content diversity. Furthermore, a recommendation algorithm based on reinforcement learning, combined with analysis from various perspectives, can recommend quick voice packs and quick reply packs that match the chat topic. This creates a social environment where intelligent recommendations and precise matching are prioritized, lowering the barrier to interaction and improving social fluency.
[0116] An embodiment of a voice packet sending device is described below.
[0117] like Figure 7 As shown, Figure 7 This is a schematic structural diagram of a voice packet sending device according to an embodiment, comprising:
[0118] The context awareness module 10 is used to perform context analysis on the real-time chat data in the chat room to obtain context awareness information when a user enters the chat room.
[0119] The voice recommendation module 20 is configured to generate a voice package recommendation queue based on the context awareness information and the user's voice package resource library; wherein the voice package resource library stores user-specific voice packages.
[0120] The voice sending module 30 is configured to obtain a first voice package selected by a user from a voice package recommendation queue and send the first voice package to the chat room.
[0121] In one embodiment, for the context awareness module 10, when a user enters a chat room, the server can perform context analysis on various chat data in the chat room to obtain context awareness information, thereby providing a reference for selecting a voice package.
[0122] As an embodiment, the context awareness module 10 may further perform the following functions:
[0123] The user's historical voice package usage preference information and social style information are analyzed to obtain user portrait information, the chat room theme and the number of online users are analyzed to obtain scene feature information, and the tone and speed of the chat room voice information are used for emotion recognition to obtain emotional atmosphere information.
[0124] Situational perception information is obtained by performing multi-dimensional fusion based on the user portrait information, scene feature information and emotional atmosphere information.
[0125] For example, after a user enters a chat room, three-dimensional real-time analysis is initiated to analyze historical voice package usage preferences and social styles to obtain a user portrait; the room theme and the number of online users are analyzed to obtain scene feature information; and the emotional atmosphere is obtained through voice emotion recognition (intonation / speed).
[0126] In one embodiment, for the voice recommendation module 20, the server can intelligently generate a voice package recommendation queue from the user's voice package resource library according to context awareness information; wherein the voice package resource library stores the user-specific voice packages.
[0127] As an example, the voice recommendation module 20 may further perform the following functions:
[0128] A recommendation algorithm model is used to calculate a recommendation score for each voice package in the user's voice package resource library according to the context awareness information, and a voice package recommendation queue is generated according to the recommendation score.
[0129] Specifically, the user portrait information, scene feature information, and emotional atmosphere information in the above-mentioned situational awareness information can be input into a recommendation algorithm to form a voice package recommendation queue, and the voice package queue can be recommended to the user.
[0130] Furthermore, in order to improve the performance of the recommendation algorithm, the voice recommendation module 20 can also periodically use a reinforcement learning model to dynamically optimize the recommendation algorithm model; adjust the recommendation weight of the recommendation algorithm model according to the likes data or dislikes data; and optimize the matching algorithm of the recommendation algorithm model according to the playback completion rate and secondary usage rate.
[0131] Exemplarily, the recommendation algorithm uses a reinforcement learning model for dynamic optimization. On the one hand, explicit feedback (active likes or dislikes, etc.) directly adjusts the recommendation weight, and implicit feedback (playback completion rate / secondary usage rate) optimizes the matching algorithm through offline analysis. Incremental training is performed every 5 minutes, and at the same time, it is monitored whether there is user usage feedback. If so, it can be dynamically optimized based on user feedback, thereby dynamically improving the recommendation accuracy.
[0132] For the voice sending module 30, the user can select a target voice package from the voice package recommendation queue and send it to the chat room according to the chat needs, so that the most suitable voice package can be automatically recommended based on the user's actual usage data and combined with the current chat topic, atmosphere and user preferences, thereby increasing the user's voice package utilization rate.
[0133] Furthermore, the voice sending module 30 may also obtain a second voice packet sent by other users in the chat room, and perform multimodal fusion based on the second voice packet to generate a matching quick reply voice packet.
[0134] Specifically, when other users send voice packets, the voice packet content is analyzed in real time based on AI, and the second voice packet sent by other users in the chat room is used for multimodal fusion to obtain the most matching quick reply voice packet.
[0135] As an embodiment, the voice sending module 30 can further extract the core keywords in the second voice package and analyze the emotional tendency corresponding to the second voice package through natural language processing technology, and thereby construct a semantic-emotional association model; calculate the voice packages with semantic correlation greater than the set threshold based on the word vector, identify the adaptation content with emotional consistency greater than the set threshold, and filter the scenes in combination with the current social scene characteristics to obtain multimodal matching results; obtain a quick reply voice package that matches the second voice package according to the semantic-emotional association model and the multimodal matching results, and present it on the user chat box.
[0136] The voice packet sending device of this embodiment can execute a voice packet sending method provided in the embodiment of the present application. The implementation principle is similar. The actions performed by each module in the voice packet sending device in each embodiment of the present application correspond to the steps in the voice packet sending method in each embodiment of the present application. For the detailed functional description of each module of the voice packet sending device, please refer to the description of the corresponding voice packet sending method shown in the previous text, and will not be repeated here.
[0137] An embodiment of the voice chat system is described below.
[0138] The voice chat system of this embodiment includes: a client and a chat server; wherein the client is connected to the chat server via a communication network; the client is used to access each user of a chat room; and the chat server is used to send voice packets in the chat room using the voice packet sending method of any of the aforementioned embodiments.
[0139] Exemplarily, the technical solution of the present application can be applied to a live broadcast system, and the chat room can be a chat channel set up in the live broadcast room, wherein the live broadcast system can include an anchor end, a chat server and an audience end, and the audience end and the anchor end can enter the voice chat system through a client installed on an electronic device. Exemplarily, the anchor end and the audience end can be computer devices, such as PDAs, smart phones, tablet computers, desktop computers or laptops, etc., which are not limited to this. They can also be software modules of an application. The chat server can be implemented using an independent server or a server cluster consisting of multiple servers.
[0140] As in the above-mentioned voice chat system, when a user enters a chat room, a scenario analysis is performed on the real-time chat data of the chat room to obtain scenario-aware information; a voice package recommendation queue is generated based on the scenario-aware information and the user's voice package resource library; the user selects a voice package from the voice package recommendation queue and sends it to the chat room; at the same time, a second voice package sent by other users in the chat room is obtained, and multimodal fusion is performed based on the second voice package to generate a matching quick reply voice package; it can make intelligent scenario-based recommendations based on the user's actual usage data and combined with the current chat topic, atmosphere and user preferences, automatically recommend the most suitable voice package, increase the user's voice package usage rate, and after other users send voice packages, obtain the most matching quick reply voice package based on AI real-time analysis of the voice package content, helping users to participate in the interaction more naturally. It also provides users with an independent creation system, allowing users to perform AI secondary creation based on the existing voice package timbre and corpus, promoting platform content creation efficiency and interaction.
[0141] It should be noted that for a detailed functional description of the voice chat system, please refer to the relevant description in the previous embodiment of the voice packet sending method, which will not be repeated here.
[0142] Embodiments of electronic devices and computer-readable storage media are described below.
[0143] The present application provides a technical solution for an electronic device for implementing functions related to a voice packet sending method. The electronic device of this embodiment includes one or more processors, a memory; one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by one or more processors, and the one or more programs are configured for the steps of the voice packet sending method of any embodiment.
[0144] like Figure 8 As shown, Figure 8 1 is a block diagram of an exemplary electronic device; the electronic device may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc. The electronic device 100 may include one or more of the following components: a processing component 102, a memory 104, a power component 106, a multimedia component 108, an audio component 109, an input / output (I / O) interface 112, a sensor component 114, and a communication component 116.
[0145] The processing component 102 generally controls the overall operation of the electronic device 100, such as operations associated with display, phone calls, data communications, camera operations, and recording operations.
[0146] The memory 104 is configured to store various types of data to support operations in the electronic device 100, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0147] The power supply component 106 provides power to various components of the electronic device 100 .
[0148] The multimedia component 109 includes a screen that provides an output interface between the electronic device 100 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). In some embodiments, the multimedia component 108 includes a front camera and / or a rear camera.
[0149] The audio component 109 is configured to output and / or input audio signals.
[0150] I / O interface 112 provides an interface between processing component 102 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a start button, and a lock button.
[0151] The sensor assembly 114 includes one or more sensors for providing various aspects of status assessment for the electronic device 100. The sensor assembly 114 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact.
[0152] The communication component 116 is configured to facilitate wired or wireless communication between the electronic device 100 and other devices. The electronic device 100 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof.
[0153] The present application provides a technical solution for a computer-readable storage medium for implementing functions related to a method for sending a voice packet. The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded by a processor to execute the method for sending a voice packet in any embodiment.
[0154] In an exemplary embodiment, the computer-readable storage medium may be a non-transitory computer-readable storage medium including instructions, such as a memory including instructions. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0155] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for sending a voice packet, characterized in that: include: When a user enters a chat room, context analysis is performed on the real-time chat data in the chat room to obtain context-aware information; Generating a voice package recommendation queue based on the context awareness information and the user's voice package resource library; wherein the voice package resource library stores voice packages exclusive to the user; A first voice package selected by a user from a voice package recommendation queue is obtained, and the first voice package is sent to the chat room.
2. The method for sending a voice packet according to claim 1, wherein: Also includes: In response to the user entering the voice creation system page, requesting and loading the voice package editing page from the server; Create a customized voice pack on the voice pack editing page; Update the created exclusive voice pack to the server's voice pack resource library and download the audition link; The number of voice packages in the voice package resource library is updated on the panel of the chat room, and a download and trial listening link is displayed for the user to audition.
3. The method for sending a voice packet according to claim 2, wherein: Create a custom voice pack on the voice pack editing page, including: Selecting a voice package to be edited, and performing corpus mixing and timbre reorganization on the voice package to be edited; Deconstruct the newly added text entered by the user into independent semantic units and dynamically render them in the editable area based on the part-of-speech tagging results; After selecting an insertion position in the editable area, the timestamp of the insertion position and the newly added text are recorded, and the newly added text is converted into a fundamental frequency feature aligned with the original voice packet spectrum; The fundamental frequency characteristics and the time domain stretching rate are dynamically adjusted, audio data of a target waveform is generated by a vocoder, and naturalness optimization is performed to obtain a customized voice package.
4. The method for sending a voice packet according to claim 1, wherein: Performing situational analysis on the real-time chat data in the chat room to obtain situational awareness information includes: Analyze the user's historical voice packet usage preference information and social style information to obtain user profile information, analyze the chat room theme and the number of online users to obtain scene feature information, and perform emotion recognition on the tone and speed of the chat room voice information to obtain emotional atmosphere information; Situational perception information is obtained by performing multi-dimensional fusion based on the user portrait information, scene feature information and emotional atmosphere information.
5. The method for sending a voice packet according to claim 1, wherein: Generating a voice package recommendation queue according to the context awareness information and the user's voice package resource library includes: A recommendation algorithm model is used to calculate a recommendation score for each voice package in the user's voice package resource library according to the context awareness information, and a voice package recommendation queue is generated according to the recommendation score.
6. The method for sending a voice packet according to claim 5, wherein: Also includes: Periodically using a reinforcement learning model to dynamically optimize the recommendation algorithm model; Adjusting the recommendation weight of the recommendation algorithm model according to the likes data or dislikes data; The matching algorithm of the recommendation algorithm model is optimized according to the playback completion rate and the secondary usage rate.
7. The method for sending a voice packet according to any one of claims 1 to 6, wherein: Also includes: Obtain a second voice packet sent by another user in the chat room, and perform multimodal fusion based on the second voice packet to generate a matching quick reply voice packet.
8. The method for sending a voice packet according to claim 7, wherein: Performing multimodal fusion on the second voice package to generate a matching quick reply voice package includes: Extracting core keywords from the second voice packet and analyzing the emotional tendency corresponding to the second voice packet through natural language processing technology, and constructing a semantic-emotional association model based on this; Based on word vectors, the system calculates voice packets with semantic relevance greater than a set threshold, identifies matching content with emotional relevance greater than a set threshold, and filters scenes based on the current social scene characteristics to obtain multimodal matching results. A quick reply voice package that matches the second voice package is obtained based on the semantic-emotional association model and the multimodal matching results, and is presented in the user chat box.
9. A voice packet sending device, characterized in that: include: A context awareness module, configured to perform context analysis on real-time chat data in a chat room to obtain context awareness information when a user enters the chat room; A voice recommendation module, configured to generate a voice package recommendation queue based on the context awareness information and the user's voice package resource library; wherein the voice package resource library stores voice packages exclusive to the user; The voice sending module is used to obtain a first voice package selected by a user from a voice package recommendation queue and send the first voice package to the chat room.
10. A voice chat system, characterized in that: include: A client and a chat server; wherein the client is connected to the chat server via a communication network; The client is used for each user to access the chat room; The chat server is used to send voice packets in a chat room using the voice packet sending method according to any one of claims 1 to 7.
11. An electronic device, characterized in that: The electronic device comprises: one or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the steps of the voice packet sending method according to any one of claims 1 to 8.
12. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded by the processor and executes the steps of the voice packet sending method according to any one of claims 1 to 8.