Session message processing method and device, electronic equipment, computer readable storage medium and computer program product

By displaying voice controls in the conversation interface and fusing them with the conversation expressions to generate fused voice, the problem of single conversation process is solved, and the diversity of conversation messages and user experience is improved.

CN120455424APending Publication Date: 2025-08-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410175891.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art has a relatively single voice transmission process in the conversation interface and lacks diversity, which makes the experience dull and boring.

Method used

Display voice controls in the conversation interface and perform fusion operations with the conversation expressions to generate fusion speech and improve the diversity of conversation messages.

Benefits of technology

Through the integration of voice controls and conversational expressions, diverse conversation messages are generated to improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455424A_ABST
    Figure CN120455424A_ABST
Patent Text Reader

Abstract

The invention provides a session message processing method and device, electronic equipment, a computer readable storage medium and a computer program product, and the method comprises the steps: displaying a voice control in a session interface, and displaying at least one session expression; and in response to a fusion operation for the voice control and a first session expression in the at least one session expression, displaying a first fusion voice generated based on the first session expression. According to the invention, the diversity of the session messages in the session process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for processing conversation messages. Background Art

[0002] When sending voice based on a conversational interface, related technologies mostly use voice controls to record a voice segment in real time and then send it. The conversation process is relatively simple and the experience is dull and boring. Summary of the Invention

[0003] Embodiments of the present application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing conversation messages, which can improve the diversity of conversation messages during a conversation.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] This embodiment of the present application provides a method for processing a session message, including:

[0006] In the conversation interface, display a voice control and at least one conversation emoticon;

[0007] In response to a fusion operation on the voice control and a first conversational expression among the at least one conversational expression, a first fused voice generated based on the first conversational expression is displayed.

[0008] An embodiment of the present application provides a device for processing a session message, including:

[0009] A display module, configured to display a voice control and at least one conversation emoticon in the conversation interface;

[0010] The fusion module is configured to, in response to a fusion operation on the voice control and a first conversation expression in the at least one conversation expression, display a first fused voice generated based on the first conversation expression.

[0011] In the above scheme, the conversation interface includes a message editing area, and the voice control is displayed in the message editing area; the fusion module is also used to respond to a drag operation on the first conversation expression and drag the first conversation expression to the voice control in the message editing area; in response to a release operation on the first conversation expression, display the first fused voice generated based on the first conversation expression in the message editing area.

[0012] In the above scheme, the device also includes a switching module, which is used to display a switching control in the associated area of the first fused voice and automatically play the first fused voice; wherein the switching control is used to switch the content of the first fused voice; in response to a trigger operation on the switching control, the content of the first fused voice is switched to new content, and the new content is associated with the first conversational expression.

[0013] In the above scheme, the first fused speech has a first speech feature, and the device also includes a second display module, which is used to display a second fused speech with a second speech feature in response to a fusion operation on the first fused speech and a second conversational expression in the at least one conversational expression; wherein the second speech feature is different from the first speech feature, and the second speech feature corresponds to the second conversational expression.

[0014] In the above scheme, the conversation interface includes a message display area and a message editing area, and the first fused voice is displayed in the message editing area; the device also includes a first sending module, which is used to send the first fused voice from the message editing area to the message display area in response to a sending instruction for the first fused voice; the second display module is also used to display a second fused voice with a second voice feature in response to a fusion operation of the first fused voice and the second conversation expression in the message display area.

[0015] In the above scheme, the fusion module is further used to, in response to a fusion operation on the voice control and the first conversational expression among the at least one conversational expression, display the first fused voice generated based on the first conversational expression in the associated area of the voice control; or, in response to a fusion operation on the voice control and the first conversational expression among the at least one conversational expression, cancel the display of the voice control and display the first fused voice generated based on the first conversational expression at the position of the voice control.

[0016] In the above scheme, the device also includes a second fusion module, which is used to switch the first fused voice to a third fused voice in response to a fusion operation on the voice control and the third conversation expression in the at least one conversation expression; wherein the content of the third fused voice includes content associated with the first conversation expression and the third conversation expression.

[0017] In the above scheme, the third fused voice is carried in a message bubble, which includes a third conversation expression; the device also includes a deletion module, which is used to switch the third fused voice to the first fused voice in response to a deletion operation on the third conversation expression.

[0018] In the above scheme, the third fused voice is carried in a message bubble, which includes a first conversational expression and a third conversational expression. The first conversational expression and the third conversational expression are in a draggable state and have a first sort. The device also includes a sorting adjustment module, which is used to respond to a sorting adjustment operation on the conversational expression in the third fused voice, adjust the first sort to the second sort, and switch the third fused voice to a fourth fused voice, and the content of the fourth fused voice is different from that of the third fused voice.

[0019] In the above scheme, the device also includes a playback module, which is used to obtain the voice features corresponding to the first conversational expression in response to the playback instruction for the first fused speech; and play the first fused speech based on the voice features corresponding to the first conversational expression.

[0020] In the above scheme, there are multiple conversational expressions, and the multiple conversational expressions belong to at least two expression categories, and different expression categories correspond to different voice features; the playback module is also used to respond to the playback instruction for the first fused speech, determine the target voice features corresponding to the target expression category to which the first conversational expression belongs; and use the target voice features corresponding to the target expression category to which the first conversational expression belongs to play the first fused speech.

[0021] In the above solution, the playback module is further used to obtain the feature value of the target voice feature associated with the first conversational expression, and the feature values of the voice features associated with different conversational expressions are different; based on the feature value of the target voice feature, the first fused voice is played.

[0022] In the above solution, the first fused voice is carried in a message bubble, the message bubble includes the first conversational expression, and the at least one conversational expression includes the first conversational expression and a fourth conversational expression. The device also includes a third fusion module, and the third fusion module is used to compare the expression category to which the fourth conversational expression belongs and the target expression category in response to the fusion operation on the first fused voice and the fourth conversational expression; when the comparison result indicates that the expression category to which the fourth conversational expression belongs is consistent with the target expression category, a fourth fused voice with a third voice feature is displayed; wherein the feature value of the third voice feature is the sum of the feature value corresponding to the fourth conversational expression and the feature value of the target voice feature; when the comparison result indicates that the expression category to which the fourth conversational expression belongs is inconsistent with the target expression category, a fifth fused voice with a fourth voice feature is displayed; wherein the fourth voice feature is different from the target voice feature, and the fourth voice feature corresponds to the fourth conversational expression.

[0023] In the above scheme, the conversation interface includes a message display area and a message editing area, and the first fused voice is displayed in the message editing area; the device also includes a third display module, and the third display module is used to respond to a sending instruction for the first fused voice and send the first fused voice from the message editing area to the message display area; in the message display area, the first fused voice is displayed through a message bubble.

[0024] In the above scheme, the display module is also used to display the first fused voice generated based on the first conversational expression when the number of conversational expressions fused for the voice control is between the first number and the second number; the device also includes a prompt module, which is used to display a prompt message when the number of conversational expressions fused for the voice control is greater than the second number, and the prompt message is used to prompt that the number of expressions fused for the voice control has reached a threshold and the fusion of the voice control and the first conversational expression cannot be performed.

[0025] In the above scheme, the fusion operation includes a voice recording operation and an expression selection operation; the fusion module is also used to perform voice recording in response to the voice recording operation triggered by the voice control; based on the recorded voice, in response to the expression selection operation for the first conversation expression in the at least one conversation expression, display the first fused voice generated based on the first conversation expression.

[0026] In the above scheme, the display module is further used to display a voice control and at least one conversation expression distributed around the voice control in the conversation interface; the device also includes a dragging module, which is used to display the process of the voice control being dragged in response to a drag operation on the voice control; when the voice control is dragged to the first conversation expression and released, the release operation on the voice control is determined as the expression selection operation.

[0027] In the above scheme, the device also includes a sliding module, which is used to display a sliding trajectory of the sliding operation in response to the sliding operation triggered by the voice control; when the end point of the sliding trajectory touches the first conversation expression, the expression selection operation is received in response to the voice control being released.

[0028] In the above scheme, the fusion module is also used to respond to the fusion operation of the voice control and the first conversation expression of the at least one conversation expression, and if the current account has the fusion permission for the voice control and the first conversation expression, display the first fused voice generated based on the first conversation expression.

[0029] In the above scheme, the device also includes a message processing module, which is used to respond to the playback instruction for the first fused voice, obtain the message processing method corresponding to the first conversation expression; play the first fused voice, and use the message processing method corresponding to the first conversation expression to process the first fused voice after playback.

[0030] In the above scheme, the conversation interface includes a message display area, and the first fused voice is displayed in the message display area; the message processing module is also used to, when the message processing mode corresponding to the first conversation expression is the read-and-destroy mode, display the process of the first fused voice disappearing from the message display area when the first fused voice is played; when the message processing mode corresponding to the first conversation expression is the loop playback mode, obtain the loop period and number of loops corresponding to the first conversation expression; after the first fused voice is played, play the first fused voice again based on the loop period and number of loops.

[0031] An embodiment of the present application provides an electronic device, including:

[0032] a memory for storing computer-executable instructions;

[0033] The processor is configured to implement the method for processing the session message provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.

[0034] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute instructions to implement a method for processing a session message provided in an embodiment of the present application.

[0035] An embodiment of the present application provides a computer program product, comprising computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the method for processing conversation messages provided in an embodiment of the present application.

[0036] The embodiments of the present application have the following beneficial effects:

[0037] First, a voice control and at least one conversation emoticon are displayed on the conversation interface. Then, the voice control is fused with a first conversation emoticon from the at least one conversation emoticon to generate a first fused voice message based on the first conversation emoticon. In this way, by fusing the voice control and the conversation emoticon, a voice message is generated, thereby increasing the diversity of conversation messages during the conversation. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 1 is a schematic diagram of the architecture of a session message processing system 100 provided in an embodiment of the present application;

[0039] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0040] Figure 3 This is a flow chart of a method for processing session messages provided in an embodiment of the present application;

[0041] Figure 4 Schematic diagram of the voice control and conversation interface provided by the embodiment of the present application;

[0042] Figure 5 is a schematic diagram showing a process of generating a first fused speech based on a first conversation expression, provided by an embodiment of the present application;

[0043] Figure 6 2 is a schematic diagram of displaying a second fused voice using a second message format provided in an embodiment of the present application;

[0044] Figure 7 is a schematic diagram of switching the first fused voice to the third fused voice provided in an embodiment of the present application;

[0045] Figure 8 This is a schematic diagram of displaying a first conversation expression in a message bubble provided by an embodiment of the present application;

[0046] Figure 9 Schematic diagram of a voice control and at least one conversational expression provided in an embodiment of the present application;

[0047] Figure 10 is a schematic diagram showing a first voice control and a second voice control provided in an embodiment of the present application;

[0048] Figure 11 1 is a schematic diagram of a process for processing the first fused speech that has been played, provided in an embodiment of the present application;

[0049] Figure 12 This is a flowchart of the operation process of a user changing the chat bubble style through expressions and gestures provided by an embodiment of the present application;

[0050] Figure 13 This is a flowchart of the operation process of floating and restoring the super bubble scene bullet screen provided in an embodiment of the present application;

[0051] Figure 14 This is a flowchart of the process of changing bubble emotions through user operations provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0053] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0054] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0056] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0057] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0058] 2) Client, also known as user end, refers to the program corresponding to the server that provides local services to users. Except for some applications that can only run locally, it is generally installed on the terminal and needs to cooperate with the server to run. That is, there must be corresponding servers and service programs on the network to provide corresponding services. In this way, specific communication connections need to be established between the client and server to ensure the normal operation of the application, such as virtual scene clients (such as game clients) and video clients.

[0059] 3) Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0060] 4) Message bubble, also known as chat bubble, refers to a style box that surrounds the text sent by the user in a chat scene.

[0061] 5) Pinch operation, a gesture operation performed on a touch screen device by using two fingers (usually the thumb and index finger) at the same time, is used to zoom in or out on the content on the screen, such as a web page, image or map. When the two fingers are in contact with the screen at the same time and move in a relatively stationary manner, if the fingers are close (approaching each other on the screen), the content on the display screen will be zoomed out; if the fingers are separated (moving away from each other on the screen), the content on the display screen will be zoomed in.

[0062] See also Figure 1 , Figure 1This is an architectural diagram of a session message processing system 100 provided in an embodiment of the present application, including a terminal (terminal 400 is shown as an example), wherein terminal 400 is connected to server 200 via network 300, wherein network 300 may be a wide area network or a local area network, or a combination of the two, and data transmission is achieved using wireless or wired links.

[0063] The server 200 is configured to send display data of the conversation interface including the voice control to the terminal 400;

[0064] Terminal 400 is used to receive display data of a conversation interface including a voice control, and based on the display data, display the voice control and at least one conversation expression in the conversation interface; in response to a fusion operation on the voice control and a first conversation expression in at least one conversation expression, display a first fused voice generated based on the first conversation expression.

[0065] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, an intelligent voice interaction device, a smart home appliance, a virtual reality device, a vehicle-mounted terminal, an aircraft, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device, an intelligent speaker, and a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0066] Next, the electronic device implementing the method for processing session messages provided in the embodiment of the present application is described. Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device can be a server or a terminal. Figure 1 Take the terminal shown in as an example, Figure 2 The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2Various buses are labeled as bus system 440 .

[0067] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0068] The user interface 430 includes one or more output devices 431 that enable display of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0069] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0070] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0071] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0072] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0073] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).

[0074] a presentation module 453 for enabling display of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0075] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0076] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 The apparatus 455 for processing conversation messages stored in the memory 450 is shown. This apparatus can be software in the form of a program or plug-in, and includes the following software modules: a display module 4551 and a fusion module 4552. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0077] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus for processing a session message provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the session message processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0078] In some embodiments, the terminal or server can implement the method for processing session messages provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a local client, that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP, a web browser APP; it can also be a small program, that is, a program that can be run only by downloading it into a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of client, module or plug-in.

[0079] Based on the above description of the session message processing system and electronic device provided by the embodiment of the present application, the following describes the session message processing method provided by the embodiment of the present application. In actual implementation, the session message processing method provided by the embodiment of the present application can be implemented by the terminal or the server alone, or by the terminal and the server in collaboration, so that Figure 1 The terminal 400 in the embodiment of the present application alone performs the method for processing the session message as an example for explanation. Figure 3 , Figure 3 This is a flow chart of the method for processing session messages provided by the embodiment of the present application. Next, Figure 3 The steps shown are explained.

[0080] Step 101: The terminal displays a voice control and at least one conversation emoticon in a conversation interface.

[0081] In actual implementation, the terminal is provided with a client that supports the processing of conversation messages, such as a video playback client, a browser client, a social client, etc. When the user opens the client on the terminal and the terminal runs the client, the terminal can display a conversation interface for the target object based on the client, and display voice controls in the conversation interface.

[0082] It should be noted that the voice control is used to generate voice messages or to record voice to generate voice messages, and the conversation emoticon refers to the expression message that can be sent for communication in the conversation interface; at the same time, the conversation interface includes a message editing area and a message display area. The message display area is an area where at least two objects engaged in a conversation in the conversation interface display the conversation messages sent, and the message editing area is an area where at least two objects engaged in a conversation in the conversation interface edit the conversation messages. The voice control is displayed in the message editing area, and the conversation emoticon can be displayed in the message editing area or in the message display area. This is not limited in the embodiments of the present application.

[0083] For example, see Figure 4 , Figure 4 This is a schematic diagram of the voice control and conversation interface provided by the embodiment of the present application, based on Figure 4 , 401 indicates the message display area, 402 indicates the message editing area, 403 indicates the voice control, the conversation expression in the dotted box 402 is the conversation expression displayed in the message editing area, and the conversation expression indicated by 404 is the conversation expression displayed in the message display area.

[0084] Step 102 : In response to a fusion operation on a voice control and a first conversation expression in at least one conversation expression, display a first fused voice generated based on the first conversation expression.

[0085] In actual implementation, the voice control and the first conversation expression can be merged only when the current account has permission. Specifically, in response to the fusion operation on the voice control and the first conversation expression of the at least one conversation expression, the process of displaying the first fused voice generated based on the first conversation expression can be, in response to the fusion operation on the voice control and the first conversation expression of the at least one conversation expression, if the current account has permission to merge the voice control and the first conversation expression, displaying the first fused voice generated based on the first conversation expression.

[0086] It should be noted that whether the current account has fusion permissions can be determined based on pre-set conditions, for example, it can be determined based on the level of the account. For example, when the level of the current account reaches the target level, it is determined that the current account has fusion permissions for voice controls and the first conversational expressions; or it can be determined based on the identity of the current account. For example, when the current account is a member account, it is determined that the current account has fusion permissions for voice controls and the first conversational expressions, etc.

[0087] In actual implementation, if the current account does not have the permission to integrate the voice control and the first conversation expression, a permission prompt message will be displayed. The permission prompt message is used to prompt that the current account does not have the permission to integrate the voice control and the first conversation expression, and the voice control and the first conversation expression cannot be integrated.

[0088] It should be noted that the fusion operation on the voice control and the first conversation expression in at least one conversation expression refers to the fusion operation on the voice control and at least one conversation expression displayed in the conversation interface, that is, the first conversation expression targeted by the fusion operation is a conversation expression or multiple conversation expressions displayed in the conversation interface, and the fusion operation is to fuse the voice control and the conversation expression or multiple conversation expressions displayed in the conversation interface, that is, the first conversation expression can indicate one conversation expression or multiple conversation expressions, which is not limited in this embodiment of the present application.

[0089] In actual implementation, as described above, the conversation interface includes a message editing area, and the voice control is displayed in the message editing area; thus, in response to a fusion operation on the voice control and a first conversation expression in at least one conversation expression, a process of displaying a first fused voice generated based on the first conversation expression can be, in response to a drag operation on the first conversation expression, dragging the first conversation expression to the voice control in the message editing area; in response to a release operation on the first conversation expression, displaying the first fused voice generated based on the first conversation expression in the message editing area.

[0090] For example, see Figure 5 , Figure 5This is a schematic diagram showing a process of generating a first fused speech based on a first conversation expression provided by an embodiment of the present application. Figure 5 , Figure 5 In a, 501 indicates a message display area, 502 indicates a message editing area, and 504 indicates a voice control. Figure 5 The drag operation of the first conversation expression indicated by 503 in a is performed to drag the first conversation expression to the message editing area. Figure 5 In response to the release operation for the first conversation expression, in the message editing area, the generated expression based on the first conversation is displayed, such as Figure 5 The first fused speech is indicated by 505 in b.

[0091] It should be noted that at least one of the speech content and speech features of the first fused speech is associated with the first conversational expression. The speech features include at least one of the speech timbre, pitch, volume, breath, and accent.

[0092] In some embodiments, when the speech content of the first fused speech is associated with the first conversational emoticon, the speech content of the first fused speech is generated based on the first conversational emoticon. Specifically, in response to a fusion operation on a speech control and the first conversational emoticon in at least one conversational emoticon, the pre-set text content corresponding to the first conversational emoticon is obtained. Here, different conversational emoticons correspond to different text content. The text content is then voice-converted to obtain the speech content of the first fused speech, thereby determining to display the first fused speech generated based on the first conversational emoticon and displaying it. For example, when the first conversational emoticon is a monkey emoticon, the corresponding text content may be "Are you the rescuer sent by the monkey?", and the text content is voice-converted to obtain the speech content of the first fused speech. When the first conversational emoticon is a crown emoticon, the corresponding text content may be "If you want to wear the crown, you must bear its weight." The text content is voice-converted to obtain the speech content of the first fused speech.

[0093] It should be noted that when the voice content of the first fused voice is associated with the first conversational expression, the voice features of the first fused voice are pre-set and have nothing to do with the first conversational expression; thus, after performing voice conversion on the text content to obtain the voice content of the first fused voice, the pre-set voice features are obtained, and then the first fused voice is generated based on the pre-set voice features and the converted voice content.

[0094] It should be noted that the process of converting text content into speech to obtain the speech content of the first fused speech can be carried out through a trained text-to-speech model, thereby converting the text content into speech based on the text-to-speech model to obtain the speech content of the first fused speech.

[0095] In some embodiments, when the voice features of the first fused speech are associated with the first conversational expression, in response to the fusion operation on the speech control and the first conversational expression in at least one conversational expression, the voice features corresponding to the pre-set first conversational expression are obtained, where different conversational expressions correspond to different voice features, or the conversational expressions can be classified, and conversational expressions of different expression categories correspond to different voice features; then, based on the obtained voice features corresponding to the first conversational expression, the first fused speech is generated.

[0096] It should be noted that when the speech features of the first fused speech are associated with the first conversational expression, the speech content of the first fused speech is pre-set and unrelated to the first conversational expression. Thus, the speech features associated with the first conversational expression and the pre-set speech content are obtained, and then the first fused speech is generated based on the pre-set speech content and the speech features associated with the first conversational expression. For example, when the first conversational expression is a monkey expression, the corresponding speech features may be the timbre of Monkey King from Journey to the West, and the first fused speech is generated based on this speech feature. When the first conversational expression is an angry expression, the corresponding speech features may be the timbre and volume of Zhang Fei from Romance of the Three Kingdoms, and the first fused speech is generated based on this speech feature. When the first conversational expression is a snowflake expression, the corresponding speech features may be a southern accent, and the first fused speech is generated based on this speech feature.

[0097] In other embodiments, when the voice features and voice content of the first fused voice are both associated with the first conversational expression, in response to the fusion operation on the voice control and the first conversational expression in at least one conversational expression, the voice content and voice features corresponding to the first conversational expression are obtained, thereby generating the first fused voice based on the voice content and voice features. For example, continue to refer to Figure 5 , when the first conversation expression is as follows Figure 5 When the monkey expression indicated by 503 in a is shown, the corresponding voice feature can be the timbre of Monkey King in Journey to the West, and the corresponding text content can be "Are you the reinforcement sent by the monkey?", thereby generating the first fused voice based on the voice content and voice feature.

[0098] In some embodiments, there may be multiple display positions for the first fused speech. Specifically, in response to a fusion operation on a voice control and a first conversational expression in at least one conversational expression, the process of displaying the first fused speech generated based on the first conversational expression may be, in response to the fusion operation on the voice control and the first conversational expression in at least one conversational expression, displaying the first fused speech generated based on the first conversational expression in the associated area of the voice control; or in response to the fusion operation on the voice control and the first conversational expression in at least one conversational expression, canceling the display of the voice control and displaying the first fused speech generated based on the first conversational expression at the position of the voice control.

[0099] It should be noted that the associated area of the voice control can be one of the left area, right area, upper area and lower area of the voice control; and canceling the display of the voice control and displaying the first fused voice generated based on the first conversation expression in the position of the voice control means switching the voice control to the first fused voice; so that in the subsequent process, the first fused voice can be played in response to a trigger operation based on the first fused voice, such as a click operation.

[0100] In some embodiments, after displaying the first fused voice generated based on the first conversation expression, a switch control may be displayed in an associated area of the first fused voice, and the first fused voice may be automatically played; wherein the switch control is used to switch the content of the first fused voice; in response to a triggering operation on the switch control, the content of the first fused voice is switched to new content, and the new content is associated with the first conversation expression.

[0101] It should be noted that the associated area of the first fused voice can be one of the left area, right area, upper area, and lower area of the first fused voice; in addition to automatically playing the first fused voice, the first fused voice can also be played in response to a play instruction for the first fused voice when a switch control is displayed in the associated area of the first fused voice; after playing the first fused voice, if the user is not satisfied with the first fused voice, a trigger operation on the switch control can be executed to switch the first fused voice, that is, in response to the trigger operation on the switch control, the content of the first fused voice is switched to new content. The content of the first fused voice includes at least one of voice content and voice features, and the new content also includes at least one of voice content and voice features.

[0102] In some embodiments, the first fused voice has a first voice feature, so after displaying the first fused voice generated based on the first conversation expression, it is also possible to display a second fused voice with a second voice feature in response to a fusion operation on the first fused voice and a second conversation expression in at least one conversation expression; wherein the second voice feature is different from the first voice feature, and the second voice feature corresponds to the second conversation expression.

[0103] It should be noted that the first voice feature and the second voice feature are the same as the voice features described above, including at least one of the voice timbre, pitch, volume, breath, accent, etc.; after the first fused voice is displayed, the user can play the first fused voice. If the user is not satisfied with the first voice feature included in the first fused voice, the user can adjust the first voice feature of the first fused voice based on the second conversational expression, that is, fuse the first fused voice with the second conversational expression to obtain a second fused voice with the second voice feature.

[0104] For example, when the first conversational expression is an angry expression, the first voice feature of the first fused voice corresponds to the timbre and volume of Zhang Fei in Romance of the Three Kingdoms. If the user is not satisfied with the current first voice feature, the first fused voice can be fused with the second conversational expression, such as a monkey expression, to obtain a second fused voice. The second voice feature of the second fused voice corresponds to the timbre of Monkey King in Journey to the West.

[0105] It should be noted that the fusion process of the first fused voice and the second conversational expression in at least one conversational expression is similar to the fusion process of the voice control and the first conversational expression in at least one conversational expression described above. This embodiment of the present application will not be elaborated on.

[0106] In actual implementation, the conversation interface includes a message display area and a message editing area. The first fused voice can be displayed in the message editing area or in the message display area. This embodiment of the present application does not limit this.

[0107] When the first fused voice is displayed in the message display area, the voice features of the first fused voice can be directly adjusted based on the second conversational expression to obtain the second fused voice; and when the first fused voice is displayed in the message editing area, the voice features of the first fused voice can be adjusted based on the second conversational expression to obtain the second fused voice, and then the second fused voice can be sent to the message display area in response to a send instruction for the second fused voice; or, the first fused voice can be first sent to the message display area, and then the voice features of the first fused voice can be adjusted based on the second conversational expression to obtain the second fused voice. Specifically, after displaying the first fused voice generated based on the first conversational expression, the first fused voice can also be sent from the message editing area to the message display area in response to a send instruction for the first fused voice. Thus, the process of displaying the second fused voice having the second voice features in response to a fusion operation on the first fused voice and the second conversational expression in at least one conversational expression can be displaying the second fused voice having the second voice features in response to a fusion operation on the first fused voice and the second conversational expression in the message display area.

[0108] In some embodiments, when the first fused voice is displayed in the message display area, the fusion may not only change the voice features of the first fused voice, but also change the message style of the first fused voice. Specifically, after the first fused voice is sent from the message editing area to the message display area, the first fused voice is displayed in the conversation interface using the first message style; thus, in response to the fusion operation of the first fused voice and the second conversation expression in the message display area, the process of displaying the second fused voice with the second voice features may be, in response to the fusion operation of the first fused voice and the second conversation expression, using the second message style to display the second fused voice obtained by fusing the first fused voice and the second conversation expression; wherein, the second message style is different from the first message style.

[0109] It should be noted that the term "message style" refers to the display format of conversational messages, such as voice messages and emoticons, in a conversational interface. For example, it can indicate a message bubble or a message card. Different message styles are used to indicate at least one of the shape, color, and size of the message bubble or message card, and this embodiment of the present application does not limit this. The first message style can be pre-set; the second message style, which is different from the first message style, can include at least one of the shape, color, and size of the message bubble or message card. The second message style is used to assist in indicating the content of the first fused speech, such as the speech content and speech characteristics.

[0110] For example, see Figure 6 , Figure 6 This is a schematic diagram of displaying the second fusion voice using the second message style provided by the embodiment of the present application, based on Figure 6 , Figure 6 In a, 601 indicates the message display area, and 602 indicates the message editing area. After displaying the first fusion voice generated based on the first conversation expression, in response to the sending instruction for the first fusion voice, the first fusion voice is sent from the message editing area to the message display area, and then in the conversation interface, the following is used: Figure 6 The first message style indicated by 603 in a is displayed, the first fusion voice is displayed, and then in response to the first message in the message display area Figure 6 The first fused speech indicated by 603 in a and Figure 6 The fusion operation of the second conversation expression indicated by 604 in a is performed as follows Figure 6 The second message style indicated by 605 in b is displayed as follows Figure 6 The second fused speech indicated by 605 in b.

[0111] It should be noted that in this application, in addition to the first fused voice, when other fused voices are displayed in the message display area, the fusion will not only change the voice characteristics of other fused voices, but also change the message style of other fused voices. This embodiment of the application will not go into details.

[0112] In some embodiments, each conversational expression has corresponding voice content. At the same time, corresponding voice content also exists after at least two conversational expressions are combined. Specifically, after displaying the first fused voice generated based on the first conversational expression, the first fused voice can be switched to a third fused voice in response to a fusion operation on the voice control and the third conversational expression in at least one conversational expression; wherein the content of the third fused voice includes content associated with the first conversational expression and the third conversational expression.

[0113] It should be noted that the voice content of the third fused speech may be generated based on the voice content corresponding to the third conversational expression and the voice content of the first fused speech after respectively determining the voice content corresponding to the third conversational expression and the voice content of the first fused speech. Specifically, upon receiving a fusion operation for the voice control and the third conversational expression in at least one conversational expression, the voice content corresponding to the third conversational expression and the voice content of the first fused speech are obtained, thereby generating the content of the third fused speech, i.e., the third fused speech, based on the voice content corresponding to the third conversational expression and the voice content of the first fused speech.

[0114] Alternatively, the voice content of the third fused speech can directly correspond to the third conversational expression and the first conversational expression. That is, after presetting the voice content corresponding to different conversational expressions, corresponding voice content is also set for an expression combination including at least two conversational expressions. Specifically, when a fusion operation is received for a voice control and a third conversational expression in at least one conversational expression, the voice content corresponding to the first conversational expression and the third conversational expression is obtained, and then the content of the third fused speech, i.e., the third fused speech, is generated based on the voice content. This embodiment of the present application is not limited to this.

[0115] It should be noted that the voice content of the third fused voice can be composed of at least two independent voice contents, such as the voice content corresponding to the third conversational expression and the voice content of the first fused voice, that is, the voice content of the third fused voice includes the voice content corresponding to the third conversational expression and the voice content of the first fused voice; or the voice content of the third fused voice can be regenerated, for example, based on the voice content corresponding to the third conversational expression and the voice content of the first fused voice, or directly obtaining the voice content corresponding to the first conversational expression and the third conversational expression. This embodiment of the application does not limit this.

[0116] For example, see Figure 7 , Figure 7 This is a schematic diagram of switching the first fusion voice to the third fusion voice provided in an embodiment of the present application, based on Figure 7 In a, 701 indicates the first fused voice, wherein the first fused voice is generated based on the monkey expression, and the voice content of the first fused voice can be "Are you the rescuer sent by the monkey?" The voice feature of the first fused voice can be the timbre of Monkey King in Journey to the West, and 702 indicates the third conversation expression, i.e., the crown expression, and the voice content corresponding to the third conversation expression can be "If you want to wear the crown, you must bear its weight." Thus, in response to the fusion operation of the first fused voice and the third conversation expression, the following is displayed: Figure 7 The third fused speech indicated by 703 in step b of the preceding text may include the speech content of the third fused speech, which may be "You are truly a monkey wearing a golden crown—pretending to be a monkey," and the speech features of the third fused speech may be the timbre of Monkey King from Journey to the West. The speech content of the third fused speech may be generated based on the speech content corresponding to the third conversational expression and the speech content of the first fused speech.

[0117] It should be noted that the voice features of the third fused speech may correspond to the first conversational expression, the third conversational expression, or a combination of the first and third conversational expressions. For example, the first conversational expression corresponds to voice feature A, the third conversational expression corresponds to voice feature B, and the combination of the first and third conversational expressions may correspond to a new voice feature C. Therefore, the voice features of the third fused speech may be voice feature A, voice feature B, or voice feature C. This is not limited in this embodiment of the present application.

[0118] In some embodiments, the first fused voice is carried in a message bubble, which includes a first conversation expression. When the first fused voice is switched to the third fused voice, the message bubble may include not only the first conversation message but also the third conversation message, thereby indicating that the first fused voice has been switched to the third fused voice based on the displayed first conversation message and the third conversation message.

[0119] In actual implementation, after the first fused voice is switched to the third fused voice, the third fused voice may be switched to the first fused voice in response to a deletion operation on the third conversation expression.

[0120] It should be noted that when a conversational expression is deleted, the speech content and / or speech features associated with the corresponding conversational expression included in the fused speech are cleared. Figure 7 In response to the deletion operation on the crown expression, the third fusion voice is switched to the first fusion voice, that is, the following is displayed: Figure 7 The first fused speech indicated by 701 in a, wherein the speech content of the first fused speech can be "Are you the rescuer sent by the monkey?", and the speech feature of the first fused speech can be the timbre of Monkey King in Journey to the West.

[0121] In actual implementation, there are multiple ways to implement the deletion operation of the third conversation expression. In some embodiments, a deletion indicator for each conversation expression integrated into the third fused speech can be displayed in the associated area of the third fused speech. The deletion indicator is used to clear the content (including voice content and / or voice features) associated with the corresponding conversation expression in the third fused speech, and multiple deletion indicators correspond to multiple conversation expressions integrated into the third fused speech. Then, a triggering operation for the deletion indicator corresponding to the third conversation expression among the multiple deletion indicators is used as the deletion operation for the third conversation expression.

[0122] In other embodiments, in response to a drag operation on the third conversation expression, a drag trajectory of the third conversation expression may be displayed; when the drag trajectory indicates that the distance between the third conversation expression and the third fused voice is not less than a target distance, in response to the third conversation expression being released, a deletion operation on the third conversation expression is received.

[0123] It should be noted that the associated area of the third fused voice can be one of the upper area, lower area, left area and right area of the third fused voice; and the target distance can also be pre-set; at the same time, the methods for implementing the deletion operation for the third conversation expression include but are not limited to the above two methods, and the embodiments of this application do not limit this.

[0124] In actual implementation, the content of the fused voice can also be changed by changing the order of the conversation messages in the fused voice; the third fused voice is carried in a message bubble, which includes the first conversation expression and the third conversation expression. At the same time, the first conversation expression and the third conversation expression are in a draggable state and have the first order; thus, after switching the first fused voice to the third fused voice, it is also possible to adjust the first order to the second order in response to the order adjustment operation of the conversation expressions in the third fused voice, and switch the third fused voice to the fourth fused voice, and the content of the fourth fused voice is different from that of the third fused voice.

[0125] It should be noted that the content of the fused speech includes at least one of the speech features and the speech content, and the content of the fused speech is associated with the arrangement order of the conversational expressions, and the content of the fused speech corresponding to different arrangement orders is different; based on this, when the arrangement order of the conversational expressions in the third fused speech is the first sorting, the content of the third fused speech is associated with the first sorting, and in response to the sorting adjustment operation of the conversational expressions in the third fused speech, after the first sorting is adjusted to the second sorting, the content of the fourth fused speech is associated with the second sorting, thereby switching the third fused speech to the fourth fused speech.

[0126] In some embodiments, the first fused voice can also be played. Specifically, after displaying the first fused voice generated based on the first conversational expression, it is also possible to obtain the voice features corresponding to the first conversational expression in response to a play instruction for the first fused voice; and play the first fused voice based on the voice features corresponding to the first conversational expression.

[0127] It should be noted that different conversational expressions can correspond to different voice features. As described above, the voice features include at least one of the voice timbre, pitch, volume, breath, accent, etc.; at the same time, as described above, the conversation interface can also include a message editing interface and a message display interface, and the first fused voice can be displayed on the message editing interface or on the message display interface. This embodiment of the application does not limit this.

[0128] In actual implementation, in addition to different conversational expressions corresponding to different voice features, different conversational expression categories can also correspond to different voice features. Specifically, there are multiple conversational expressions, and the multiple conversational expressions belong to at least two expression categories, and different expression categories correspond to different voice features. Therefore, in response to the playback instruction for the first fused voice, the process of obtaining the voice features corresponding to the first conversational expression can be, in response to the playback instruction for the first fused voice, determining the target voice features corresponding to the target expression category to which the first conversational expression belongs; and the process of playing the first fused voice based on the voice features corresponding to the first conversational expression can be, using the target voice features corresponding to the target expression category to which the first conversational expression belongs to play the first fused voice.

[0129] It should be noted that before determining the target speech feature corresponding to the target expression category to which the first conversational expression belongs, it is first necessary to determine the target expression category to which the first conversational expression belongs. Here, the correspondence between the conversational expression and the expression category is pre-set, that is, after the first conversational expression is determined, the expression category to which the first conversational expression belongs can be determined based on the pre-set correspondence. At the same time, as mentioned above, the speech feature includes at least one of pitch, timbre, volume, breath and accent. Different conversational expressions correspond to different speech features. For example, some conversational expressions correspond to timbre, some conversational expressions correspond to volume, and some conversational expressions correspond to timbre and volume. For example, for conversational expressions A, B, C, and D, these four conversational expressions correspond to different speech features. The speech feature of conversational expression A is timbre, the speech feature of conversational expression B is timbre, the speech feature of conversational expression C is timbre and volume, and the speech feature of conversational expression D is volume.

[0130] Different expression categories correspond to different voice features. For example, conversational expressions are divided into four expression categories. The conversational expressions of the first expression category correspond to volume, the conversational expressions of the second expression category correspond to timbre, the conversational expressions of the third expression category correspond to tone, and the conversational expressions of the fourth expression category correspond to timbre and volume. For example, conversational expressions A and conversational expressions B belong to the same expression category, and conversational expressions C and conversational expressions D belong to the same expression category. Then there are two expression categories here, and different expression categories correspond to different voice features, while the same expression category corresponds to corresponding voice features, that is, the voice features corresponding to conversational expressions A and conversational expressions B are both timbre, and the voice features corresponding to conversational expressions C and conversational expressions D are both timbre and volume.

[0131] In actual implementation, the process of playing the first fused voice using the target voice feature corresponding to the target expression category to which the first conversational expression belongs can be to obtain the feature value of the target voice feature associated with the first conversational expression, where the feature values of the voice feature associated with different conversational expressions are different; and playing the first fused voice based on the feature value of the target voice feature.

[0132] It should be noted that when the target speech feature corresponds to volume, the characteristic value of the target speech feature indicates the volume when the first fused speech is played; when the target speech feature corresponds to pitch, the characteristic value of the target speech feature indicates the pitch when the first fused speech is played; when the target speech feature corresponds to timbre, the characteristic value of the target speech feature indicates the identifier of the timbre when the first fused speech is played, that is, it is used to determine the timbre when the first fused speech is played; when the target speech feature corresponds to timbre and volume, the characteristic value of the target speech feature indicates the volume and timbre identifier when the first fused speech is played, so that based on the characteristic value of the target speech feature such as the volume and timbre identifier, the timbre indicated by the timbre identifier and the corresponding volume are used to play the first fused speech.

[0133] It should be noted that when different conversational expressions correspond to different voice features, in response to the playback instruction for the first fused voice, the target voice features corresponding to the first conversational expression are used to play the first fused voice. At the same time, the process is similar to the process of playing the first fused voice in response to the playback instruction for the first fused voice, using the target voice features corresponding to the target expression category to which the first conversational expression belongs, as described above. The embodiments of the present application do not limit this.

[0134] For example, when the first conversational expression is an angry expression, the target speech features corresponding to the target expression category to which the first conversational expression belongs are timbre and volume. The characteristic value of the target speech feature indicates that the volume when the first fused speech is played is XX decibels. At the same time, the timbre identifier is A. The A identifier indicates that the timbre when the first fused speech is played is the timbre of Zhang Fei in Romance of the Three Kingdoms. Based on this, the timbre and XX decibels of Zhang Fei in Romance of the Three Kingdoms are used to play the first fused speech.

[0135] In actual implementation, the first fused voice is carried in a message bubble, which includes a first conversational expression, and at least one conversational expression includes the first conversational expression and the fourth conversational expression. After displaying the first fused voice generated based on the first conversational expression, it is also possible to, in response to the fusion operation of the first fused voice and the fourth conversational expression, compare the expression category to which the fourth conversational expression belongs with the target expression category; when the comparison result indicates that the expression category to which the fourth conversational expression belongs is consistent with the target expression category, display a fourth fused voice with a third voice feature; wherein the feature value of the third voice feature is the sum of the feature value corresponding to the fourth conversational expression and the feature value of the target voice feature; when the comparison result indicates that the expression category to which the fourth conversational expression belongs is inconsistent with the target expression category, display a fifth fused voice with a fourth voice feature; wherein the fourth voice feature is different from the target voice feature, and the fourth voice feature corresponds to the fourth conversational expression.

[0136] It should be noted that when the conversation interface includes a message display area and a message editing area, the first fused voice here is displayed in the message display area, and similar to the first conversation expression described above, the fourth conversation expression can also be displayed in the message display area, or can also be displayed in the message editing area, and this embodiment of the present application does not limit this; at the same time, the fusion process of the fourth conversation expression and the first fused voice is also similar to the fusion process of the first conversation expression and the voice control, and this embodiment of the present application does not elaborate on this. When the fourth conversation expression is displayed in the message display area, the fusion process of the fourth conversation expression and the first fused voice can also be achieved by kneading the fourth conversation expression and the first fused voice. There are many ways to implement the fusion process of the fourth conversation expression and the first fused voice, and this embodiment of the present application does not limit this.

[0137] At the same time, when it is determined that the expression category to which the fourth conversation expression belongs is consistent with the target expression category, it is first necessary to determine the expression category to which the fourth conversation expression belongs. Here, the process of determining the expression category to which the fourth conversation expression belongs is similar to the process of determining the expression category to which the first conversation expression belongs as described above. This embodiment of the present application will not be elaborated on.

[0138] In actual implementation, when the comparison result indicates that the expression category to which the fourth conversation expression belongs is consistent with the target expression category, the characteristic value of the speech feature corresponding to the expression category to which the fourth conversation expression belongs is obtained, and then the characteristic value is weightedly summed with the characteristic value of the target speech feature to obtain the characteristic value of the third speech feature, thereby determining the fourth fused speech with the third speech feature.

[0139] It should be noted that when the speech parameters corresponding to the target speech feature, such as timbre, volume, pitch, etc., are not exactly the same as the speech parameters corresponding to the speech feature corresponding to the expression category to which the fourth conversational expression belongs, the characteristic values of the same speech parameters are weighted and summed, and different speech parameters are used directly. For example, when the target speech feature corresponds to volume, the corresponding characteristic value, that is, the size of the volume, is a, and the speech feature corresponding to the expression category to which the fourth conversational expression belongs corresponds to timbre and volume, the corresponding characteristic value, that is, the identification of the timbre is a, and the size of the volume is b, then the volume a and the volume b are weighted and summed to obtain the volume c, so that the third speech feature corresponds to volume and timbre, and the characteristic value of the third speech feature, that is, the size of the volume, is c, and the identification of the timbre is a.

[0140] In actual implementation, when the comparison result indicates that the expression category to which the fourth conversational expression belongs is inconsistent with the target expression category, the target speech feature in the first fused speech is directly converted into the fourth speech feature. For example, when the target speech feature corresponds to volume, the corresponding characteristic value, that is, the size of the volume, is a, and the speech feature corresponding to the expression category to which the fourth conversational expression belongs corresponds to timbre and volume, the corresponding characteristic value, that is, the identifier of the timbre is a, and the size of the volume is b, then after the first fused speech is fused with the fourth conversational expression, the fourth speech feature corresponds to timbre and volume, the characteristic value of the fourth speech feature, that is, the identifier of the timbre is a, and the size of the volume is b.

[0141] In some embodiments, as described above, the conversation interface includes a message display area and a message editing area, and the first fused voice is displayed in the message editing area; thus, after displaying the first fused voice generated based on the first conversation expression, it is also possible to send the first fused voice from the message editing area to the message display area in response to a sending instruction for the first fused voice; in the message display area, the first fused voice is displayed through a message bubble.

[0142] It should be noted that when the first fused voice is displayed through a message bubble, the message bubble may not include the first conversation expression, or may also include the first conversation expression, which is not limited in this embodiment of the present application; for example, see Figure 8 , Figure 8 This is a schematic diagram showing a first conversation expression in a message bubble provided by an embodiment of the present application, based on Figure 8 , the first fused voice as indicated by 801 is displayed through a message bubble, wherein the message bubble may include a first conversation expression as indicated by 802.

[0143] In some embodiments, the process of displaying the first fused speech generated based on the first conversational expression may be, when the number of conversational expressions fused for the voice control is between a first number and a second number, displaying the first fused speech generated based on the first conversational expression; thereby, when the number of conversational expressions fused for the voice control is greater than the second number, displaying a prompt message, the prompt message being used to prompt that the number of expressions fused for the voice control has reached a threshold and the fusion of the voice control and the first conversational expression cannot be performed.

[0144] It should be noted that when displaying the first fused speech generated based on the first conversational expression, the number of conversational expressions fused by the voice control can be displayed in the voice control or the associated area of the first fused speech, or the number of conversational expressions fused by the voice control can be not displayed; at the same time, the first number, the second number and the threshold are all pre-set, and the second number is greater than the first number.

[0145] In some embodiments, the fusion operation includes a voice recording operation and an expression selection operation; thus, in response to the fusion operation on the voice control and the first conversation expression in at least one conversation expression, the process of displaying the first fused voice generated based on the first conversation expression may be, in response to the voice recording operation triggered by the voice control, performing voice recording; based on the recorded voice, in response to the expression selection operation on the first conversation expression in at least one conversation expression, displaying the first fused voice generated based on the first conversation expression.

[0146] It should be noted that the conversation interface includes a message display area and a message editing area, and the voice control and conversation message are displayed in the message editing area; the voice recording operation can refer to a click operation on the voice control or a press operation on the voice control, which is not limited here; at the same time, as mentioned above, the voice feature can be at least one of timbre, accent, pitch, volume, breath, etc. Here, different expressions correspond to different voice features. Therefore, the voice features of the first fused voice obtained by fusing the voice control and the first conversation expression are the voice features corresponding to the first conversation message, and the voice content of the first fused voice is the recorded voice content. For example, when the first conversation expression is a puppy expression, in response to the expression selection operation for the first conversation expression in at least one conversation expression, the first fused voice generated based on the first conversation expression, that is, the voice features of the recorded voice message, are displayed, and the voice features corresponding to the puppy expression are switched.

[0147] It should be noted that, in response to the expression selection operation for the first conversational expression in at least one conversational expression, after displaying the first fused voice generated based on the first conversational expression, it is also possible to respond to the play instruction for the first fused voice and play the first fused voice, that is, to use the voice features corresponding to the first conversational expression to play the recorded voice, and then respond to the sending instruction for the first fused voice and send the first fused voice to the message display area of the conversation interface.

[0148] For example, see Figure 9 , Figure 9 This is a schematic diagram of a voice control and at least one conversational expression provided by an embodiment of the present application, based on Figure 9 The dotted box 901 indicates the message display area, the dotted box 902 indicates the message editing area, 903 indicates the voice control, at least one conversational expression includes 6 conversational expressions distributed around the voice control indicated by 903, and 904 indicates the first conversational expression, i.e., the puppy expression. Thus, in response to the voice recording operation triggered by the voice control indicated by 903, voice recording is performed; based on the recorded voice, in response to the expression selection operation for the first conversational expression indicated by 904 in the at least one conversational expression, the first fused voice generated based on the first conversational expression is displayed.

[0149] In actual implementation, the process of displaying the voice control and at least one conversation expression in the conversation interface may be, in the conversation interface, displaying the voice control and at least one conversation expression distributed around the voice control; and there are multiple ways to receive the expression selection operation. Specifically, after voice recording, in response to a drag operation on the voice control, the process of the voice control being dragged may be displayed; when the voice control is dragged to the first conversation expression and released, the release operation on the voice control is determined as an expression selection operation; or, after voice recording, in response to a sliding operation triggered by the voice control, a sliding track of the sliding operation may be displayed; when the end point of the sliding track touches the first conversation expression, the expression selection operation is received in response to the voice control being released.

[0150] It should be noted that the voice control is in a draggable state, so that the drag operation of the voice control can be realized; at the same time, for the sliding operation triggered by the voice control, the sliding area can be displayed in response to the pressing operation on the voice control, and then the sliding track of the sliding operation can be displayed in response to the sliding operation triggered in the sliding area.

[0151] In some embodiments, the voice features of the first fused speech can also be adjusted. Specifically, the number of at least one conversational expression is multiple. After displaying the first fused speech generated based on the first conversational expression, it is also possible to, in response to a drag operation on a fifth conversational expression, drag the fifth conversational expression to the first fused speech, where the fifth conversational expression is any one of the at least one conversational expression except the first conversational expression; in response to a release operation on the fifth conversational expression, display a sixth fused speech obtained by changing the voice features of the first fused speech to the fifth voice features; wherein the fifth voice features are different from the voice features of the first fused speech, and the fifth voice features correspond to the fifth conversational expression.

[0152] It should be noted that the sixth fused voice obtained by changing the voice features of the first fused voice into the fifth voice features is similar to the process of displaying the second fused voice with the second voice features in response to the fusion operation of the first fused voice and the second conversation expression in the message display area as described above. This embodiment of the present application will not be elaborated on.

[0153] In some embodiments, the conversation interface includes a message editing area, and the voice control is displayed in the message editing area and includes a first voice control and a second voice control. The first voice control can be the voice control for generating the first fused voice as described above, which is used to display the voice recording interface. The voice recording interface includes a second voice control, and the second voice control is the voice control for recording voice messages as described above. At the same time, the message editing area also includes a message input box. The first voice control is displayed in the message input box of the message editing area and is used to display the voice recording interface. Therefore, after the first voice control is displayed in the message input box of the message editing area, in response to a triggering operation for the first voice control, a voice recording interface including the second voice control is displayed, wherein the voice recording interface also includes at least one conversational expression distributed around the second voice control, so that voice recording is performed in response to the voice recording operation triggered based on the second voice control.

[0154] For example, see Figure 10 , Figure 10 This is a schematic diagram showing the first voice control and the second voice control provided by an embodiment of the present application, based on Figure 10 , in such Figure 10 After the message editing area indicated by the dotted box 1001 in a includes the message input box indicated by 1002 and the first voice control indicated by 1003 is displayed, in response to the triggering operation for the first voice control, the message input box indicated by the dotted box 1005 in b includes the message input box indicated by 1003 and the first voice control indicated by 1003. Figure 10The voice recording interface of the second voice control indicated by 1004 in b, wherein the voice recording interface also includes at least one conversational expression distributed around the second voice control, so as to perform voice recording in response to the voice recording operation triggered based on the second voice control.

[0155] In some embodiments, after displaying the first fused voice generated based on the first conversational expression, it is also possible to obtain a message processing method corresponding to the first conversational expression in response to a playback instruction for the first fused voice; play the first fused voice, and use the message processing method corresponding to the first conversational expression to process the played first fused voice.

[0156] It should be noted that the message processing method corresponding to the conversational expressions is pre-set, and the message processing methods corresponding to different conversational expressions may be the same or different, and this embodiment of the present application does not limit this.

[0157] In actual implementation, the conversation interface includes a message display area, and the first fused voice is displayed in the message display area; the process of processing the first fused voice after playback using the message processing method corresponding to the first conversation expression can be, when the message processing method corresponding to the first conversation expression is a read-and-destroy method, when the first fused voice is played, the process of showing the first fused voice disappearing from the message display area; when the message processing method corresponding to the first conversation expression is a loop playback method, the loop period and number of loops corresponding to the first conversation expression are obtained; after the first fused voice is played, the first fused voice is played again based on the loop period and the number of loops.

[0158] It should be noted that when the first fused voice is sent by the target object, playing the first fused voice here may refer to the target object playing the first fused voice itself, or other objects playing the first fused voice. When the message processing method corresponding to the conversation emoticon is the read-and-destroy method, when the first fused voice is played, the process of displaying the first fused voice disappearing from the message display area includes: when the other object completes playing the first fused voice, the process of displaying the first fused voice disappearing from the message display area in the conversation interface of the target object, wherein the other object is the object with which the target object conducts a conversation based on the conversation interface, and here, the first fused voice in the conversation interface of the other object will also gradually disappear; or, it may also be the process of displaying the first fused voice disappearing from the message display area in the conversation interface of the target object when the target object completes playing the first fused voice. This is not limited in this embodiment of the present application.

[0159] For example, see Figure 11 , Figure 11 This is a schematic diagram of a process for processing the first fused voice after playback provided by an embodiment of the present application, based on Figure 11 , Figure 11 In a, 1101 is a message display area, 1102 indicates a message editing area, and 1103 indicates a message display area based on the above. Figure 11 The first fusion voice generated by the first conversation expression indicated by 1104 in a responds to the sending instruction for the first fusion voice, such as Figure 11 As shown in b, the first fusion voice indicated by 1105 is displayed in the message display area, and when the message processing mode corresponding to the first conversation expression is the self-destructing mode, when the first fusion voice is played, the process of the first fusion voice disappearing from the message display area is shown, as shown in FIG. Figure 11 As shown in c.

[0160] It should be noted that, for the case where the message processing method corresponding to the conversational expression is a loop playback method, the number of loops and the loop period are both pre-set. Similarly, when the first fused voice is sent by the target object, playing the first fused voice here can refer to the target object playing the first fused voice itself, or other objects playing the first fused voice. When other objects finish playing the first fused voice, the first fused voice is played again in the conversation interface of the other objects based on the loop period and the number of loops; or, when the target object finishes playing the first fused voice, the first fused voice is played again in the conversation interface of the target object based on the loop period and the number of loops.

[0161] In some embodiments, in response to a fusion operation on a voice control and a first conversational expression in at least one conversational expression, after displaying a first fused voice generated based on the first conversational expression, it is also possible to, in response to a play instruction for the first fused voice, play the first fused voice, and during the playback of the first fused voice, play an animation effect associated with the first conversational expression.

[0162] It should be noted that the animation effect associated with the first conversation expression is pre-set, and the playback method of the animation effect can also be pre-set, for example, it can be played on the entire screen or in the form of a floating window, which is not limited in this embodiment of the application. The animation effect can include a flashing display of the target conversation expression, a color change display, or a wave display such as ripples or spots in the expression, which is not limited in this embodiment of the application.

[0163] In some embodiments, in response to a fusion operation on a voice control and a first conversation expression in at least one conversation expression, after displaying a first fused voice generated based on the first conversation expression, it is also possible to present a text display area including the first conversation expression in response to a text conversion instruction for the first fused voice; in the text display area, present the text content corresponding to the first fused voice.

[0164] It should be noted that the text display area can be in one of the upper area, lower area, left area and right area of the first fused voice, and the position of the first conversation expression in the text display area can also be pre-set.

[0165] In some embodiments, the conversation interface includes a message display area, and the first fused voice is displayed in the message display area; in response to the fusion operation of the voice control and the first conversation expression in at least one conversation expression, after displaying the first fused voice generated based on the first conversation expression, it is also possible to, in response to the play instruction for the first fused voice, play the first fused voice, and when the playback of the first fused voice is completed, convert the first conversation expression into a barrage, and fix the barrage at the target position of the message display area; when the display time of the barrage reaches the first target time, cancel the display of the barrage.

[0166] It should be noted that the target position is used to indicate the display area occupied by the bullet screen. The display area here can be the entire conversation interface or a portion of the conversation interface, or the entire message display area or a portion of the message display area. This is not limited in the embodiments of the present application. When the display area is a portion of the area, the area ratio of the display area to the corresponding conversation interface or message display area is pre-set, such as the display area occupying half of the conversation interface, or the display area occupying half of the message display area.

[0167] It should be noted that the first target duration can be pre-set, such as 5 seconds. When the display duration of the barrage reaches the first target duration, the barrage is canceled. In this way, while ensuring the effect of the barrage, excessive interference with the user's conversation process is prevented, thereby improving the user's conversation experience.

[0168] In actual implementation, the first fused voice is sent by the target object; after the barrage is fixedly displayed at the target position of the message display area, the reply barrage of other objects to the barrage can also be displayed floating above the barrage; among them, the other objects are the objects with which the target object conducts conversations based on the conversation interface.

[0169] It should be noted that the other objects can be one or more. When the other objects see the barrage sent by the target object, they send a reply message in response to the barrage. When the reply message is sent to the message display area in the conversation interface, the reply message is converted into a reply barrage and displayed floating above the barrage sent by the target object.

[0170] In actual implementation, after the barrage is fixedly displayed at the target location in the message display area, if the target party continues to send conversation messages such as text messages and voice messages, the continued conversation messages can also be displayed in the barrage. Specifically, the conversation interface also includes a message editing area, so that after the barrage is fixedly displayed at the target location in the message display area, the edited message content can also be displayed in response to a message editing operation triggered based on the message editing area; in response to a send instruction for the edited message content, the edited message content can be displayed in the barrage at the target location. The message content includes at least one of a text message, an emoticon message, and a voice message.

[0171] Using the above-described embodiments of the present application, a voice control and at least one conversational expression are first displayed on the conversation interface. Then, the voice control is fused with a first conversational expression from the at least one conversational expression to generate a first fused voice message based on the first conversational expression. In this way, by fusing the voice control and the conversational expression, a voice message is generated, thereby increasing the diversity of conversational messages during the conversation.

[0172] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0173] During the conversation process of related technologies, the chat bubble shape is fixed, and the conversation experience is dull and boring. At the same time, user emotions are expressed entirely through text or emoticons, which is a relatively simple method. In addition, misunderstandings of emotions may occur in many scenarios, further reducing the conversation experience.

[0174] Based on this, the present invention provides a new way of text and voice expression for chat dialogues, aiming to express user emotions more accurately and unlock more interactive gameplay. Specifically, by integrating text messages with text messages (text conversation messages), integrating emoticons (expressions) with bubbles (target messages), and integrating emoticons with voice messages in chat scenarios, the bubble state (display style) of voice messages and voice messages is transformed to be more in line with physical intuition. In this way, the emotional transmission between users can be expressed more clearly, and user interaction can be increased in a more interesting way. At the same time, users can enhance the expression of text messages by pinching text messages with both hands and changing the shape of text bubbles. Moreover, by long pressing a text message and dragging it to another text message, text messages can also be merged to change the shape of text bubbles to enhance the expression of text messages. In addition, dragging an emoticon package into a text message or voice message can give the message different forms, such as bubble shape, color, text, sound, emoticon package, etc. In addition, when the user is not speaking, he can directly drag the emoticon package onto the voice control, so that the current buzzwords are intelligently generated according to the meaning of the emoticon package and converted into AI voice to be sent.

[0175] Next, the technical solution of this application is explained from the product side.

[0176] For text messages, in response to a pinch operation on the text message, the chat bubble color, shape, expression, etc. are changed, thereby enhancing the user's emotional expression; or, in response to dragging an expression into the bubble, the chat bubble color, shape, expression, etc. are changed, thereby enhancing the user's emotional expression.

[0177] For a voice message, in response to a drag operation on an emoticon, the emoticon is dragged into the voice message, thereby changing the timbre of the voice; or Figure 5 、 6 As shown in Figure 7, in response to dragging an expression to an empty voice (voice control), a related voice is automatically generated according to the expression, and then when a playback instruction for the voice is received, it is read out intelligently by AI.

[0178] Next, the technical solution of this application is explained from a technical perspective.

[0179] In actual implementation, the technical solution of this application mainly includes three processes, namely, the operation process of users changing the chat bubble style through expressions and gestures, the operation process of super bubble scene barrage floating screen and recovery, and the process of changing bubble emotions through user operations.

[0180] For the user's operation process of changing the chat bubble style through emoticons and gestures, see Figure 12 , Figure 12 This is a flowchart of the operation process of the user changing the chat bubble style through expressions and gestures provided by the embodiment of the present application, based on Figure 12 , the user changes the chat bubble style through expressions and gestures Figure 12 Specifically, you can first enter a normal text message, and then drag the emoticon to the chat bubble of the text message to change the chat bubble style, or pinch adjacent text messages to change the chat bubble style.

[0181] For inputting ordinary text messages, first, in response to the message sender's sending operation for the input chat conversation text, the chat conversation text input by the message sender is sent, wherein the sending client, that is, the client corresponding to the message sender, sends the chat conversation text to the server, so that after the server receives the chat conversation text, it generates a unique message identifier, writes the chat conversation text to the storage, and pushes the chat conversation text to the receiving client, that is, the client corresponding to the message recipient. After the chat conversation text is successfully displayed on the receiving client, the sending client receives the unique message identifier of the chat conversation text sent by the server, and displays the chat bubble and text content corresponding to the chat conversation text on the session interface.

[0182] Regarding the process of changing the chat bubble style by dragging an emoticon into the chat bubble of the above-mentioned chat conversation text, in response to the message sender dragging the system emoticon into the chat bubble of the chat conversation text, the sending client sends a chat mood change request carrying the emoticon dragged by the message sender and the identifier of the chat conversation text to the server. The server receives the chat mood change request, determines the emotion expressed by the emoticon, and calculates a new bubble style result in combination with the current emotion of the message, updates it to the cloud storage, and then pushes a style change result of the message to the receiving client, that is, the client corresponding to the message recipient. After successfully displaying the style change result on the receiving client, the sending client receives the new bubble style of the chat conversation text sent by the server, thereby displaying the new chat bubble and text content corresponding to the chat conversation text on the session interface.

[0183] Regarding the process of changing the chat bubble style by pinching adjacent text messages by gesture, in response to the message sender's pinching operation on the chat conversation text and the text messages adjacent to the chat conversation text, the sending client extracts the message identifier of the pinched message, and sends a chat mood change request carrying a list of identifiers of messages associated with the pinching operation (if emoticons are transmitted together), to the server. The server receives the chat mood change request, analyzes the list of pinched identifiers or system emoticons in the request, calculates the new bubble style result according to the number of messages and the meaning of the emoticons, updates it to the cloud storage, and then pushes a style change result of the message to the receiving client, that is, the client corresponding to the message recipient. After the style change result is successfully displayed on the receiving client, the sending client receives the new bubble style of the chat conversation text sent by the server, thereby displaying the new chat bubble and text content corresponding to the chat conversation text on the session interface.

[0184] For the operation process of floating and restoring the super bubble scene bullet screen, see Figure 13 , Figure 13 This is a flow chart of the operation process of floating and restoring the super bubble scene barrage provided by the embodiment of the present application, based on Figure 13 The super bubble scene barrage floating screen and recovery operation process is through Figure 13 Specifically, when the changed chat bubble is a super bubble such as a screen-dominating bullet screen, when the message sender continues to send messages, the first to fourth messages will be displayed in the super bubble on the screen, and will also be displayed in the super bubble of the recipient's client interface on the screen; when the fifth message is sent, the super bubble is restored, that is, the messages included in the super bubble are restored to separately displayed messages and displayed in the style of ordinary bubbles.

[0185] For the process of changing the bubble's mood through user operations, see Figure 14 , Figure 14 This is a flow chart of the process of changing the bubble emotion through user operation provided by the embodiment of the present application, based on Figure 14 The process of changing the bubble mood through user operation in this application is Figure 14 Specifically, in the conversation interface between the message sender and the message recipient, in response to the message sender's sending operation on the chat text, the sent chat text is displayed in the conversation interface, and then in response to the message sender dragging the angry expression to the chat bubble, an angry emotion bubble is generated, that is, based on the angry emotion bubble, the sent chat text is displayed, and then when the message recipient replies to the chat text, in response to the message sender dragging the happy expression to the angry emotion bubble, the angry emotion bubble is changed to a happy emotion bubble, and then based on the happy emotion bubble, the sent chat text is displayed.

[0186] In some embodiments, when emotions cannot be directly identified, chat bubble forms with different emotions can be generated by simply determining keywords, key punctuation marks, user-defined emotions, etc.; at the same time, when emotions are recognized through voice and voice is converted into text, user emotions can be judged to generate different emotional chat bubble forms.

[0187] In this way, through this application, the user's true emotions can be understood to better convey emotional information to other users, allowing users to feel the other party's emotions and semantics more directly, while increasing the fun of the chat.

[0188] Using the above-described embodiments of the present application, a voice control and at least one conversational expression are first displayed on the conversation interface. Then, the voice control is fused with a first conversational expression from the at least one conversational expression to generate a first fused voice message based on the first conversational expression. In this way, by fusing the voice control and the conversational expression, a voice message is generated, thereby increasing the diversity of conversational messages during the conversation.

[0189] The following continues to describe the exemplary structure of the session message processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the session message processing device 455 of the memory 450 may include:

[0190] Display module 4551, used to display the voice control and at least one conversation emoticon in the conversation interface;

[0191] The fusion module 4552 is configured to, in response to a fusion operation on the voice control and a first conversation expression in the at least one conversation expression, display a first fused voice generated based on the first conversation expression.

[0192] In some embodiments, the conversation interface includes a message editing area, and the voice control is displayed in the message editing area; the fusion module 4552 is further used to, in response to a drag operation on the first conversation emoticon, drag the first conversation emoticon to the voice control in the message editing area; in response to a release operation on the first conversation emoticon, display a first fused voice generated based on the first conversation emoticon in the message editing area.

[0193] In some embodiments, the device further includes a switching module, which is used to display a switching control in an associated area of the first fused voice and automatically play the first fused voice; wherein the switching control is used to switch the content of the first fused voice; in response to a triggering operation on the switching control, the content of the first fused voice is switched to new content, and the new content is associated with the first conversational expression.

[0194] In some embodiments, the first fused speech has a first speech feature, and the device further includes a second display module, which is used to display a second fused speech with a second speech feature in response to a fusion operation on the first fused speech and a second conversational expression in the at least one conversational expression; wherein the second speech feature is different from the first speech feature, and the second speech feature corresponds to the second conversational expression.

[0195] In some embodiments, the conversation interface includes a message display area and a message editing area, and the first fused voice is displayed in the message editing area; the device also includes a first sending module, and the first sending module is used to send the first fused voice from the message editing area to the message display area in response to a sending instruction for the first fused voice; the second display module is also used to display a second fused voice with second voice features in response to a fusion operation of the first fused voice and the second conversation expression in the message display area.

[0196] In some embodiments, the fusion module 4552 is further used to, in response to a fusion operation on the voice control and the first conversational expression in the at least one conversational expression, display the first fused voice generated based on the first conversational expression in the associated area of the voice control; or, in response to a fusion operation on the voice control and the first conversational expression in the at least one conversational expression, cancel the display of the voice control and display the first fused voice generated based on the first conversational expression at the position of the voice control.

[0197] In some embodiments, the device also includes a second fusion module, which is used to switch the first fused voice to a third fused voice in response to a fusion operation on the voice control and the third conversation expression in the at least one conversation expression; wherein the content of the third fused voice includes content associated with the first conversation expression and the third conversation expression.

[0198] In some embodiments, the third fused voice is carried in a message bubble, which includes a third conversation expression; the device also includes a deletion module, which is used to switch the third fused voice to the first fused voice in response to a deletion operation on the third conversation expression.

[0199] In some embodiments, the third fused voice is carried in a message bubble, which includes a first conversational expression and a third conversational expression. The first conversational expression and the third conversational expression are in a draggable state and have a first sort. The device also includes a sorting adjustment module, which is used to respond to a sorting adjustment operation on the conversational expression in the third fused voice, adjust the first sort to the second sort, and switch the third fused voice to a fourth fused voice, and the content of the fourth fused voice is different from that of the third fused voice.

[0200] In some embodiments, the device further includes a playback module, which is configured to obtain voice features corresponding to the first conversational expression in response to a playback instruction for the first fused speech; and play the first fused speech based on the voice features corresponding to the first conversational expression.

[0201] In some embodiments, there are multiple conversational expressions, and the multiple conversational expressions belong to at least two expression categories, and different expression categories correspond to different voice features; the playback module is also used to respond to the playback instruction for the first fused speech, determine the target voice features corresponding to the target expression category to which the first conversational expression belongs; and use the target voice features corresponding to the target expression category to which the first conversational expression belongs to play the first fused speech.

[0202] In some embodiments, the playback module is further configured to obtain a feature value of the target voice feature associated with the first conversational expression, where different conversational expressions have different feature values of voice features; and play the first fused voice based on the feature value of the target voice feature.

[0203] In some embodiments, the first fused speech is carried in a message bubble, the message bubble includes the first conversational expression, and the at least one conversational expression includes the first conversational expression and a fourth conversational expression. The device further includes a third fusion module, and the third fusion module is configured to, in response to a fusion operation on the first fused speech and the fourth conversational expression, compare the expression category to which the fourth conversational expression belongs with the target expression category; when the comparison result indicates that the expression category to which the fourth conversational expression belongs is consistent with the target expression category, display a fourth fused speech with a third speech feature; wherein the feature value of the third speech feature is the sum of the feature value corresponding to the fourth conversational expression and the feature value of the target speech feature; when the comparison result indicates that the expression category to which the fourth conversational expression belongs is inconsistent with the target expression category, display a fifth fused speech with a fourth speech feature; wherein the fourth speech feature is different from the target speech feature, and the fourth speech feature corresponds to the fourth conversational expression.

[0204] In some embodiments, the conversation interface includes a message display area and a message editing area, and the first fused voice is displayed in the message editing area; the device also includes a third display module, and the third display module is used to respond to a sending instruction for the first fused voice and send the first fused voice from the message editing area to the message display area; in the message display area, the first fused voice is displayed through a message bubble.

[0205] In some embodiments, the display module 4551 is further used to display a first fused voice generated based on the first conversational expression when the number of conversational expressions fused for the voice control is between a first number and a second number; the device also includes a prompt module, which is used to display a prompt message when the number of conversational expressions fused for the voice control is greater than the second number, and the prompt message is used to prompt that the number of expressions fused for the voice control has reached a threshold and the fusion of the voice control and the first conversational expression cannot be performed.

[0206] In some embodiments, the fusion operation includes a voice recording operation and an expression selection operation; the fusion module 4552 is also used to perform voice recording in response to the voice recording operation triggered based on the voice control; based on the recorded voice, in response to the expression selection operation for the first conversation expression in the at least one conversation expression, display the first fused voice generated based on the first conversation expression.

[0207] In some embodiments, the display module 4551 is further used to display a voice control and at least one conversation expression distributed around the voice control in a conversation interface; the device also includes a dragging module, which is used to display the process of the voice control being dragged in response to a drag operation on the voice control; when the voice control is dragged to the first conversation expression and released, the release operation on the voice control is determined as the expression selection operation.

[0208] In some embodiments, the device further includes a sliding module, which is used to display a sliding trajectory of the sliding operation in response to the sliding operation triggered by the voice control; when the end point of the sliding trajectory touches the first conversation expression, the expression selection operation is received in response to the voice control being released.

[0209] In some embodiments, the fusion module 4552 is further used to respond to a fusion operation on the voice control and the first conversation expression in the at least one conversation expression, and if the current account has the fusion permission for the voice control and the first conversation expression, display a first fusion voice generated based on the first conversation expression.

[0210] In some embodiments, the device also includes a message processing module, which is used to obtain a message processing method corresponding to the first conversational expression in response to a playback instruction for the first fusion voice; play the first fusion voice, and use the message processing method corresponding to the first conversational expression to process the first fusion voice that has been played.

[0211] In some embodiments, the conversation interface includes a message display area, and the first fused voice is displayed in the message display area; the message processing module is also used to, when the message processing method corresponding to the first conversation expression is a read-and-destroy method, display the process of the first fused voice disappearing from the message display area when the first fused voice is played; when the message processing method corresponding to the first conversation expression is a loop playback method, obtain the loop period and number of loops corresponding to the first conversation expression; after the first fused voice is played, play the first fused voice again based on the loop period and number of loops.

[0212] An embodiment of the present application provides a computer program product comprising computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the method for processing conversation messages described in the embodiment of the present application.

[0213] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the method for processing a session message provided in the embodiment of the present application, for example, Figure 3 The method for processing session messages is shown.

[0214] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disk, or a CD-ROM; or various devices including one or any combination of the above memories.

[0215] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0216] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0217] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0218] In summary, the embodiments of the present application have the following beneficial effects:

[0219] By integrating voice controls and conversational expressions, voice messages are generated, thereby increasing the diversity of conversational messages during the conversation.

[0220] It should be noted that in the embodiments of the present application, when obtaining conversation messages such as emoticons, user operation data and other related data, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0221] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A method for processing a conversation message, characterized in that: The method comprises: In the conversation interface, display a voice control and at least one conversation emoticon; In response to a fusion operation on the voice control and a first conversational expression among the at least one conversational expression, a first fused voice generated based on the first conversational expression is displayed.

2. The method according to claim 1, wherein The conversation interface includes a message editing area, and the voice control is displayed in the message editing area; and in response to a fusion operation on the voice control and a first conversation expression of the at least one conversation expression, displaying a first fused voice generated based on the first conversation expression, comprises: In response to a drag operation on the first conversation emoticon, dragging the first conversation emoticon to the voice control in the message editing area; In response to a release operation on the first conversation expression, a first fused voice generated based on the first conversation expression is displayed in the message editing area.

3. The method according to claim 1, wherein After displaying the first fused speech generated based on the first conversation expression, the method further includes: Displaying a switching control in an area associated with the first fused voice and automatically playing the first fused voice; Wherein, the switching control is used to switch the content of the first fused voice; In response to a triggering operation on the switching control, the content of the first fused speech is switched to new content, where the new content is associated with the first conversation expression.

4. The method according to claim 1, wherein The first fused speech has a first speech feature. After displaying the first fused speech generated based on the first conversation expression, the method further includes: In response to a fusion operation on the first fused speech and a second conversational expression in the at least one conversational expression, displaying a second fused speech having a second speech feature; The second voice feature is different from the first voice feature, and the second voice feature corresponds to the second conversation expression.

5. The method according to claim 4, wherein The conversation interface includes a message display area and a message editing area, and the first fused voice is displayed in the message editing area; After displaying the first fused speech generated based on the first conversation expression, the method further includes: In response to a sending instruction for the first fused voice, sending the first fused voice from the message editing area to the message display area; The displaying of a second fused speech having a second speech feature in response to a fusion operation on the first fused speech and a second conversation expression in the at least one conversation expression includes: In response to a fusion operation on the first fused voice and the second conversation emoticon in the message display area, a second fused voice having a second voice feature is displayed.

6. The method according to claim 1, wherein The step of displaying a first fused voice generated based on the first conversation expression in response to a fusion operation on the voice control and the first conversation expression in the at least one conversation expression includes: In response to a fusion operation on the voice control and a first conversation expression in the at least one conversation expression, displaying the first fused voice generated based on the first conversation expression in an associated area of the voice control; or In response to a fusion operation on the voice control and a first conversational expression in the at least one conversational expression, the voice control is canceled from being displayed, and the first fused voice generated based on the first conversational expression is displayed at the location of the voice control.

7. The method according to claim 1, wherein After displaying the first fused speech generated based on the first conversation expression, the method further includes: In response to a fusion operation on the voice control and a third conversation expression in the at least one conversation expression, switching the first fused voice to a third fused voice; The content of the third fused voice includes content associated with the first conversation expression and the third conversation expression.

8. The method according to claim 7, wherein The third fused voice is carried in a message bubble, and the message bubble includes a third conversation emoticon; after switching the first fused voice to the third fused voice, the method further includes: In response to a deletion operation on the third conversation emoticon, the third fused voice is switched to the first fused voice.

9. The method according to claim 7, wherein: The third fused voice is carried in a message bubble, which includes the first conversation emoticon and the third conversation emoticon, wherein the first conversation emoticon and the third conversation emoticon are in a draggable state and have a first order; After switching the first fused voice to the third fused voice, the method further includes: In response to the sorting adjustment operation on the conversational expressions in the third fused speech, the first sorting is adjusted to the second sorting, and the third fused speech is switched to a fourth fused speech, the content of the fourth fused speech being different from that of the third fused speech.

10. The method according to claim 1, wherein After displaying the first fused speech generated based on the first conversation expression, the method further includes: In response to a play instruction for the first fused speech, obtaining a speech feature corresponding to the first conversational expression; The first fused voice is played based on the voice feature corresponding to the first conversation expression.

11. The method according to claim 10, wherein There are multiple conversation expressions, each of which belongs to at least two expression categories, and different expression categories correspond to different voice features; The acquiring, in response to the play instruction for the first fused speech, a speech feature corresponding to the first conversational expression, includes: In response to a play instruction for the first fused speech, determining a target speech feature corresponding to a target expression category to which the first conversational expression belongs; Playing the first fused voice based on the voice feature corresponding to the first conversation expression includes: The first fused voice is played using a target voice feature corresponding to a target expression category to which the first conversational expression belongs.

12. The method according to claim 11, wherein Playing the first fused speech using a target speech feature corresponding to a target expression category to which the first conversation expression belongs includes: Obtaining a feature value of the target voice feature associated with the first conversational expression, where different voice features are associated with different conversational expressions; Based on the feature value of the target speech feature, the first fused speech is played.

13. The method according to claim 12, wherein: The first fused speech is carried in a message bubble, the message bubble includes the first conversation emoticon, the at least one conversation emoticon includes the first conversation emoticon and a fourth conversation emoticon, and after displaying the first fused speech generated based on the first conversation emoticon, the method further includes: In response to a fusion operation on the first fused speech and the fourth conversation expression, comparing the expression category to which the fourth conversation expression belongs with the target expression category; When the comparison result indicates that the expression category to which the fourth conversational expression belongs is consistent with the target expression category, a fourth fused speech having a third speech feature is displayed; wherein the feature value of the third speech feature is the sum of the feature value corresponding to the fourth conversational expression and the feature value of the target speech feature; When the comparison result indicates that the expression category to which the fourth conversational expression belongs is inconsistent with the target expression category, a fifth fused speech having a fourth speech feature is displayed; wherein the fourth speech feature is different from the target speech feature, and the fourth speech feature corresponds to the fourth conversational expression.

14. The method according to claim 1, wherein The conversation interface includes a message display area and a message editing area, and the first fused voice is displayed in the message editing area; After displaying the first fused speech generated based on the first conversation expression, the method further includes: In response to a sending instruction for the first fused voice, sending the first fused voice from the message editing area to the message display area; In the message display area, the first fused voice is displayed through a message bubble.

15. The method according to claim 1, wherein The displaying of the first fused speech generated based on the first conversation expression includes: When the number of conversation expressions fused for the voice control is between a first number and a second number, displaying a first fused voice generated based on the first conversation expression; The method further comprises: When the number of conversational expressions to be fused for the voice control is greater than the second number, a prompt message is displayed, wherein the prompt message is used to prompt that the number of expressions to be fused for the voice control has reached a threshold and the fusion of the voice control and the first conversational expression cannot be performed.

16. The method according to claim 1, wherein The fusion operation includes a voice recording operation and an expression selection operation; The step of displaying a first fused voice generated based on the first conversation expression in response to a fusion operation on the voice control and the first conversation expression in the at least one conversation expression includes: performing voice recording in response to a voice recording operation triggered by the voice control; Based on the recorded speech, in response to an expression selection operation on a first conversation expression among the at least one conversation expression, a first fused speech generated based on the first conversation expression is displayed.

17. The method according to claim 16, wherein The method of displaying a voice control and at least one conversation expression in a conversation interface includes: In the conversation interface, displaying a voice control and at least one conversation emoticon distributed around the voice control; After the voice recording, the method further includes: In response to a drag operation on the voice control, displaying a process of the voice control being dragged; When the voice control is dragged to the first conversation expression and then released, the release operation on the voice control is determined as the expression selection operation.

18. The method according to claim 17, wherein After the voice recording, the method further includes: In response to a sliding operation triggered by the voice control, displaying a sliding track of the sliding operation; When the end point of the sliding track reaches the first conversation expression, in response to the voice control being released, the expression selection operation is received.

19. The method according to claim 1, wherein The step of displaying a first fused voice generated based on the first conversation expression in response to a fusion operation on the voice control and the first conversation expression in the at least one conversation expression includes: In response to a fusion operation on the voice control and a first conversation expression among the at least one conversation expression, if the current account has fusion permission for the voice control and the first conversation expression, a first fused voice generated based on the first conversation expression is displayed.

20. The method of claim 1, wherein After displaying the first fused speech generated based on the first conversation expression, the method further includes: In response to a play instruction for the first fused voice, obtaining a message processing method corresponding to the first conversation emoticon; The first fused voice is played, and the played first fused voice is processed using a message processing method corresponding to the first conversational expression.

21. The method according to claim 20, wherein The conversation interface includes a message display area, and the first fused voice is displayed in the message display area; and the processing of the played first fused voice using a message processing method corresponding to the first conversation expression includes: When the message processing mode corresponding to the first conversation emoticon is the self-destructing mode, after the first fused voice is played, the process of the first fused voice disappearing from the message display area is displayed; When the message processing mode corresponding to the first conversation expression is a loop playback mode, obtaining a loop period and a loop count corresponding to the first conversation expression; After the first fused voice is played, the first fused voice is played again based on the loop period and the number of loops.

22. A device for processing conversation messages, characterized in that: The device comprises: A display module, configured to display a voice control and at least one conversation emoticon in the conversation interface; The fusion module is configured to, in response to a fusion operation on the voice control and a first conversation expression in the at least one conversation expression, display a first fused voice generated based on the first conversation expression.

23. An electronic device, characterized in that: include: a memory for storing computer-executable instructions; The processor is configured to implement the method for processing a conversation message according to any one of claims 1 to 21 when executing the computer-executable instructions stored in the memory.

24. A computer-readable storage medium, characterized in that Computer executable instructions are stored, which are used to cause a processor to execute and implement the method for processing a conversation message according to any one of claims 1 to 21.

25. A computer program product comprising computer executable instructions, characterized in that When the computer executable instructions are executed by a processor, the method for processing a conversation message according to any one of claims 1 to 21 is implemented.