Session message processing method and device, electronic equipment, computer readable storage medium and computer program product

By integrating voice messages and conversation expressions in the conversation interface, changing the style and characteristics of voice messages, the problem of single conversation process is solved, and the diversity of conversation messages and the improvement of user experience is achieved.

CN120455423APending Publication Date: 2025-08-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410175786.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art has a relatively single conversation process in the conversation interface, and the experience is dull and boring, and lacks diversity.

Method used

By displaying the fusion operation of voice messages and conversation expressions in the conversation interface, the message style and voice characteristics of voice messages are changed to correspond to the target conversation expressions, and diversification of voice messages is achieved.

Benefits of technology

Improves the diversity of conversation messages during the conversation process and enhances the fun and interactiveness of the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455423A_ABST
    Figure CN120455423A_ABST
Patent Text Reader

Abstract

The invention provides a session message processing method and device, electronic equipment, a computer readable storage medium and a computer program product, and the method comprises the steps: employing a first message style in a session interface, displaying a voice message with a first voice feature, and displaying at least one session expression; in response to a fusion operation for the voice message and a target session expression in the at least one session expression, displaying a fused voice message; wherein the fused voice message satisfies at least one of the following conditions: the message style of the fused voice message is changed from a first message style to a second message style, and the voice feature of the fused voice message is changed from a first voice feature to a second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target session expression. According to the invention, the diversity of the session messages in the session process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for processing conversation messages. Background Art

[0002] When conducting conversations based on a conversation interface, related technologies mostly directly send a voice message or an emoticon message, which results in a relatively simple conversation process and a dull and boring experience. Summary of the Invention

[0003] Embodiments of the present application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing conversation messages, which can improve the diversity of conversation messages during a conversation.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] This embodiment of the present application provides a method for processing a session message, including:

[0006] In the conversation interface, a first message style is used to display a voice message having a first voice feature and at least one conversation emoticon;

[0007] In response to a fusion operation on the voice message and a target conversation emoticon in the at least one conversation emoticon, displaying a fused voice message;

[0008] The fused voice message satisfies at least one of the following conditions: the message style of the fused voice message changes from the first message style to the second message style, and the voice feature of the fused voice message changes from the first voice feature to the second voice feature;

[0009] The second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversation expression.

[0010] An embodiment of the present application provides a device for processing a session message, including:

[0011] A display module, configured to display, in a conversation interface, a voice message having a first voice feature using a first message style and displaying at least one conversation emoticon;

[0012] A fusion module is used to display a fused voice message in response to a fusion operation on the voice message and a target conversation expression in the at least one conversation expression; wherein the fused voice message satisfies at least one of the following: the message style of the fused voice message is changed from the first message style to the second message style, and the voice feature of the fused voice message is changed from the first voice feature to the second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversation expression.

[0013] In the above scheme, the conversation interface includes a message display area and a message editing area, the voice message is displayed in the message display area, and the target conversation emoticon is displayed in the message editing area; the fusion module is also used to respond to a drag operation on the target conversation emoticon, drag the target conversation emoticon from the message editing area to the voice message in the message display area; in response to a release operation on the target conversation emoticon, display a fused voice message obtained by fusing the target conversation emoticon and the voice message.

[0014] In the above scheme, the device also includes a switching module, which is used to display a switching control in the associated area of the fused voice message and automatically play the fused voice message; wherein the switching control is used to switch the message style and at least one of the voice features of the fused voice message; in response to a trigger operation on the switching control, the fused voice message is switched to an associated fused voice message, and the associated fused voice message is associated with the target conversation expression.

[0015] In the above scheme, there are multiple conversational expressions, and the multiple conversational expressions belong to at least two expression categories, and different expression categories correspond to different voice features; the device also includes a first playback module, which is used to respond to the playback instruction for the fused voice message and use the target voice features corresponding to the target expression category to which the target conversational expression belongs to play the fused voice message.

[0016] In the above scheme, the first playback module is also used to determine the target voice feature corresponding to the target expression category to which the target conversational expression belongs; obtain the feature value of the target voice feature associated with the target conversational expression, and the feature values of the voice feature associated with conversational expressions of different expression categories are different; and play the fused voice message based on the feature value of the target voice feature.

[0017] In the above scheme, the fused voice message is carried in a message bubble, which includes the target conversation expression. The device also includes a comparison module, which is used to compare the expression category to which the other conversation expressions belong and the target expression category in response to the fusion operation on the fused voice message and the other conversation expressions; when the comparison result indicates that the expression category to which the other conversation expressions belong is consistent with the target expression category, a first new fused voice message with a third voice feature is displayed; wherein the feature value of the third voice feature is the sum of the feature value corresponding to the other conversation expression and the feature value of the target voice feature; when the comparison result indicates that the expression category to which the other conversation expressions belong is inconsistent with the target expression category, a second new fused voice message with a fourth voice feature is displayed; wherein the fourth voice feature is different from the second voice feature, and the fourth voice feature corresponds to the other conversation expression.

[0018] In the above scheme, the at least one conversational expression includes the target conversational expression and other conversational expressions, the fused voice message is carried in a message bubble, and the message bubble includes the target conversational expression; the device also includes a second fusion module, which is used to respond to the fusion operation on the fused voice message and the other conversational expressions, and display a third new fused voice message obtained by replacing the target conversational expression fused in the message bubble with the other conversational expressions; wherein the third new fused voice message includes a fifth voice feature, the fifth voice feature is different from the second voice feature, and the fifth voice feature corresponds to the other conversational expressions.

[0019] In the above scheme, the device also includes a second playback module, which is used to respond to the playback instruction for the fused voice message, obtain the message processing method corresponding to the target conversation expression; play the fused voice message, and use the message processing method corresponding to the target conversation expression to process the fused voice message that has been played.

[0020] In the above scheme, the conversation interface includes a message display area, and the fused voice message is displayed in the message display area; the second playback module is also used to, when the message processing mode corresponding to the target conversation expression is the read-and-destroy mode, display the process of the fused voice message disappearing from the message display area when the fused voice message is played; when the message processing mode corresponding to the target conversation expression is the loop playback mode, obtain the loop period and number of loops corresponding to the target conversation expression; after the fused voice message is played, play the fused voice message again based on the loop period and number of loops.

[0021] In the above scheme, the device also includes a third playback module, which is used to play the fused voice message in response to a playback instruction for the fused voice message, and play the animation effects associated with the target conversation expression during the playback of the fused voice message.

[0022] In the above scheme, the device also includes a text conversion module, which is used to present a text display area including the target conversation emoticon in response to a text conversion instruction for the fused voice message; in the text display area, the text content corresponding to the fused voice message is presented.

[0023] In the above scheme, the conversation interface includes a message display area, and the fused voice message is displayed in the message display area; the device also includes a fourth playback module, which is used to respond to a playback instruction for the fused voice message, play the fused voice message, and when the fused voice message is played, convert the target conversation expression into a barrage, and fix the barrage to the target position of the message display area; when the display time of the barrage reaches the first target time, the barrage is canceled.

[0024] In the above scheme, the fused voice message is sent by the target object; the device also includes a reply module, which is used to float and display reply barrages of other objects to the barrages above the barrage; wherein, the other objects are the objects with which the target object conducts conversations based on the conversation interface.

[0025] In the above scheme, the conversation interface also includes a message editing area, and the device also includes an editing module, which is used to display the edited message content in response to the message editing operation triggered based on the message editing area; in response to the sending instruction for the edited message content, the edited message content is displayed in the barrage at the target location.

[0026] In the above scheme, the device also includes a restoration module, which is used to display the process of restoring the fused voice message to obtain the voice message and the target conversation expression in response to the restoration condition of the fused voice message being met; wherein the restoration condition includes at least one of the following: the fused voice message is played; the display time of the fused voice message reaches the second target time; a restoration instruction for the fused voice message is received, and the restoration instruction is used to instruct the fused voice message to be restored to the voice message and the target conversation expression.

[0027] In the above scheme, the conversation interface includes a message display area, and the voice message and the target conversation expression are displayed in the message display area; the fusion module is further used to display the process of merging the voice message and the target conversation expression in response to a pinch operation on the voice message and the target conversation expression; when the ratio of the first area to the second area reaches a target ratio, the voice message and the target conversation expression are canceled; wherein, the first area is the area of the overlapping area of the voice message and the target conversation expression, and the second area is the display area of the voice message or the display area of the target conversation expression; the fused voice message obtained by fusing the voice message and the target conversation expression is displayed.

[0028] An embodiment of the present application provides an electronic device, including:

[0029] a memory for storing computer-executable instructions;

[0030] The processor is configured to implement the method for processing the session message provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.

[0031] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute instructions to implement a method for processing a session message provided in an embodiment of the present application.

[0032] An embodiment of the present application provides a computer program product, comprising computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the method for processing conversation messages provided in an embodiment of the present application.

[0033] The embodiments of the present application have the following beneficial effects:

[0034] First, a first message style is used on a conversation interface to display a voice message having a first voice feature and at least one conversational expression. Then, the voice message and a target conversational expression from the at least one conversational expression are fused to obtain a fused voice message. The fused voice message satisfies at least one of the following conditions: the message style of the fused voice message changes from the first message style to the second message style, and the voice feature of the fused voice message changes from the first voice feature to the second voice feature; the second message style differs from the first message style, the second voice feature differs from the first voice feature, and both the second message style and the second voice feature correspond to the target conversational expression. In this way, by fusing the voice message and the expression, at least one of the voice feature and the message style of the voice message is changed, thereby increasing the diversity of conversational messages during the conversation. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 1 is a schematic diagram of the architecture of a session message processing system 100 provided in an embodiment of the present application;

[0036] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0037] Figure 3 This is a flow chart of a method for processing session messages provided in an embodiment of the present application;

[0038] Figure 4 Schematic diagram of a fused voice message obtained by fusing a voice message and a target conversation expression using a second message style provided by an embodiment of the present application;

[0039] Figure 5 1 is a schematic diagram of a process of displaying a fused voice message obtained by fusing a target conversation expression and a voice message, provided by an embodiment of the present application;

[0040] Figure 6 2 is a schematic diagram of a process for displaying a third new fused voice message provided in an embodiment of the present application;

[0041] Figure 7 Schematic diagram of a process for processing a fused voice message that has been played, provided in an embodiment of the present application;

[0042] Figure 8 This is a flowchart of the operation process of a user changing the chat bubble style through expressions and gestures provided by an embodiment of the present application;

[0043] Figure 9 This is a flowchart of the operation process of floating and restoring the super bubble scene bullet screen provided in an embodiment of the present application;

[0044] Figure 10This is a flowchart of the process of changing bubble emotions through user operations provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0046] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0047] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0049] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0050] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0051] 2) Client, also known as user end, refers to the program corresponding to the server that provides local services to users. Except for some applications that can only run locally, it is generally installed on the terminal and needs to cooperate with the server to run. That is, there must be corresponding servers and service programs on the network to provide corresponding services. In this way, specific communication connections need to be established between the client and server to ensure the normal operation of the application, such as virtual scene clients (such as game clients) and video clients.

[0052] 3) Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0053] 4) Message bubble, also known as chat bubble, refers to a style box that surrounds the text sent by the user in a chat scene.

[0054] 5) Pinch operation, a gesture operation performed on a touch screen device by using two fingers (usually the thumb and index finger) at the same time, is used to zoom in or out on the content on the screen, such as a web page, image or map. When the two fingers are in contact with the screen at the same time and move in a relatively stationary manner, if the fingers are close (approaching each other on the screen), the content on the display screen will be zoomed out; if the fingers are separated (moving away from each other on the screen), the content on the display screen will be zoomed in.

[0055] See also Figure 1 , Figure 1 This is an architectural diagram of a session message processing system 100 provided in an embodiment of the present application, including a terminal (terminal 400 is shown as an example), wherein terminal 400 is connected to server 200 via network 300, wherein network 300 may be a wide area network or a local area network, or a combination of the two, and data transmission is achieved using wireless or wired links.

[0056] The server 200 is configured to send display data of a conversation interface including a voice message displayed in a first message style and having a first voice feature to the terminal 400;

[0057] Terminal 400 is used to receive display data of a conversation interface including a voice message displayed in a first message style and having a first voice feature, and based on the display data, display the voice message with the first voice feature in the conversation interface in the first message style and at least one conversation expression; in response to a fusion operation on the voice message and a target conversation expression in at least one conversation expression, display a fused voice message; wherein the fused voice message satisfies at least one of the following: the message style of the fused voice message changes from the first message style to the second message style, and the voice feature of the fused voice message changes from the first voice feature to the second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversation expression.

[0058] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, an intelligent voice interaction device, a smart home appliance, a virtual reality device, a vehicle-mounted terminal, an aircraft, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device, an intelligent speaker, and a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0059] Next, the electronic device implementing the method for processing session messages provided in the embodiment of the present application is described. Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device can be a server or a terminal. Figure 1 Take the terminal shown in as an example, Figure 2 The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .

[0060] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0061] The user interface 430 includes one or more output devices 431 that enable display of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0062] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0063] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0064] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0065] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0066] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).

[0067] a presentation module 453 for enabling display of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0068] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0069] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 The apparatus 455 for processing conversation messages stored in the memory 450 is shown. This apparatus can be software in the form of a program or plug-in, and includes the following software modules: a display module 4551 and a fusion module 4552. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0070] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus for processing a session message provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the session message processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0071] In some embodiments, the terminal or server can implement the method for processing session messages provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a local client, that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP, a web browser APP; it can also be a small program, that is, a program that can be run only by downloading it into a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of client, module or plug-in.

[0072] Based on the above description of the session message processing system and electronic device provided by the embodiment of the present application, the following describes the session message processing method provided by the embodiment of the present application. In actual implementation, the session message processing method provided by the embodiment of the present application can be implemented by the terminal or the server alone, or by the terminal and the server in collaboration, so that Figure 1 The terminal 400 in the embodiment of the present application alone performs the method for processing the session message as an example for explanation. Figure 3 , Figure 3 This is a flow chart of the method for processing session messages provided by the embodiment of the present application. Next, Figure 3 The steps shown are explained.

[0073] Step 101: The terminal displays a voice message having a first voice feature in a conversation interface using a first message style and displays at least one conversation emoticon.

[0074] In actual implementation, the terminal is provided with a client that supports the processing of conversation messages, such as a video playback client, a browser client, a social client, etc. When the user opens the client on the terminal and the terminal runs the client, the terminal can display a conversation interface for the target object based on the client, and display a voice message with a first voice feature in the conversation interface.

[0075] It should be noted that voice features include at least one of the timbre, pitch, volume, breath, accent, etc. of the voice, and conversation expressions refer to expression messages that can be sent for communication in the conversation interface; at the same time, the voice message can be sent by the target object, or it can be sent by other objects that have a conversation with the target object based on the conversation interface. Moreover, the conversation interface includes a message editing area and a message display area. The message display area is an area where at least two objects having a conversation in the conversation interface display the conversation messages sent. The message editing area is an area where at least two objects having a conversation in the conversation interface edit conversation messages. The voice message is displayed in the message display area, and the conversation expression can be displayed in the message editing area or in the message display area. This is not limited in the embodiments of the present application.

[0076] It should be noted that the message style refers to the display format of conversation messages such as voice messages, text messages, and emoticons in the conversation interface. For example, it can indicate a message bubble or a message card. Different message styles are used to indicate at least one of the shape, color, and size of the message bubble or message card. This embodiment of the present application does not limit this. Among them, the first message style can be pre-set.

[0077] Step 102: Display a fused voice message in response to a fusion operation on a voice message and a target conversational expression in at least one conversational expression; wherein the fused voice message satisfies at least one of the following: a message style of the fused voice message changes from a first message style to a second message style, and a voice feature of the fused voice message changes from a first voice feature to a second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversational expression.

[0078] In actual implementation, after the voice message is fused with the target conversational expression, at least one of the message style and voice features of the voice message will change.

[0079] It should be noted that the fusion operation on the voice message and the target conversation expression in at least one conversation expression refers to the fusion operation on the voice message and at least one conversation expression displayed in the conversation interface, that is, the target conversation expression targeted by the fusion operation is a conversation expression or multiple conversation expressions displayed in the conversation interface, and the fusion operation is to fuse the voice message and the conversation expression or multiple conversation expressions displayed in the conversation interface, that is, the target conversation expression can indicate one conversation expression or multiple conversation expressions, which is not limited in this embodiment of the present application.

[0080] In some embodiments, when a voice feature changes, different conversational expressions may correspond to different voice features. When a fusion operation is received for a voice message and a target conversational expression, the voice feature corresponding to the target conversational expression is obtained, and then the first voice feature of the voice message is switched to the second voice feature corresponding to the target conversational expression to obtain a fused voice message, thereby displaying the fused voice message. For example, when the voice feature is timbre, when a fusion operation is received for a voice message and a target conversational expression, the timbre corresponding to the target conversational expression is obtained, and then the timbre of the voice message is switched to the timbre corresponding to the target conversational expression to obtain a fused voice message. When the voice feature is an accent, when a fusion operation is received for a voice message and a target conversational expression, the accent corresponding to the target conversational expression is obtained, and then the accent of the voice message is switched to the accent corresponding to the target conversational expression to obtain a fused voice message. For example, when the target conversational expression is a monkey expression, when a fusion operation is received for a voice message and a monkey expression, the timbre of Monkey King from Journey to the West corresponding to the monkey expression is obtained, and then the timbre of the voice message is switched to the timbre of Monkey King from Journey to the West.

[0081] In other embodiments, in the case where the message style changes, different conversation expressions may correspond to different message styles. When a fusion operation for a voice message and a target conversation expression is received, the message style corresponding to the target conversation expression is obtained, and then the first message style of the voice message is switched to the second message style corresponding to the target conversation expression, thereby using the second message style to display the fused voice message.

[0082] The message style refers to the display format of conversational messages, such as voice messages and emoticons, in a conversational interface. For example, it can indicate a message bubble or a message card. Different message styles are used to indicate that at least one of the shape, color, and size of the message bubble or message card is different. This embodiment of the present application does not limit this. The first message style can be pre-set; and the second message style is different from the first message style. The message style can include at least one of the shape, color, and size of the message bubble or message card. The second message style can assist in indicating the voice features in the fused voice message. For example, when the target conversational emoticon is a monkey emoticon, the timbre of the voice message switches to the timbre of Monkey King in Journey to the West. The second message style can indicate that the message bubble carried by the fused voice message includes the target conversational emoticon, i.e., the monkey emoticon. In this way, based on the second message style, the timbre in the fused voice message can be determined to be the timbre corresponding to the monkey emoticon.

[0083] For example, see Figure 4 , Figure 4 This is a schematic diagram of a fused voice message obtained by fusion of a voice message and a target conversation expression using a second message style provided by an embodiment of the present application, based on Figure 4 , the conversational interface includes Figure 4 The message display area indicated by 401 and the message editing area indicated by 402 in a are first opened in the conversation interface using the following method: Figure 4 The first message style indicated by 403 in a is displayed as follows Figure 4 Then, in response to the voice message indicated by 403 in a; Figure 4 The target conversation expression indicated by 404 in a and Figure 4 The fusion operation of the voice message indicated in 403 in a is performed as follows Figure 4 The second message style indicated by 405 in b shows the result of fusing the voice message and the target conversation expression, such as Figure 4 The fused voice message indicated by 405 in b.

[0084] In other embodiments, for the case where both the message style and the voice features are changed, the two cases described above are combined, and this embodiment of the present application will not go into details.

[0085] In some embodiments, the message style and / or voice features of the fused voice message can also be switched. Specifically, in response to the fusion operation of the voice message and the target conversation expression in at least one conversation expression, after the fused voice message is displayed, a switching control can be displayed in the associated area of the fused voice message, and the fused voice message can be automatically played; wherein the switching control is used to switch at least one of the message style and voice features of the fused voice message; in response to the triggering operation of the switching control, the fused voice message is switched to an associated fused voice message, and the associated fused voice message is associated with the target conversation expression.

[0086] It should be noted that the switch control can be displayed in an associated area of the fused voice message, such as one of the left area, right area, upper area, and lower area. In addition to automatically playing the first fused voice, the fused voice message can also be played in response to a play instruction for the fused voice message when the switch control is displayed in the associated area of the fused voice message. After the fused voice message is played, if the user is not satisfied with the message style and / or voice characteristics of the fused voice message, the switch control can be triggered to switch the message style and / or voice characteristics of the fused voice message. That is, in response to the triggering operation of the switch control, the fused voice message is switched to an associated fused voice message. The associated fused voice message is associated with the target conversational emoticon.

[0087] It should be noted that the switch control can be set to switch the message style and voice features of the fused voice message, or it can be set to switch the message style of the fused voice message, or it can be set to switch the voice features of the fused voice message. This embodiment of the present application does not limit this.

[0088] Alternatively, the switch control includes a first switch control and a second switch control, the first switch control is used to switch the message style of the fused voice message, and the second switch control is used to switch the voice features of the fused voice message. After the fused voice message is played, if the user is not satisfied with the message style or voice features of the fused voice message, a trigger operation can be executed on the first switch control or the second switch control to switch the corresponding message style or voice features of the fused voice message. This embodiment of the present application does not limit this.

[0089] In actual implementation, the fusion operation may include a pinch operation on the voice message and the target conversation emoticon, or the fusion operation may also include a drag operation and a release operation on the voice message and the target conversation emoticon.

[0090] In some embodiments, when the fusion operation includes a drag operation and a release operation on a voice message and a target conversation emoticon, the conversation interface includes a message display area and a message editing area, the voice message is displayed in the message display area, and the target conversation emoticon is displayed in the message editing area; thus, in response to the fusion operation on the voice message and the target conversation emoticon in at least one conversation emoticon, the process of displaying the fused voice message may be, in response to the drag operation on the target conversation emoticon, dragging the target conversation emoticon from the message editing area to the voice message in the message display area; in response to the release operation on the target conversation emoticon, displaying the fused voice message obtained by fusing the target conversation emoticon and the voice message.

[0091] For example, see Figure 5 , Figure 5 This is a schematic diagram of a process of displaying a fused voice message obtained by fusing a target conversation expression and a voice message, provided by an embodiment of the present application, based on Figure 5 , the conversational interface includes Figure 5 The message display area indicated by 501 and the message editing area indicated by 502 in a, when the fusion operation includes the following Figure 5 The voice message indicated by 503 in a and Figure 5 In response to the drag operation on the target conversation expression indicated by 504 in a and the release operation, the target conversation expression is dragged from the message editing area to the voice message in the message display area in response to the drag operation on the target conversation expression; in response to the release operation on the target conversation expression, the result obtained by merging the target conversation expression and the voice message is displayed. Figure 5 The fused voice message indicated by 505 in b.

[0092] It should be noted that the target conversation emoticon can be displayed in the message editing area or in the message display area. When the target conversation emoticon is displayed in the message display area, the fusion operation can still include a drag operation and a release operation for the voice message and the target conversation emoticon, thereby dragging the target conversation emoticon from the message display area to the voice message in the message display area for fusion. The fusion process here is similar to that when the target conversation emoticon is displayed in the message editing area. This embodiment of the present application does not elaborate on this. Among them, the voice message is sent by the target object corresponding to the terminal, and when the target conversation message is displayed in the message display area, the target conversation emoticon can be sent by the target object or by other objects with which the target object conducts a conversation based on the conversation interface. This embodiment of the present application does not limit this.

[0093] In other embodiments, when the fusion operation includes a drag operation and a release operation on the voice message and the target conversation expression, the conversation interface includes a message display area, and the voice message and the target conversation expression are displayed in the message display area; thus, in response to the fusion operation on the voice message and the target conversation expression in at least one conversation expression, the process of displaying the fused voice message may be, in response to a pinch operation on the voice message and the target conversation expression, displaying the process of merging the voice message and the target conversation expression; when the ratio of the first area to the second area reaches a target ratio, the display of the voice message and the target conversation expression is cancelled; wherein the first area is the area of the overlapping area of the voice message and the target conversation expression, and the second area is the display area of the voice message or the display area of the target conversation expression; and the fused voice message obtained by fusing the voice message and the target conversation expression is displayed.

[0094] It should be noted that the target ratio can be pre-set; at the same time, the voice message and the target conversation expression can be carried in the message bubble or message card, and the display area of the voice message and the target conversation expression is used to indicate the area occupied by the corresponding message bubble or message card, and the overlapping area of the voice message and the target conversation expression also refers to the overlapping area between the corresponding message bubbles or message cards.

[0095] In actual implementation, when the voice message and the target conversation expression are displayed in the message display area, in addition to fusing the voice message and the target conversation expression through a fusion operation, the voice message and the target conversation expression can also be automatically fused. Specifically, in the conversation interface, the continuously received voice messages and target conversation expressions, as well as the sending time of the voice messages and the target conversation expressions are displayed; in response to the difference between the sending time of the voice message and the target conversation expression being less than a first difference threshold, when the emotion represented by the voice message and the emotion represented by the target conversation expression belong to the same emotion category, the process of fusing the voice message and the target conversation expression is displayed, and the fused voice message obtained by fusing the voice message and the target conversation expression is displayed.

[0096] It should be noted that if a voice message and a target conversation emoticon are sent to the same recipient and no other messages are received between them, the voice message and the target conversation emoticon are considered received consecutively. The term "received" here refers to the client on the terminal. Therefore, it can be a voice message and a target conversation emoticon sent by the target recipient corresponding to the current terminal, and thus received by the client. The term "no other messages received" between the voice message and the target conversation emoticon refers to messages sent by the target recipient themselves, as well as messages sent by other recipients of the conversation with the target recipient using the conversation interface.

[0097] It should be noted that the first difference threshold may be preset, for example, 1 second or 2 seconds. Meanwhile, the process of determining whether the emotion represented by the voice message and the emotion represented by the target conversational expression belong to the same emotion category may include first determining the emotion represented by the voice message and the emotion represented by the target conversational expression. Specifically, the voice message is converted into a text message, and then emotion keywords of the text message are obtained. The emotion keywords are used to indicate the emotion represented by the text message, and may be words that directly indicate emotions, such as sad, happy, crying, and angry, or emoticons such as TT and QAQ, for example, TT and QAQ can indicate crying, or punctuation marks such as ???, for example, ???. It can indicate doubts; then when the emotional keywords are successfully obtained, the mapping relationship between the pre-set emotional keywords and emotions is obtained, and based on the obtained emotional keywords and the mapping relationship, the emotion represented by the text message, that is, the emotion represented by the voice message, is determined; when the emotional keywords are not obtained, the pre-trained emotion recognition model is called to perform emotion recognition on the text message to obtain the emotion represented by the text message, that is, the emotion represented by the voice message; and for the process of determining the emotion represented by the target conversational expression, each conversational expression will have a corresponding emotion pre-set, and when the conversational expression is determined, the emotion corresponding to the conversational expression is also determined. Among them, after determining the emotion represented by the voice message and the emotion represented by the target conversational expression, it can be determined whether the emotion represented by the voice message and the emotion represented by the target conversational expression are of the same emotion category;

[0098] It should be noted that the emotion categories can be pre-set, for example, they can be classified based on the tendency of emotions, such as positive, negative, etc., that is, positive emotions are one category and negative emotions are another category, or they can be divided based on specific emotions, such as sad, happy, angry, etc., which are not limited in this embodiment of the present application. The correspondence between specific emotions and emotion categories can also be pre-set. After determining the emotion represented by the voice message and the emotion represented by the target conversation expression, that is, determining the emotion category to which the emotion represented by the voice message belongs and the emotion category to which the target conversation expression belongs, it is determined whether the emotion represented by the voice message and the emotion represented by the target conversation expression belong to the same emotion category.

[0099] In some embodiments, as described above, different conversational expressions correspond to different voice features. In addition, different categories of conversational expressions may correspond to different voice features. Specifically, there are multiple conversational expressions, and the multiple conversational expressions belong to at least two expression categories, and different expression categories correspond to different voice features. Thus, after displaying the fused voice message, the fused voice message can be played in response to a play instruction for the fused voice message using the target voice feature corresponding to the target expression category to which the target conversational expression belongs.

[0100] It should be noted that, before playing the fused voice message, the target speech feature corresponding to the target expression category to which the target conversational expression belongs is used, and first, the target expression category to which the target conversational expression belongs needs to be determined. Here, the correspondence between the conversational expression and the expression category is preset, that is, after the target conversational expression is determined, the expression category to which the target conversational expression belongs can be determined based on the preset correspondence. At the same time, as mentioned above, the speech feature includes at least one of pitch, timbre, volume, breath and accent. Different conversational expressions correspond to different speech features. For example, some conversational expressions correspond to timbre, some conversational expressions correspond to volume, and some conversational expressions correspond to timbre and volume. For example, for conversational expressions A, B, C, and D, these four conversational expressions correspond to different speech features. The speech feature of conversational expression A is timbre, the speech feature of conversational expression B is timbre, the speech feature of conversational expression C is timbre and volume, and the speech feature of conversational expression D is volume.

[0101] Different expression categories correspond to different voice features. For example, conversational expressions are divided into four expression categories. The conversational expressions of the first expression category correspond to volume, the conversational expressions of the second expression category correspond to timbre, the conversational expressions of the third expression category correspond to tone, and the conversational expressions of the fourth expression category correspond to timbre and volume. For example, conversational expressions A and conversational expressions B belong to the same expression category, and conversational expressions C and conversational expressions D belong to the same expression category. Then there are two expression categories here, and different expression categories correspond to different voice features, while the same expression category corresponds to corresponding voice features, that is, the voice features corresponding to conversational expressions A and conversational expressions B are both timbre, while the voice features corresponding to conversational expressions C and conversational expressions D are both timbre and volume.

[0102] In actual implementation, since different expression categories correspond to different voice features, the process of playing the fused voice message using the target voice features corresponding to the target expression category to which the target conversational expression belongs can be as follows: determining the target voice features corresponding to the target expression category to which the target conversational expression belongs; obtaining the feature value of the target voice feature associated with the target conversational expression, the feature values of the voice features associated with conversational expressions of different expression categories are different; and playing the fused voice message based on the feature value of the target voice feature.

[0103] It should be noted that when the target voice feature corresponds to volume, the characteristic value of the target voice feature indicates the volume when the fused voice message is played; when the target voice feature corresponds to pitch, the characteristic value of the target voice feature indicates the pitch when the fused voice message is played; when the target voice feature corresponds to timbre, the characteristic value of the target voice feature indicates the identifier of the timbre when the fused voice message is played, that is, it is used to determine the timbre when the fused voice message is played; when the target voice feature corresponds to timbre and volume, the characteristic value of the target voice feature indicates the volume and timbre identifier when the fused voice message is played, so that based on the characteristic value of the target voice feature such as volume and timbre identifier, the timbre indicated by the timbre identifier and the corresponding volume are used to play the fused voice message.

[0104] It should be noted that when different conversational expressions correspond to different voice features, in response to the playback instruction for the fused voice message, the target voice features corresponding to the target conversational expressions are used to play the fused voice message. At the same time, the process is similar to the process of playing the fused voice message in response to the playback instruction for the fused voice message as described above, using the target voice features corresponding to the target expression category to which the target conversational expressions belong. The embodiments of the present application do not limit this.

[0105] For example, see Figure 4 ,based on Figure 4 , Figure 4 The target speech features corresponding to the target expression category to which the target conversational expression indicated by 404 in a belongs are timbre and volume. The characteristic value of the target speech feature indicates that the volume when the fused voice message is played is XX decibels. At the same time, the timbre identifier is A. The A identifier indicates that the timbre when the fused voice message is played is the timbre of Zhang Fei in the Romance of the Three Kingdoms. Based on this, the timbre and XX decibels of Zhang Fei in the Romance of the Three Kingdoms are used to play the fused voice message.

[0106] In actual implementation, the fused voice message is carried in a message bubble, which includes a target conversational expression, and at least one conversational expression includes a target conversational expression and other conversational expressions. Therefore, after displaying the fused voice message, it is also possible to, in response to the fusion operation of the fused voice message and the other conversational expressions, compare the expression category to which the other conversational expressions belong and the target expression category; when the comparison result indicates that the expression category to which the other conversational expressions belong is consistent with the target expression category, display a first new fused voice message with a third voice feature; wherein the feature value of the third voice feature is the sum of the feature value corresponding to the other conversational expressions and the feature value of the target voice feature; when the comparison result indicates that the expression category to which the other conversational expressions belong is inconsistent with the target expression category, display a second new fused voice message with a fourth voice feature; wherein the fourth voice feature is different from the second voice feature, and the fourth voice feature corresponds to the other conversational expressions.

[0107] It should be noted that when the conversation interface includes a message display area and a message editing area, the fused voice message here is displayed in the message display area, and similar to the target conversation expression mentioned above, other conversation expressions can also be displayed in the message display area, or can also be displayed in the message editing area, and this embodiment of the present application does not limit this; at the same time, the fusion process of other conversation expressions and fused voice messages is also similar to the fusion process of target conversation expressions and voice messages, and this embodiment of the present application does not elaborate on this. At the same time, when determining that the expression category to which other conversation expressions belong is consistent with the target expression category, it is first necessary to determine the expression category to which other conversation expressions belong. Here, the process of determining the expression category to which other conversation expressions belong is similar to the process of determining the expression category to which the target conversation expression belongs as described above, and this embodiment of the present application does not elaborate on this.

[0108] In actual implementation, when the comparison result indicates that the expression category to which other conversation expressions belong is consistent with the target expression category, the characteristic value of the voice feature corresponding to the expression category to which other conversation expressions belong is obtained, and then the characteristic value is weightedly summed with the characteristic value of the target voice feature to obtain the characteristic value of the third voice feature, thereby determining the first new fused voice message with the third voice feature.

[0109] It should be noted that when the speech parameters corresponding to the target speech feature, such as timbre, volume, pitch, etc., are not exactly the same as the speech parameters corresponding to the speech features corresponding to the expression categories to which other conversational expressions belong, the characteristic values of the same speech parameters are weighted and summed, and different speech parameters are used directly. For example, when the target speech feature corresponds to volume, the corresponding characteristic value, that is, the volume size, is a, and the speech features corresponding to the expression categories to which other conversational expressions belong correspond to timbre and volume, the corresponding characteristic value, that is, the timbre identifier, is a, and the volume size is b, then the volume a and the volume b are weighted and summed to obtain the volume c, so that the third speech feature corresponds to volume and timbre, and the characteristic value of the third speech feature, that is, the volume size, is c and the timbre identifier is a.

[0110] In actual implementation, when the comparison result indicates that the expression category to which other conversational expressions belong is inconsistent with the target expression category, the target voice feature in the fused voice message is directly converted into a fourth voice feature. For example, when the target voice feature corresponds to volume, the corresponding characteristic value, that is, the size of the volume, is a, and the voice feature corresponding to the expression category to which other conversational expressions belong corresponds to timbre and volume, the corresponding characteristic value, that is, the identification of the timbre is a, and the size of the volume is b. After the fused voice message is fused with other conversational expressions, the fourth voice feature corresponds to timbre and volume, the characteristic value of the fourth voice feature, that is, the identification of the timbre is a, and the size of the volume is b.

[0111] It should be noted that, since the voice message with the first voice feature can also be obtained by fusion, and when different expression categories correspond to different voice features, the target voice feature is also the second voice feature.

[0112] In some embodiments, at least one conversational expression includes a target conversational expression and other conversational expressions, and the fused voice message is carried in a message bubble, which includes the target conversational expression; thus, after displaying the fused voice message, it is also possible, in response to a fusion operation on the fused voice message and other conversational expressions, to display a third new fused voice message obtained by replacing the target conversational expression fused in the message bubble with other conversational expressions; wherein the third new fused voice message includes a fifth voice feature, the fifth voice feature is different from the second voice feature, and the fifth voice feature corresponds to the other conversational expressions.

[0113] For example, see Figure 6 , Figure 6 This is a schematic diagram of the process of displaying the third new fusion voice message provided by the embodiment of the present application, based on Figure 6 , the conversational interface includes Figure 6 The message display area indicated by 601 and the message editing area indicated by 602 in a are as shown in FIG. Figure 6The fused voice message indicated by 603 in a is carried in a message bubble, and the message bubble includes the following: Figure 6 The target conversation expression indicated in 606 in a, in response to the target conversation expression indicated in 606, Figure 6 The fused voice message indicated by 603 in a and Figure 6 The fusion operation of other conversation expressions indicated by 604 in a shows that the fusion of the message bubble Figure 6 The target conversation expression indicated by 606 in a is replaced by Figure 6 The other conversation expressions indicated by 607 in b are obtained, such as Figure 6 The third new fused voice message indicated by 605 in b.

[0114] It should be noted that displaying the third new fused voice message obtained by replacing the target conversation expression fused in the message bubble with other conversation expressions means converting the second voice feature included in the fused voice message into the fifth voice feature corresponding to the other conversation expression. The fifth voice feature here also includes at least one of volume, timbre, pitch, breath and accent as described above.

[0115] It should be noted that the message bubble may not include the target conversational expression, so that when a fusion operation for a fused voice message and other conversational expressions is received, a third new fused voice message with the fifth voice feature can be directly displayed, wherein the third new fused voice message is also carried in the message bubble, and the message bubble may include other conversational expressions or may not include other conversational expressions. This embodiment of the present application does not limit this.

[0116] In other embodiments, the voice features of the third new fused voice message can also be determined based on whether the other conversation expressions and the target conversation expressions belong to the same expression category. Specifically, when the other conversation expressions and the target conversation expressions have the same expression category, or represent the same emotions such as sadness, happiness, anger, etc., the fifth voice feature is similar to the third voice feature described above, so that the third new fused voice message is similar to the first new fused voice message described above; when the other conversation expressions and the target conversation expressions have different expression categories, or represent different emotions, the fifth voice feature is similar to the fourth voice feature described above, so that the third new fused voice message is similar to the second new fused voice message described above.

[0117] In some embodiments, after displaying the fused voice message, it is also possible to obtain a message processing method corresponding to the target conversation expression in response to a play instruction for the fused voice message; play the fused voice message, and use the message processing method corresponding to the target conversation expression to process the played fused voice message.

[0118] It should be noted that the message processing method corresponding to the conversational expressions is pre-set, and the message processing methods corresponding to different conversational expressions may be the same or different, and this embodiment of the present application does not limit this.

[0119] In actual implementation, the conversation interface includes a message display area, and the fused voice message is displayed in the message display area; thus, the process of processing the fused voice message that has been played using the message processing method corresponding to the target conversation expression can be, when the message processing method corresponding to the target conversation expression is the read-and-destroy method, when the fused voice message is played, the process of displaying the fused voice message disappearing from the message display area; when the message processing method corresponding to the target conversation expression is the loop playback method, the loop period and number of loops corresponding to the target conversation expression are obtained; after the fused voice message is played, the fused voice message is played again based on the loop period and number of loops.

[0120] It should be noted that when the fused voice message is sent by the target object, playing the fused voice message here may refer to the target object playing the fused voice message itself, or other objects playing the fused voice message. When the message processing mode corresponding to the target conversation expression is the read-and-burn mode, when the fused voice message is played, the process of displaying the fused voice message disappearing from the message display area includes the process of displaying the fused voice message disappearing from the message display area in the conversation interface of the target object when the other object finishes playing the fused voice message, wherein the other object is the object with which the target object conducts a conversation based on the conversation interface. Here, the fused voice message in the conversation interface of the other object will also gradually disappear; or, it may also be the process of displaying the fused voice message disappearing from the message display area in the conversation interface of the target object when the target object finishes playing the fused voice message. This is not limited in this embodiment of the present application.

[0121] For example, see Figure 7 , Figure 7 This is a schematic diagram of the process of processing the played fusion voice message provided by the embodiment of the present application, based on Figure 7 , in such Figure 7 The message display area indicated by the dotted box 701 in a displays the voice message indicated by 703. Figure 7 The message editing area indicated by the dotted box 702 in a displays the target conversation expression indicated by 704. Thus, in response to the fusion operation of the voice message and the target conversation expression, the fused voice message indicated by 705 in b is displayed. When the message processing mode corresponding to the target conversation expression is the self-destructing mode, when the fused voice message is played, the process of the fused voice message disappearing from the message display area is displayed, as shown in FIG. Figure 7 As shown in c.

[0122] It should be noted that when the message processing method corresponding to the conversational expression is a loop playback method, the number of loops and the loop period are both pre-set. Similarly, when the fused voice message is sent by the target object, playing the fused voice message here can refer to the target object playing the fused voice message itself, or other objects playing the fused voice message. When other objects finish playing the fused voice message, the fused voice message is played again in the conversation interface of the other objects based on the loop period and the number of loops; or, when the target object finishes playing the fused voice message, the fused voice message is played again in the conversation interface of the target object based on the loop period and the number of loops.

[0123] In some embodiments, in response to a fusion operation on a voice message and a target conversation expression in at least one conversation expression, after displaying the fused voice message, it is also possible to play the fused voice message in response to a play instruction for the fused voice message, and during the playback of the fused voice message, play animation effects associated with the target conversation expression.

[0124] It should be noted that the animation effects associated with the target conversational expression are pre-set, and the playback method of the animation effects can also be pre-set, for example, it can be played on the entire screen or in the form of a floating window, which is not limited in this embodiment of the application. The animation effects can include flashing display of the target conversational expression, color change display, ripples, spots, and other fluctuations in the expression, which are not limited in this embodiment of the application.

[0125] In some embodiments, in response to a fusion operation on a voice message and a target conversation expression in at least one conversation expression, after displaying the fused voice message, it is also possible to present a text display area including the target conversation expression in response to a text conversion instruction for the fused voice message; in the text display area, the text content corresponding to the fused voice message is presented.

[0126] It should be noted that the text display area can be in one of the upper area, lower area, left area and right area of the fused voice message, and the position of the target conversation expression in the text display area can also be pre-set.

[0127] In some embodiments, the conversation interface includes a message display area, and the fused voice message is displayed in the message display area; in response to the fusion operation of the voice message and the target conversation expression in at least one conversation expression, after displaying the fused voice message, it is also possible to play the fused voice message in response to the play instruction of the fused voice message, and when the playback of the fused voice message is completed, the target conversation expression is converted into a barrage, and the barrage is fixedly displayed at the target position in the message display area; when the display time of the barrage reaches the first target time, the barrage is cancelled.

[0128] It should be noted that the target position is used to indicate the display area occupied by the bullet screen. The display area here can be the entire conversation interface or a portion of the conversation interface, or the entire message display area or a portion of the message display area. This is not limited in the embodiments of the present application. When the display area is a portion of the area, the area ratio of the display area to the corresponding conversation interface or message display area is pre-set, such as the display area occupying half of the conversation interface, or the display area occupying half of the message display area.

[0129] It should be noted that the first target duration can be pre-set, such as 5 seconds. When the display duration of the barrage reaches the first target duration, the barrage is canceled. In this way, while ensuring the effect of the barrage, excessive interference with the user's conversation process is prevented, thereby improving the user's conversation experience.

[0130] In actual implementation, the fused voice message is sent by the target object; after the barrage is fixedly displayed at the target position of the message display area, the reply barrage of other objects to the barrage can also be displayed floating above the barrage; among them, the other objects are the objects with which the target object conducts conversations based on the conversation interface.

[0131] It should be noted that the other objects can be one or more. When the other objects see the barrage sent by the target object, they send a reply message in response to the barrage. When the reply message is sent to the message display area in the conversation interface, the reply message is converted into a reply barrage and displayed floating above the barrage sent by the target object.

[0132] In actual implementation, after the barrage is fixedly displayed at the target location in the message display area, if the target party continues to send conversation messages such as text messages and voice messages, the continued conversation messages can also be displayed in the barrage. Specifically, the conversation interface also includes a message editing area, so that after the barrage is fixedly displayed at the target location in the message display area, the edited message content can also be displayed in response to a message editing operation triggered based on the message editing area; in response to a send instruction for the edited message content, the edited message content can be displayed in the barrage at the target location. The message content includes at least one of a text message, an emoticon message, and a voice message.

[0133] In some embodiments, the fused voice message obtained by fusion can also be restored. Specifically, in response to the fusion operation on the voice message and the target conversation expression in at least one conversation expression, after the fused voice message is displayed, it is also possible to display the process of restoring the fused voice message to obtain the voice message and the target conversation expression in response to the restoration condition of the fused voice message being met; wherein the restoration condition includes at least one of the following: the fused voice message is played; the display time of the fused voice message reaches the second target time; a restoration instruction for the fused voice message is received, and the restoration instruction is used to instruct the fused voice message to be restored to the voice message and the target conversation expression.

[0134] It should be noted that when a fused voice message is sent to a target object, the reply message to the fused voice message is to an object other than the target object in the conversation interface; and the second target duration can also be pre-set, for example, 5 seconds; at the same time, the restoration instruction for the fused voice message can be triggered by the target object corresponding to the conversation interface, and a restoration control for restoring the fused voice message can be displayed on the conversation interface, thereby receiving the restoration instruction for the fused voice message in response to the triggering operation of the restoration control, or receiving the restoration instruction for the fused voice message in response to receiving a restoration operation for the fused voice message triggered on the conversation interface, such as a pressing operation or a sliding operation.

[0135] In some embodiments, the first new fused voice message, the second new fused voice message and the third new fused voice message mentioned above can also be restored, wherein the restoration conditions for the first new fused voice message, the second new fused voice message and the third new fused voice message are similar to the restoration conditions for the fused voice message, wherein the restoration process for the first new fused voice message, the second new fused voice message and the third new fused voice message is similar to the restoration process for the fused voice message, and this embodiment of the present application will not go into details; at the same time, after the first new fused voice message, the second new fused voice message and the third new fused voice message are restored to obtain the fused voice message, the fused voice message can also be restored, and the restoration process is as described above, and this embodiment of the present application will not go into details.

[0136] Applying the above-mentioned embodiment of the present application, a first message style is first adopted in the conversation interface to display a voice message having a first voice feature and at least one conversational expression. Then, a fused voice message is obtained by fusing the voice message and the target conversational expression in the at least one conversational expression. The fused voice message satisfies at least one of the following conditions: the message style of the fused voice message changes from the first message style to the second message style, and the voice feature of the fused voice message changes from the first voice feature to the second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversational expression. In this way, by fusing the voice message and the expression, at least one of the voice feature and the message style of the voice message is changed, thereby increasing the diversity of conversational messages during the conversation.

[0137] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0138] During the conversation process of related technologies, the chat bubble shape is fixed, and the conversation experience is dull and boring. At the same time, user emotions are expressed entirely through text or emoticons, which are relatively simple. In addition, misunderstandings of emotions may occur in many scenarios, further reducing the conversation experience.

[0139] Based on this, the present invention provides a new way of text and voice expression for chat dialogues, aiming to express user emotions more accurately and unlock more interactive gameplay. Specifically, by integrating text messages with text messages (text conversation messages), integrating emoticons (expressions) with bubbles (target messages), and integrating emoticons with voice messages in chat scenarios, the bubble state (display style) of voice messages and voice messages is transformed to be more in line with physical intuition. In this way, the emotional transmission between users can be expressed more clearly, and user interaction can be increased in a more interesting way. At the same time, users can enhance the expression of text messages by pinching text messages with both hands and changing the shape of text bubbles. Moreover, by long pressing a text message and dragging it to another text message, text messages can also be merged to change the shape of text bubbles to enhance the expression of text messages. In addition, dragging an emoticon package into a text message or voice message can give the message different forms, such as bubble shape, color, text, sound, emoticon package, etc. In addition, when the user is not speaking, he can directly drag the emoticon package onto the voice control, so that the current buzzwords are intelligently generated according to the meaning of the emoticon package and converted into AI voice to be sent.

[0140] Next, the technical solution of this application is explained from the product side.

[0141] For text messages, in response to a pinch operation on the text message, the chat bubble color, shape, expression, etc. are changed, thereby enhancing the user's emotional expression; or, in response to dragging an expression into the bubble, the chat bubble color, shape, expression, etc. are changed, thereby enhancing the user's emotional expression.

[0142] For voice messages, such as Figure 4 、 5 6, in response to the drag operation on the emoticon, the emoticon is dragged into the voice message, thereby changing the timbre of the voice; or, in response to dragging the emoticon into an empty voice (the first voice control), the relevant voice is automatically generated according to the emoticon, and then when a playback instruction for the voice is received, it is read out intelligently by the AI.

[0143] Next, the technical solution of this application is explained from a technical perspective.

[0144] In actual implementation, the technical solution of this application mainly includes three processes, namely, the operation process of users changing the chat bubble style through expressions and gestures, the operation process of super bubble scene barrage floating screen and recovery, and the process of changing bubble emotions through user operations.

[0145] For the user's operation process of changing the chat bubble style through emoticons and gestures, see Figure 8 , Figure 8This is a flow chart of the operation process of the user changing the chat bubble style through expressions and gestures provided by the embodiment of the present application, based on Figure 8 , the user changes the chat bubble style through expressions and gestures Figure 8 Specifically, you can first enter a normal text message, and then drag the emoticon to the chat bubble of the text message to change the chat bubble style, or pinch adjacent text messages to change the chat bubble style.

[0146] For inputting ordinary text messages, first, in response to the message sender's sending operation for the input chat conversation text, the chat conversation text input by the message sender is sent, wherein the sending client, that is, the client corresponding to the message sender, sends the chat conversation text to the server, so that after the server receives the chat conversation text, it generates a unique message identifier, writes the chat conversation text to the storage, and pushes the chat conversation text to the receiving client, that is, the client corresponding to the message recipient. After the chat conversation text is successfully displayed on the receiving client, the sending client receives the unique message identifier of the chat conversation text sent by the server, and displays the chat bubble and text content corresponding to the chat conversation text on the session interface.

[0147] Regarding the process of changing the chat bubble style by dragging an emoticon into the chat bubble of the above-mentioned chat conversation text, in response to the message sender dragging the system emoticon into the chat bubble of the chat conversation text, the sending client sends a chat mood change request carrying the emoticon dragged by the message sender and the identifier of the chat conversation text to the server. The server receives the chat mood change request, determines the emotion expressed by the emoticon, and calculates a new bubble style result in combination with the current emotion of the message, updates it to the cloud storage, and then pushes a style change result of the message to the receiving client, that is, the client corresponding to the message recipient. After successfully displaying the style change result on the receiving client, the sending client receives the new bubble style of the chat conversation text sent by the server, thereby displaying the new chat bubble and text content corresponding to the chat conversation text on the session interface.

[0148] Regarding the process of changing the chat bubble style by pinching adjacent text messages by gesture, in response to the message sender's pinching operation on the chat conversation text and the text messages adjacent to the chat conversation text, the sending client extracts the message identifier of the pinched message, and sends a chat mood change request carrying a list of identifiers of messages associated with the pinching operation (if emoticons are transmitted together), to the server. The server receives the chat mood change request, analyzes the list of pinched identifiers or system emoticons in the request, calculates the new bubble style result according to the number of messages and the meaning of the emoticons, updates it to the cloud storage, and then pushes a style change result of the message to the receiving client, that is, the client corresponding to the message recipient. After the style change result is successfully displayed on the receiving client, the sending client receives the new bubble style of the chat conversation text sent by the server, thereby displaying the new chat bubble and text content corresponding to the chat conversation text on the session interface.

[0149] For the operation process of floating and restoring the super bubble scene bullet screen, see Figure 9 , Figure 9 This is a flow chart of the operation process of floating and restoring the super bubble scene barrage provided by the embodiment of the present application, based on Figure 9 The super bubble scene barrage floating screen and recovery operation process is through Figure 9 Specifically, when the changed chat bubble is a super bubble such as a screen-dominating bullet screen, when the message sender continues to send messages, the first to fourth messages will be displayed in the super bubble on the screen, and will also be displayed in the super bubble of the recipient's client interface on the screen; when the fifth message is sent, the super bubble is restored, that is, the messages included in the super bubble are restored to separately displayed messages and displayed in the style of ordinary bubbles.

[0150] For the process of changing the bubble's mood through user operations, see Figure 10 , Figure 10 This is a flow chart of the process of changing the bubble emotion through user operation provided by the embodiment of the present application, based on Figure 10 The process of changing the bubble mood through user operation in this application is Figure 10Specifically, in the conversation interface between the message sender and the message recipient, in response to the message sender's sending operation on the chat text, the sent chat text is displayed in the conversation interface, and then in response to the message sender dragging the angry expression to the chat bubble, an angry emotion bubble is generated, that is, based on the angry emotion bubble, the sent chat text is displayed, and then when the message recipient replies to the chat text, in response to the message sender dragging the happy expression to the angry emotion bubble, the angry emotion bubble is changed to a happy emotion bubble, and then based on the happy emotion bubble, the sent chat text is displayed.

[0151] In some embodiments, when emotions cannot be directly identified, chat bubble forms with different emotions can be generated by simply determining keywords, key punctuation marks, user-defined emotions, etc.; at the same time, when emotions are recognized through voice and voice is converted into text, user emotions can be judged to generate different emotional chat bubble forms.

[0152] In this way, through this application, the user's true emotions can be understood to better convey emotional information to other users, allowing users to feel the other party's emotions and semantics more directly, while increasing the fun of the chat.

[0153] Applying the above-mentioned embodiment of the present application, a first message style is first adopted in the conversation interface to display a voice message having a first voice feature and at least one conversational expression. Then, a fused voice message is obtained by fusing the voice message and the target conversational expression in the at least one conversational expression. The fused voice message satisfies at least one of the following conditions: the message style of the fused voice message changes from the first message style to the second message style, and the voice feature of the fused voice message changes from the first voice feature to the second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversational expression. In this way, by fusing the voice message and the expression, at least one of the voice feature and the message style of the voice message is changed, thereby increasing the diversity of conversational messages during the conversation.

[0154] The following continues to describe the exemplary structure of the session message processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the session message processing device 455 of the memory 450 may include:

[0155] Display module 4551, configured to display a voice message having a first voice feature in a conversation interface using a first message style and at least one conversation emoticon;

[0156] The fusion module 4552 is used to display a fused voice message in response to a fusion operation on the voice message and the target conversation expression in the at least one conversation expression; wherein the fused voice message satisfies at least one of the following: the message style of the fused voice message is changed from the first message style to the second message style, and the voice feature of the fused voice message is changed from the first voice feature to the second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversation expression.

[0157] In some embodiments, the conversation interface includes a message display area and a message editing area, the voice message is displayed in the message display area, and the target conversation emoticon is displayed in the message editing area; the fusion module 4552 is also used to respond to a drag operation on the target conversation emoticon, drag the target conversation emoticon from the message editing area to the voice message in the message display area; in response to a release operation on the target conversation emoticon, display a fused voice message obtained by fusing the target conversation emoticon and the voice message.

[0158] In some embodiments, the device also includes a switching module, which is used to display a switching control in the associated area of the fused voice message and automatically play the fused voice message; wherein the switching control is used to switch the message style and at least one of the voice features of the fused voice message; in response to a triggering operation on the switching control, the fused voice message is switched to an associated fused voice message, and the associated fused voice message is associated with the target conversation expression.

[0159] In some embodiments, there are multiple conversational expressions, and the multiple conversational expressions belong to at least two expression categories, and different expression categories correspond to different voice features; the device also includes a first playback module, which is used to respond to a playback instruction for the fused voice message and use the target voice feature corresponding to the target expression category to which the target conversational expression belongs to play the fused voice message.

[0160] In some embodiments, the first playback module is further used to determine the target voice feature corresponding to the target expression category to which the target conversational expression belongs; obtain the feature value of the target voice feature associated with the target conversational expression, the feature values of the voice feature associated with conversational expressions of different expression categories are different; and play the fused voice message based on the feature value of the target voice feature.

[0161] In some embodiments, the fused voice message is carried in a message bubble, which includes the target conversation expression. The device also includes a comparison module, which is used to compare the expression category to which the other conversation expressions belong and the target expression category in response to the fusion operation on the fused voice message and the other conversation expressions; when the comparison result indicates that the expression category to which the other conversation expressions belong is consistent with the target expression category, a first new fused voice message with a third voice feature is displayed; wherein the feature value of the third voice feature is the sum of the feature value corresponding to the other conversation expression and the feature value of the target voice feature; when the comparison result indicates that the expression category to which the other conversation expressions belong is inconsistent with the target expression category, a second new fused voice message with a fourth voice feature is displayed; wherein the fourth voice feature is different from the second voice feature, and the fourth voice feature corresponds to the other conversation expression.

[0162] In some embodiments, the at least one conversational expression includes the target conversational expression and other conversational expressions, the fused voice message is carried in a message bubble, and the message bubble includes the target conversational expression; the device also includes a second fusion module, which is used to respond to a fusion operation on the fused voice message and the other conversational expressions and display a third new fused voice message obtained by replacing the target conversational expression fused in the message bubble with the other conversational expressions; wherein the third new fused voice message includes a fifth voice feature, the fifth voice feature is different from the second voice feature, and the fifth voice feature corresponds to the other conversational expressions.

[0163] In some embodiments, the device also includes a second playback module, which is used to respond to a playback instruction for the fused voice message, obtain a message processing method corresponding to the target conversation expression; play the fused voice message, and use the message processing method corresponding to the target conversation expression to process the fused voice message that has been played.

[0164] In some embodiments, the conversation interface includes a message display area, and the fused voice message is displayed in the message display area; the second playback module is also used to, when the message processing method corresponding to the target conversation expression is the read-and-destroy method, display the process of the fused voice message disappearing from the message display area when the fused voice message is played; when the message processing method corresponding to the target conversation expression is the loop playback method, obtain the loop period and number of loops corresponding to the target conversation expression; after the fused voice message is played, play the fused voice message again based on the loop period and number of loops.

[0165] In some embodiments, the device also includes a third playback module, which is used to play the fused voice message in response to a playback instruction for the fused voice message, and in the process of playing the fused voice message, play the animation effects associated with the target conversation expression.

[0166] In some embodiments, the device also includes a text conversion module, which is used to present a text display area including the target conversation emoticon in response to a text conversion instruction for the fused voice message; in the text display area, the text content corresponding to the fused voice message is presented.

[0167] In some embodiments, the conversation interface includes a message display area, and the fused voice message is displayed in the message display area; the device also includes a fourth playback module, which is used to respond to a playback instruction for the fused voice message, play the fused voice message, and when the fused voice message is played, convert the target conversation expression into a barrage, and fix the barrage at the target position of the message display area; when the display time of the barrage reaches the first target time, cancel the display of the barrage.

[0168] In some embodiments, the fused voice message is sent by the target object; the device also includes a reply module, which is used to float and display reply barrages of other objects to the barrages above the barrage; wherein, the other objects are the objects with which the target object conducts conversations based on the conversation interface.

[0169] In some embodiments, the conversation interface also includes a message editing area, and the device also includes an editing module, wherein the editing module is used to display the edited message content in response to a message editing operation triggered based on the message editing area; in response to a sending instruction for the edited message content, the edited message content is displayed in the barrage at the target location.

[0170] In some embodiments, the device also includes a restoration module, which is used to display the process of restoring the fused voice message to obtain the voice message and the target conversation expression in response to the restoration condition of the fused voice message being met; wherein the restoration condition includes at least one of the following: the fused voice message is played; the display time of the fused voice message reaches a second target time; a restoration instruction for the fused voice message is received, and the restoration instruction is used to instruct the fused voice message to be restored to the voice message and the target conversation expression.

[0171] In some embodiments, the conversation interface includes a message display area, and the voice message and the target conversation expression are displayed in the message display area; the fusion module 4552 is also used to display the process of merging the voice message and the target conversation expression in response to a pinch operation on the voice message and the target conversation expression; when the ratio of the first area to the second area reaches a target ratio, the voice message and the target conversation expression are canceled; wherein the first area is the area of the overlapping area of the voice message and the target conversation expression, and the second area is the display area of the voice message or the display area of the target conversation expression; the fused voice message obtained by fusing the voice message and the target conversation expression is displayed.

[0172] An embodiment of the present application provides a computer program product comprising computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the method for processing conversation messages described in the embodiment of the present application.

[0173] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the method for processing a session message provided in the embodiment of the present application, for example, Figure 3 The method for processing session messages is shown.

[0174] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disk, or a CD-ROM; or various devices including one or any combination of the above memories.

[0175] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0176] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0177] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0178] In summary, the embodiments of the present application have the following beneficial effects:

[0179] By fusing the voice message and the emoticon, at least one of the voice feature and the message style of the voice message is changed, thereby improving the diversity of the conversation message during the conversation.

[0180] It should be noted that in the embodiments of the present application, when obtaining conversation messages such as voice messages and emoticon messages, user operation data and other related data, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0181] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A method for processing a conversation message, characterized in that: The method comprises: In the conversation interface, a first message style is used to display a voice message having a first voice feature and at least one conversation emoticon; In response to a fusion operation on the voice message and a target conversation emoticon in the at least one conversation emoticon, displaying a fused voice message; The fused voice message satisfies at least one of the following conditions: the message style of the fused voice message changes from the first message style to the second message style, and the voice feature of the fused voice message changes from the first voice feature to the second voice feature; The second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversation expression.

2. The method according to claim 1, wherein The conversation interface includes a message display area and a message editing area, the voice message is displayed in the message display area, and the target conversation emoticon is displayed in the message editing area; The displaying of the fused voice message in response to the fusion operation on the voice message and the target conversation expression in the at least one conversation expression includes: In response to a drag operation on the target conversation emoticon, dragging the target conversation emoticon from the message editing area to the voice message in the message display area; In response to a release operation on the target conversation emoticon, a fused voice message obtained by fusing the target conversation emoticon and the voice message is displayed.

3. The method according to claim 1, wherein After displaying the fused voice message in response to the fusion operation on the voice message and the target conversation emoticon in the at least one conversation emoticon, the method further includes: Displaying a switching control in an area associated with the fused voice message and automatically playing the fused voice message; The switching control is used to switch at least one of the message style and the voice feature of the fused voice message; In response to a triggering operation on the switching control, the fused voice message is switched to an associated fused voice message, where the associated fused voice message is associated with the target conversation emoticon.

4. The method according to claim 1, wherein There are multiple conversation expressions, each of which belongs to at least two expression categories, and different expression categories correspond to different voice features; After displaying the fused voice message, the method further includes: In response to a play instruction for the fused voice message, the fused voice message is played using a target voice feature corresponding to a target expression category to which the target conversational expression belongs.

5. The method according to claim 4, wherein The step of playing the fused voice message by using a target voice feature corresponding to a target expression category to which the target conversation expression belongs includes: Determine the target speech feature corresponding to the target expression category to which the target conversational expression belongs; Obtaining a feature value of the target voice feature associated with the target conversational expression, wherein the feature values of the voice feature associated with conversational expressions of different expression categories are different; The fused voice message is played based on the feature value of the target voice feature.

6. The method according to claim 5, wherein The fused voice message is carried in a message bubble, the message bubble includes the target conversation emoticon, the at least one conversation emoticon includes the target conversation emoticon and other conversation emoticons, and after displaying the fused voice message, the method further includes: In response to a fusion operation on the fused voice message and the other conversation expressions, comparing the expression category to which the other conversation expressions belong with the target expression category; When the comparison result indicates that the expression category to which the other conversation expressions belong is consistent with the target expression category, a first new fused voice message having a third voice feature is displayed; wherein the feature value of the third voice feature is the sum of the feature value corresponding to the other conversation expressions and the feature value of the target voice feature; When the comparison result indicates that the expression category to which the other conversation expressions belong is inconsistent with the target expression category, a second new fused voice message with a fourth voice feature is displayed; wherein the fourth voice feature is different from the second voice feature, and the fourth voice feature corresponds to the other conversation expressions.

7. The method according to claim 1, wherein The at least one conversation expression includes the target conversation expression and other conversation expressions, the fused voice message is carried in a message bubble, and the message bubble includes the target conversation expression; After displaying the fused voice message, the method further includes: In response to a fusion operation on the fused voice message and the other conversation emoticons, displaying a third new fused voice message obtained by replacing the target conversation emoticon fused in the message bubble with the other conversation emoticons; The third new fused voice message includes a fifth voice feature, the fifth voice feature is different from the second voice feature, and the fifth voice feature corresponds to the other conversation expressions.

8. The method according to claim 1, wherein After displaying the fused voice message, the method further includes: In response to a play instruction for the fused voice message, obtaining a message processing method corresponding to the target conversation emoticon; The fused voice message is played, and the fused voice message that has been played is processed using a message processing method corresponding to the target conversational expression.

9. The method according to claim 8, wherein The conversation interface includes a message display area, and the fused voice message is displayed in the message display area; and the message processing method corresponding to the target conversation expression is used to process the played fused voice message, including: When the message processing mode corresponding to the target conversation emoticon is the self-destructing mode, when the playback of the fused voice message is completed, the process of the fused voice message disappearing from the message display area is displayed; When the message processing mode corresponding to the target conversation expression is a loop playback mode, obtaining the loop period and loop count corresponding to the target conversation expression; After the fused voice message is played, the fused voice message is played again based on the loop period and the number of loops.

10. The method according to claim 1, wherein After displaying the fused voice message in response to the fusion operation on the voice message and the target conversation emoticon in the at least one conversation emoticon, the method further includes: In response to a play instruction for the fused voice message, play the fused voice message, and During the playback of the fused voice message, an animation effect associated with the target conversation expression is played.

11. The method according to claim 1, wherein After displaying the fused voice message in response to the fusion operation on the voice message and the target conversation emoticon in the at least one conversation emoticon, the method further includes: In response to a text conversion instruction for the fused voice message, presenting a text display area including the target conversation emoticon; In the text display area, the text content corresponding to the fused voice message is presented.

12. The method according to claim 1, wherein The conversation interface includes a message display area, and the fused voice message is displayed in the message display area; After displaying the fused voice message in response to the fusion operation on the voice message and the target conversation emoticon in the at least one conversation emoticon, the method further includes: In response to a play instruction for the fused voice message, play the fused voice message, and when the fused voice message is played, convert the target conversation emoticon into a barrage, and display the barrage fixedly at a target position in the message display area; When the display duration of the bullet comment reaches a first target duration, the display of the bullet comment is canceled.

13. The method according to claim 12, wherein: The fused voice message is sent by a target object; after the barrage is fixedly displayed at a target position in the message display area, the method further includes: Above the bullet screen, a floating display of reply bullet screens from other objects in response to the bullet screen; The other objects are objects with which the target object conducts a conversation based on the conversation interface.

14. The method according to claim 12, wherein: The conversation interface further includes a message editing area. After the barrage is fixedly displayed at a target position in the message display area, the method further includes: In response to a message editing operation triggered based on the message editing area, displaying the edited message content; In response to a sending instruction for the edited message content, the edited message content is displayed in the bullet screen at the target location.

15. The method according to claim 1, wherein After displaying the fused voice message in response to the fusion operation on the voice message and the target conversation emoticon in the at least one conversation emoticon, the method further includes: In response to a restoration condition of the fused voice message being met, displaying a process of restoring the fused voice message to obtain the voice message and the target conversation emoticon; Wherein, the reducing conditions include at least one of the following: The fusion voice message is played; The display duration of the fused voice message reaches a second target duration; A restoration instruction for the fused voice message is received, where the restoration instruction is used to instruct to restore the fused voice message to the voice message and the target conversation emoticon.

16. The method according to claim 1, wherein The conversation interface includes a message display area, and the voice message and the target conversation emoticon are displayed in the message display area; The displaying of the fused voice message in response to the fusion operation on the voice message and the target conversation expression in the at least one conversation expression includes: In response to a pinch operation on the voice message and the target conversation emoticon, displaying a process of merging the voice message and the target conversation emoticon; When the ratio of the first area to the second area reaches a target ratio, canceling the display of the voice message and the target conversation emoticon; The first area is the area of the overlapping region of the voice message and the target conversation emoticon, and the second area is the display area of the voice message or the display area of the target conversation emoticon; The fused voice message obtained by fusing the voice message and the target conversation expression is displayed.

17. A device for processing conversation messages, characterized in that: The device comprises: A display module, configured to display, in a conversation interface, a voice message having a first voice feature using a first message style and displaying at least one conversation emoticon; A fusion module is used to display a fused voice message in response to a fusion operation on the voice message and a target conversation expression in the at least one conversation expression; wherein the fused voice message satisfies at least one of the following: the message style of the fused voice message is changed from the first message style to the second message style, and the voice feature of the fused voice message is changed from the first voice feature to the second voice feature; the second message style is different from the first message style, the second voice feature is different from the first voice feature, and both the second message style and the second voice feature correspond to the target conversation expression.

18. An electronic device, characterized in that: include: a memory for storing computer-executable instructions; The processor is configured to implement the method for processing a conversation message according to any one of claims 1 to 16 when executing the computer-executable instructions stored in the memory.

19. A computer-readable storage medium, characterized in that Computer executable instructions are stored, which are used to cause a processor to execute and implement the method for processing a conversation message according to any one of claims 1 to 16.

20. A computer program product comprising computer executable instructions, characterized in that When the computer executable instructions are executed by a processor, the method for processing a conversation message according to any one of claims 1 to 16 is implemented.