Emotional interaction method and apparatus, device, and storage medium
By creating voice expressions containing visual and pronunciation parts, the problem of insufficient expression resources of traditional expressions is solved, and the efficiency and expression ability of information interaction are improved.
Patent Information
- Application Number
- PCT/CN2025/070461
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-10
AI Technical Summary
Traditional expression resources are difficult to fully express the information people want to convey, resulting in low information interaction efficiency.
By determining the target media resource, voice expressions are created based on the received editing operation, including visual and voice parts, and allow the user to choose to send to the associated session.
It improves the efficiency of voice expression production, enables users to express information conveniently, and enhances the effect of information interaction.
Smart Images

Figure CN2025070461_10072025_PF_FP_ABST
Abstract
Description
Method, device, equipment and storage medium for expression interaction
[0001] This application claims priority to Chinese invention patent application number 202410023197.4, filed on January 5, 2024, and entitled “Method, apparatus, device and storage medium for expression interaction”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for expression interaction. Background Art
[0003] With the development of computer technology, the Internet has become an important platform for people to exchange information. In the process of people exchanging information through the Internet, such as in chat sessions, various types of emojis have become an important medium for people to express themselves socially and exchange information. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for expression interaction is provided. The method includes: determining a target media resource for generating a voice expression; creating a voice expression based on a received editing operation, wherein the voice expression includes a visual portion and a voice portion, wherein the editing operation includes at least one of the following: a first operation for modifying first visual content and / or first voice content included in the target media resource, and a second operation for inputting second visual content and / or second voice content; and publishing the voice expression, such that the voice expression can be sent to a conversation associated with the target user based on a selection of the target user.
[0005] In a second aspect of the present disclosure, a method for emoticon interaction is provided. The method includes: receiving a selection of a voice emoticon; and sending the voice emoticon to a target conversation, wherein the voice emoticon includes a visual portion and a voice portion, wherein the voice emoticon is created based on an edit operation associated with a target media resource, wherein the edit operation includes at least one of the following: a first operation for modifying first visual content and / or first voice content included in the target media resource; and a second operation for adding second visual content and / or second voice content.
[0006] In a third aspect of the present disclosure, a method for creating an emoticon is provided. The method includes: presenting, based on target speech content, a first set of candidate visual content determined based on the target speech content; and creating a speech emoticon based on a selection of first visual content from the first set of candidate visual content, the speech emoticon comprising a visual portion and a speech portion, wherein the visual portion is determined based on the first visual content and the speech portion is determined based on the target speech content.
[0007] In a fourth aspect of the present disclosure, a method for creating an emoticon is provided. The method includes: presenting, based on target text content, a second set of candidate visual content determined based on the target text content; and creating a voice emoticon based on a selection of second visual content from the second set of candidate visual content, the voice emoticon comprising a visual portion and a voice portion, wherein the visual portion is determined based on the target visual content and the voice portion is generated based on the target text content.
[0008] In a fifth aspect of the present disclosure, a device for expression interaction is provided. The device includes: a determination module configured to determine a target media resource for generating a voice expression; a creation module configured to create a voice expression based on a received editing operation, wherein the voice expression includes a visual part and a voice part, and the editing operation includes at least one of the following: a first operation for modifying first visual content and / or first voice content included in the target media resource, and a second operation for inputting second visual content and / or second voice content; and a publishing module configured to publish the voice expression, so that the voice expression can be sent to a session associated with the target user based on a selection of the target user.
[0009] In a sixth aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of any one of aspects 1 to 4.
[0010] In a seventh aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of any one of the first to fourth aspects.
[0011] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0013] FIG1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure may be implemented;
[0014] FIG2 shows a flowchart of an example process of expression interaction according to some embodiments of the present disclosure;
[0015] 3A to 3E illustrate example interfaces according to some embodiments of the present disclosure;
[0016] FIG4 shows a flowchart of an example process of expression interaction according to some embodiments of the present disclosure;
[0017] 5A-5B illustrate example interfaces according to some embodiments of the present disclosure;
[0018] FIG6 shows a flowchart of an example process of emoticon creation according to some embodiments of the present disclosure;
[0019] FIG7 shows a flowchart of an example process of emoticon creation according to some embodiments of the present disclosure;
[0020] FIG8 shows a schematic structural block diagram of an example apparatus for expression interaction according to some embodiments of the present disclosure; and
[0021] FIG9 illustrates a block diagram of an electronic device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0022] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0023] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0024] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0025] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0026] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.
[0027] When people interact with each other online, they expect to use high-quality emoticons to express their desired information conveniently. However, traditional emoticon resources are unable to fully express the information people want to convey.
[0028] Embodiments of the present disclosure provide an expression interaction solution. According to the solution, a target media resource for generating a voice expression can be determined. Furthermore, a voice expression can be created based on received editing operations, where the voice expression includes a visual component and a voice component.
[0029] Such an editing operation may include at least one of the following: a first operation for modifying the first visual content and / or first voice content included in the target media resource, and a second operation for inputting second visual content and / or second voice content.
[0030] Furthermore, the voice expression may be published so that the voice expression can be sent to a conversation associated with the target user based on a selection of the target user.
[0031] In this way, the embodiments of the present disclosure can support users in editing media resources to create voice expressions, thereby improving the efficiency of producing voice expressions.
[0032] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.
[0033] Sample Environment
[0034] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the example environment 100 may include an electronic device 110 .
[0035] In this example environment 100, electronic device 110 may run an application 120 that supports interface interaction. Application 120 may be any suitable type of application for interface interaction, examples of which may include, but are not limited to, video applications, social applications, or other suitable applications. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.
[0036] In the environment 100 of FIG. 1 , if the application 120 is in an active state, the electronic device 110 may present an interface 150 for supporting interface interaction through the application 120 .
[0037] In some embodiments, the electronic device 110 communicates with the server 130 to enable the provision of services for the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0038] The server 130 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. For example, the server 130 may include a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like. The server 130 may provide background services for the application 120 that supports virtual scenes in the electronic device 110.
[0039] A communication connection may be established between the server 130 and the electronic device 110. The communication connection may be established in a wired or wireless manner. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the electronic device 110 may implement signaling interaction through the communication connection between the two.
[0040] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0041] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0042] First example process
[0043] FIG2 shows a flow chart of an example interaction process 200 according to some embodiments of the present disclosure. The process 200 may be implemented at the electronic device 110. The process 200 is described below with reference to FIG1.
[0044] As shown in FIG. 2 , at block 210 , the electronic device 110 determines a target media resource for generating a voice expression.
[0045] In some embodiments, such target media resources may include any appropriate type of resources including visual content and / or audio content, such as pictures, videos, audio, and emoticons.
[0046] In some embodiments, the target media resource may include a media resource input by the user. For example, the user may shoot a video, a picture, or record a voice as the target media resource for generating a voice expression.
[0047] In some embodiments, the electronic device 110 may also determine a target media resource for generating a voice expression based on a user's selection of a target media resource from the candidate media resources.
[0048] The process 200 will be described below with reference to Figures 3A to 3E . Figures 3A to 3E show example interfaces 300A to 300E according to some embodiments of the present disclosure. Interfaces 300A to 300E may be provided by, for example, the electronic device 110 shown in Figure 1 .
[0049] As shown in Figure 3A, interface 300A may be, for example, a conversation interface. Electronic device 110 may display an expression panel 305 in interface 300A. For example, expression panel 305 may display added voice expressions, such as voice expression 315.
[0050] Additionally, the expression panel 305 may further include, for example, a creation portal 310. Upon receiving a selection of the creation portal 310, the electronic device 110 may present an interface 300B as shown in FIG3B.
[0051] In interface 300B, electronic device 110 may display a set of candidate images, such as candidate image 320-1, candidate image 320-2, candidate image 320-3, and candidate image 320-4 (individually or collectively referred to as candidate image 320). In some embodiments, such candidate images may include static images and / or static images.
[0052] In some embodiments, such candidate images 320 may be associated with the current user. For example, candidate images 320 may include images collected by the current user. As another example, candidate images 320 may also include images associated with works published by the current user, such as image works published by the current user or video frames of video works published by the current user.
[0053] In some embodiments, with the user's knowledge and authorization, the candidate images 320 may also include images local to the electronic device 110 .
[0054] Furthermore, the electronic device 110 may, for example, receive a user's selection of the candidate image 320 - 3 , and may accordingly determine the candidate image 320 - 3 as a target media resource for generating a voice expression.
[0055] The electronic device 110 may also determine the target media resource for generating the voice expression based on other appropriate methods. In some embodiments, the electronic device 110 may also support the user to select an available expression as the target media resource.
[0056] In some examples, the emoticons available to the user may include image emoticons, voice emoticons, etc. that the user has added; or image emoticons, voice emoticons, etc. shared by other users.
[0057] In some embodiments, the electronic device 110 may present a set of candidate image expressions in the expression panel 305. As an example, such candidate image expressions may have only a visual part and not include a speech part.
[0058] Furthermore, the electronic device 110 may receive a user's selection of a target image expression from the group of candidate image expressions, and determine the target image expression as a target media resource.
[0059] In some further embodiments, the electronic device 110 may also receive a preset operation associated with the first work, and may accordingly determine the first work as the target media resource.
[0060] For example, the electronic device 110 may provide a voice expression creation portal in association with the viewing interface of the first work, and may determine the first work as the target media resource based on the user's selection of the creation portal.
[0061] Continuing with reference to Figure 2, in box 220, the electronic device 110 creates a voice expression based on the received editing operation, where the voice expression includes a visual part and a voice part, and the editing operation includes at least one of the following: a first operation for modifying the first visual content and / or first voice content included in the target media resource, and a second operation for inputting second visual content and / or second voice content.
[0062] In some embodiments, the electronic device 110 may provide the user with an editing interface for creating voice expressions.
[0063] In some embodiments, the electronic device 110 may modify the first visual content and / or first voice content included in the target media resource accordingly based on the user's first operation in the editing interface.
[0064] 3B is used as an example, when the target media resource is a candidate image 320-3, the electronic device 110 may provide an image editing interface for editing the candidate image 320-3. It should be understood that such an image editing interface may support any appropriate type of image editing operation, such as cropping an image, adding text, adding stickers, applying filters, and the like.
[0065] For example, as shown in FIG3C , the electronic device 110 may display the edited image 325 in the interface 300C. It should be understood that such an image editing process is not necessary. The image 325 may also be the original candidate image 320 - 3 .
[0066] 3C , the electronic device 110 may provide a control 330 for adding voice content. For example, when the user triggers the control 330, the electronic device 110 may obtain the voice content input by the user.
[0067] It should be understood that such voice content may include any appropriate audio content, regardless of whether such audio content represents corresponding text information. For example, the voice content may include a recorded speech of a user, specific sounds in the environment, etc. For another example, the voice content may also include music content with or without lyrics.
[0068] 3D , after the voice content input is completed, the electronic device 110 may present an interface 300D, for example. In the interface 300D, the electronic device 110 may provide a preview entry 335 for playing the input voice content.
[0069] Additionally, the electronic device 110 may also apply a specific voice style to the input voice content. As shown in FIG3D , the electronic device 110 may, for example, provide a set of candidate voice styles 340 and, based on the user's selection, apply a designated target voice style to the input voice content. Such a voice style may correspond to, for example, different timbres, intonations, and the like.
[0070] Taking FIG. 3D as an example, when the user selects a target voice style (eg, “style 1”), the voice content triggered to be played by the preview entry 335 may match the target voice style.
[0071] In some other embodiments, the electronic device 110 may also support the user to select a target voice style to be applied before inputting the voice content. Accordingly, after the voice content input is completed, the voice content triggered by the preview entry 335 may be, for example, the voice content to which the target voice style has been applied.
[0072] Furthermore, the electronic device 110 may create a corresponding voice expression based on the user's selection of the completion control 345. The voice expression may include a visual portion and a voice portion.
[0073] Accordingly, the visual portion may correspond to the image 325 shown in FIG3C , for example. The audio portion may correspond to the audio content played by the preview entry 335 .
[0074] In some embodiments, when the image 325 is a dynamic image, the voice content of the voice expression can be aligned with the dynamic image in time, so that the voice content can be played from the starting frame of the dynamic image, for example.
[0075] In some embodiments, the electronic device 110 may also add voice content in other appropriate ways. As an example, the electronic device 110 may provide a set of candidate voice content and add the candidate voice content to the voice expression based on the user's selection of the candidate voice content.
[0076] As an example, the electronic device 110 may obtain a set of voice materials, which may be created and shared by other users, or may include voice materials provided by the platform. Accordingly, the user may select a specific voice material and combine it with an existing image to generate a corresponding voice expression.
[0077] In some embodiments, voice content may also be added based on text content input by the user. For example, the electronic device 110 may receive text content input by the user and may generate voice content that matches the text content to add to the voice emoticon. Alternatively, the electronic device 110 may search for voice content that matches the text content from existing voice materials to add to the voice emoticon.
[0078] The above describes the process of creating a voice emoticon using an example where the target media resource includes an image. As mentioned above, the target media resource may also include, for example, a user-selected image emoticon. It should be understood that similar processes described above can be used to support editing the visual portion of the image emoticon and adding audio content to create a voice emoticon.
[0079] In some embodiments, the target media resource may further include a voice expression. In this case, the electronic device 110 may support the user to edit or replace the visual portion and / or the voice portion of the voice expression.
[0080] For example, the user can choose to retain the visual portion of an existing voice emoticon and record a new voice portion to generate a voice emoticon. Alternatively, the user can choose to retain the visual portion of an existing voice emoticon and apply a specific voice style to the existing voice portion to generate a voice emoticon.
[0081] In some embodiments, the target media resource may also include audio content, such as a voice clip. For example, electronic device 110 may receive a user's selection of a voice clip in a conversation to determine the voice clip as the target media resource. Alternatively, electronic device 110 may receive a user-recorded voice clip as the target media resource.
[0082] Accordingly, electronic device 110 may, for example, support the user in adding corresponding visual content. Similar to the discussion above regarding FIG. 3B , electronic device 110 may, for example, provide a set of candidate images for the user to select. In some examples, such candidate images may include, for example, images associated with the user. Alternatively, such candidate images may also include images generated based on the speech segment or searched for.
[0083] Based on the process discussed above, the embodiments of the present disclosure can support users in editing media resources to create voice expressions, thereby improving the efficiency of producing voice expressions.
[0084] 2 , in block 230 , the electronic device 110 publishes the voice expression so that the voice expression can be sent to a conversation associated with the target user based on the target user's selection.
[0085] In some embodiments, the electronic device 110 may publish the created voice expression accordingly based on the user's publishing operation.
[0086] In some examples, such publishing operation may include adding the voice emoticon to the emoticon panel. As shown in FIG3E , the electronic device 110 may correspondingly display the published voice emoticon 350 in the emoticon panel 300E so that the voice emoticon 350 may be sent to the associated conversation.
[0087] In some examples, such a publishing operation may include a sharing operation for a voice emoticon. For example, electronic device 110 may share the voice emoticon with other users or groups based on the user's sharing request. As an example, a user may share the voice emoticon with other users or groups, for example, through a private message or comment.
[0088] Alternatively, the electronic device 110 may also publish the voice emoticon as a resource available to other users based on a user's sharing request. Other users may further add the voice emoticon to support the use of such a voice emoticon in a conversation.
[0089] In another example, such a publishing operation may include publishing the voice expression 350 as a work (also referred to as a second work), so that users can add the voice expression 350 through the work. For example, an adding entry for adding an expression can be provided in the viewing interface of the work, and the current user or others can add the voice expression through the adding entry.
[0090] The process of using voice emoticons will be described in detail below with reference to FIG4 .
[0091] Second example process
[0092] FIG4 shows a flow chart of an example process 400 of expression interaction according to some embodiments of the present disclosure. Process 400 may be implemented at electronic device 110. Process 400 is described below with reference to FIG1.
[0093] At block 410 , the electronic device 110 receives a selection of a voice emoticon.
[0094] The process 400 will be described below with reference to Figures 5A and 5B. Figures 5A and 5B illustrate example interfaces 500A and 500B according to some embodiments of the present disclosure.
[0095] As shown in FIG5A , interface 500A may be, for example, a conversation interface (e.g., a conversation with “User B”). As shown, electronic device 110 may display an expression panel 505 in interface 500A. Expression panel 505 may, for example, allow selection of a set of voice expressions, such as voice expression 510-1 and voice expression 510-2 (individually or collectively referred to as voice expressions 510). Such voice expressions 510 may, for example, be created based on the process described above with reference to FIG2 .
[0096] As shown in FIG. 5A , the electronic device 110 may provide corresponding preview entries 515 - 1 and 515 - 2 in association with the voice expression 510 - 1 and the voice expression 510 - 2 .
[0097] For example, upon receiving a selection of the preview entry 515 - 1 , the electronic device 110 may play the voice portion of the voice expression 510 - 1 .
[0098] 4 , in block 420 , the electronic device 110 may send a voice expression to the target conversation, where the voice expression includes a visual portion and a voice portion.
[0099] Taking FIG. 5B as an example, upon receiving a selection of the voice expression 515 - 2 shown in FIG. 5A , the electronic device 110 may send the voice expression 515 - 2 to the conversation.
[0100] Specifically, the electronic device 110 may display the visual portion 520 of the voice expression 510-2 in the message window of the conversation. In the case where the visual portion 520 includes a dynamic image, the visual portion 520 may be automatically played once or multiple times.
[0101] In some embodiments, the electronic device 110 may further display a play entry 525 of the voice emoticon 510 - 2 in the message window. Further, upon receiving a selection of the play entry 525 , the electronic device 110 may play the voice portion of the voice emoticon 510 - 2 .
[0102] In some embodiments, when the visual portion 520 includes a dynamic image, when the play entry 525 is triggered, the dynamic image can be replayed, for example, so that the dynamic image can be synchronized with the playback of the voice portion.
[0103] It should be understood that the recipient of the voice expression can similarly view the visual portion 520 of the voice expression and can play the voice portion of the voice expression based on the selection of the play entry 525 .
[0104] Third example process
[0105] FIG6 shows a flow chart of an example process 600 for creating an emoticon according to some embodiments of the present disclosure. The process 600 may be implemented at the electronic device 110. The process 600 is described below with reference to FIG1.
[0106] As shown in the figure, in block 610 , the electronic device 110 presents a first group of candidate visual contents determined based on the target voice content, based on the target voice content.
[0107] In some embodiments, the target voice content may include voice content input by a user. For example, the electronic device 110 may obtain the voice content input by the user through a voice acquisition device.
[0108] In some embodiments, the target voice content may also include voice content selected by the user in other ways. For example, the electronic device 110 may receive a selection of a voice segment in a conversation and use that voice segment as the target voice content. Alternatively, the electronic device 110 may receive a selection of voice material provided by another user or platform and use that voice material as the target voice content.
[0109] In some embodiments, the first set of candidate visual content provided by the electronic device 110 is generated based on the text content corresponding to the target speech content. For example, the electronic device 110 can use any appropriate machine learning model to generate a set of candidate visual content corresponding to the text content, such as static images, dynamic images, or videos.
[0110] In box 620, the electronic device 110 creates a voice expression based on the selection of the first visual content in the first set of candidate visual contents, where the voice expression includes a visual part and a voice part, where the visual part is determined based on the first visual content and the voice part is determined based on the target voice content.
[0111] In some embodiments, the electronic device 110 may determine the speech portion of the voice emoticon based on the target voice content. For example, the target voice content itself may be used as the speech portion of the voice emoticon. Alternatively, the electronic device 110 may apply a specific voice style to the target voice content to generate the speech portion of the voice emoticon.
[0112] In some embodiments, electronic device 110 may determine the visual portion of the voice emoticon based on the selected first visual content. For example, the first visual content itself may be used as the visual portion of the voice emoticon. Alternatively, electronic device 110 may receive a user's editing operation on the first visual content and generate the visual portion of the voice emoticon accordingly.
[0113] Based on this approach, the embodiments of the present disclosure can support users to efficiently create matching voice expressions by recording voice.
[0114] Fourth example process
[0115] FIG7 shows a flow chart of an example process 700 for creating an emoticon according to some embodiments of the present disclosure. The process 700 may be implemented at the electronic device 110. The process 700 is described below with reference to FIG1.
[0116] As shown in the figure, in block 710 , the electronic device 110 presents a second set of candidate visual contents determined based on the target text content, based on the target text content.
[0117] In some embodiments, the electronic device 110 may receive target text content input or selected by a user, and may accordingly provide a second set of candidate visual content determined based on the target text content.
[0118] In some embodiments, the second set of candidate visual content may be, for example, searched from a visual content library based on the target text content. Alternatively or additionally, the second set of candidate visual content may be generated based on the target text content. For example, electronic device 110 may utilize any appropriate machine learning model to generate a set of candidate visual content corresponding to the text content, such as static images, dynamic images, or videos.
[0119] In box 720, the electronic device 110 creates a voice expression based on the selection of the second visual content in the second set of candidate visual contents, where the voice expression includes a visual part and a voice part, where the visual part is determined based on the target visual content and the voice part is generated based on the target text content.
[0120] In some embodiments, electronic device 110 may determine the visual portion of the voice emoticon based on the selected first visual content. For example, the first visual content itself may be used as the visual portion of the voice emoticon. Alternatively, electronic device 110 may receive a user's editing operation on the first visual content and generate the visual portion of the voice emoticon accordingly.
[0121] In some embodiments, electronic device 110 may generate the speech portion of the speech expression based on the target text content. For example, electronic device 110 may provide a set of candidate speech styles. Furthermore, electronic device 110 may receive a selection of a target speech style from the set of candidate speech styles, so that the speech portion of the speech expression is generated based on the target speech style.
[0122] Based on this approach, the embodiments of the present disclosure can support users to efficiently create matching voice expressions by inputting text.
[0123] Example devices and equipment
[0124] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG8 shows a schematic structural block diagram of an example apparatus 800 for expression interaction according to certain embodiments of the present disclosure. Apparatus 800 may be implemented as or included in electronic device 110. Each module / component in apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0125] As shown in Figure 8, the device 800 includes a determination module 810, which is configured to determine the target media resource for generating a voice emoticon; a creation module 820, which is configured to create a voice emoticon based on the received editing operation, where the voice emoticon includes a visual part and a voice part, and the editing operation includes at least one of the following: a first operation for modifying the first visual content and / or the first voice content included in the target media resource, and a second operation for inputting the second visual content and / or the second voice content; and a publishing module 830, which is configured to publish the voice emoticon so that the voice emoticon can be sent to a session associated with the target user based on the selection of the target user.
[0126] In some embodiments, the determination module 810 is further configured to: present a group of candidate images, the group of candidate images including static images and / or dynamic images; and determine the target image as the target media resource based on the selection of the target image from the group of candidate images.
[0127] In some embodiments, the set of candidate images includes at least one of the following: images collected by the current user; and images associated with works published by the current user.
[0128] In some embodiments, the visual portion of the voice emoticon is determined based on the target image.
[0129] In some embodiments, the apparatus 800 further includes a first receiving module configured to receive a second operation for inputting a second voice content, wherein the voice portion of the voice expression is determined based on the second voice content.
[0130] In some embodiments, the apparatus 800 further includes a second receiving module configured to: receive a first operation for editing the target image, wherein the visual portion of the voice expression is determined based on the edited target image.
[0131] In some embodiments, the apparatus 800 further includes a style determination module configured to: provide a set of candidate voice styles; and receive a selection of a target voice style from the set of candidate voice styles, such that the voice portion of the voice expression matches the target voice style.
[0132] In some embodiments, the target image is a target dynamic image, and the speech portion is configured to be played from a starting frame of the target dynamic image.
[0133] In some embodiments, the determination module 810 is further configured to present a set of candidate images based on the selection of the creation entry in the expression panel.
[0134] In some embodiments, the determination module 810 is further configured to: present a group of candidate image expressions; and determine the target image expression as the target media resource based on a selection of the target image expression from the group of candidate image expressions.
[0135] In some embodiments, the device 800 also includes a second receiving module configured to: receive a second operation for inputting a second voice content via an editing interface, wherein the visual part of the voice expression is determined based on the target image expression, and the voice part is determined based on the second voice content.
[0136] In some embodiments, the determination module 810 is further configured to: determine the first work as a target media resource based on a preset operation associated with the first work; or obtain a picture or video taken by the user as the target media content.
[0137] In some embodiments, the apparatus 800 further includes a sending module configured to: display the published voice expressions in the expression panel; and send the voice expressions to the conversation based on a selection of the voice expressions.
[0138] In some embodiments, the device 800 also includes a playback module configured to: display the visual part of the sent voice expression and the playback entrance of the voice expression in the message window of the conversation; and play the voice part of the voice expression based on the selection of the playback entrance.
[0139] In some embodiments, the apparatus 800 further includes a preview module configured to: provide a preview entry associated with the voice expression in the expression panel; and play the voice portion of the voice expression based on a selection of the preview entry.
[0140] In some embodiments, the publishing module 830 is further configured to perform at least one of the following: publishing the voice expression as a second work, wherein at least one user is enabled to add the voice expression via the second work; adding the voice expression so that the voice expression can be used for the current user; sharing the voice expression with at least one user.
[0141] FIG9 shows a block diagram of an electronic device 900 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 900 shown in FIG9 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 900 shown in FIG9 can be used to implement the electronic device 110 of FIG1 .
[0142] As shown in FIG9 , electronic device 900 is a general-purpose electronic device. Components of electronic device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. Processing unit 910 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 900.
[0143] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any accessible media that the electronic device 900 can access, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 930 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 900.
[0144] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0145] The communication unit 940 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0146] The input device 950 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 960 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 900 may also communicate with one or more external devices (not shown) through the communication unit 940 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 900, or with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0147] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0148] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0149] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0150] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0151] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0152] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for expression interaction, comprising: Determining a target media resource for generating a voice expression; Creating a voice expression based on a received editing operation, the voice expression including a visual part and a voice part, the editing operation including at least one of the following: a first operation for modifying first visual content and / or first voice content included in the target media resource, a second operation for inputting second visual content and / or second voice content; And Publishing the voice expression such that the voice expression can be sent to a session associated with the target user based on the selection of the target user.
2. The method according to claim 1, wherein determining the target media resource for generating the voice expression includes: Presenting a set of candidate images, the set of candidate images including static images and / or dynamic images; And Based on the selection of a target image from the set of candidate images, determining the target image as the target media resource.
3. The method according to claim 2, wherein the set of candidate images includes at least one of the following: Images collected by the current user; Images associated with works published by the current user.
4. The method according to claim 2, wherein the visual part of the voice expression is determined based on the target image.
5. The method according to claim 2, further comprising: Receiving the second operation for inputting the second voice content, wherein the voice part of the voice expression is determined based on the second voice content.
6. The method according to claim 5, further comprising: Receiving the first operation for editing the target image, wherein the visual part of the voice expression is determined based on the edited target image.
7. The method according to claim 5, further comprising: Providing a set of candidate voice styles; And Receiving the selection of a target voice style from the set of candidate voice styles such that the voice part of the voice expression matches the target voice style.
8. The method according to claim 5, wherein the target image is a target dynamic image, and the voice part is configured to play from the starting frame of the target dynamic image.
9. The method according to claim 2, wherein presenting a set of candidate images includes: Presenting the set of candidate images based on the selection of a creation entry in an expression panel.
10. The method according to claim 1, wherein determining the target media resource for generating the voice expression includes: Presenting a set of candidate image expressions; And Based on the selection of a target image expression from the set of candidate image expressions, determining the target image expression as the target media resource.
11. The method according to claim 10, further comprising: Receiving, via the editing interface, the second operation for inputting the second voice content, wherein the visual part of the voice expression is determined based on the target image expression, and the voice part is determined based on the second voice content.
12. The method according to claim 1, wherein determining the target media resource for generating the voice expression includes: Determine the first work as the target media resource based on a preset operation associated with the first work; Or Obtain a picture or video taken by the user as the target media content.
13. The method according to claim 1, further comprising: Display the published voice expression in the expression panel; And Send the voice expression to the session based on the selection of the voice expression.
14. The method according to claim 13, further comprising: Display the visual part of the sent voice expression and the playback entry of the voice expression in the message window of the session; And Play the voice part of the voice expression based on the selection of the playback entry.
15. The method according to claim 13, further comprising: Provide a preview entry associated with the voice expression in the expression panel; And Play the voice part of the voice expression based on the selection of the preview entry.
16. The method according to claim 1, wherein publishing the target expression includes at least one of the following: Publish the voice expression as a second work, enabling at least one user to add the voice expression via the second work; Add the voice expression so that the voice expression can be used by the current user; Share the voice expression with at least one user.
17. A method for expression interaction, comprising: Receive a selection of a voice expression; And Send the voice expression to a target session, the voice expression including a visual part and a voice part, wherein the voice expression is created based on an editing operation associated with a target media resource, and the editing operation includes at least one of the following: a first operation for modifying the first visual content and / or the first voice content included in the target media resource, and a second operation for adding a second visual content and / or a second voice content.
18. The method according to claim 17, wherein receiving the selection of the voice expression includes: Present an expression panel in the session interface of the target session; And Receive the selection of the voice expression in the expression panel.
19. The method according to claim 18, further comprising: Provide a preview entry associated with the voice expression in the expression panel; And Play the voice part of the voice expression based on the selection of the preview entry.
20. The method according to claim 17, further comprising: Display the visual part of the sent voice expression and the playback entry of the voice expression in the message window of the target session; And Play the voice part of the voice expression based on the selection of the playback entry.
21. A method for expression production, comprising: Present a first set of candidate visual contents determined based on the target voice content based on the target voice content; And Create a voice expression based on the selection of the first visual content in the first set of candidate visual contents, the voice expression including a visual part and a voice part, wherein the visual part is determined based on the first visual content and the voice part is determined based on the target voice content.
22. The method according to claim 21, wherein the first set of candidate visual contents is generated based on the text content corresponding to the target speech content.
23. A method for creating an expression, comprising: Presenting a second set of candidate visual contents determined based on the target text content based on the target text content; And Creating a speech expression based on a selection of a second visual content in the second set of candidate visual contents, the speech expression including a visual part and a speech part, wherein the visual part is determined based on the target visual content, and the speech part is generated based on the target text content.
24. The method according to claim 23, further comprising: Providing a set of candidate speech styles; And Receiving a selection of a target speech style in the set of candidate speech styles, such that the speech part of the speech expression is generated based on the target speech style.
25. An apparatus for expression interaction, comprising: A determination module configured to determine a target media resource for generating a speech expression; A creation module configured to create a speech expression based on a received editing operation, the speech expression including a visual part and a speech part, the editing operation including at least one of the following: a first operation for modifying a first visual content and / or a first speech content included in the target media resource, a second operation for inputting a second visual content and / or a second speech content; And A publishing module configured to publish the speech expression such that the speech expression can be sent to a session associated with the target user based on a selection of the target user.
26. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the electronic device to execute the method according to any one of claims 1 to 16, 17 to 20, 21 to 22 or 23 to 24.
27. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 16, 17 to 20, 21 to 22 or 23 to 24.
Citation Information
Patent Citations
Expression interaction method and device, equipment and storage medium
CN117834576A
Method and equipment for sending session message
CN110417641A
Sound expression application method, device and equipment and readable storage medium
CN111724799A
Voice expression display method and device and voice expression generation method and device
CN112910752A
Method and device for realizing voice message visualization service
WO2015117373A1