Information exchange method and apparatus, and electronic device and storage medium

By identifying and displaying information associated with audio, the problem of single functionality of instant messaging software is solved, and a more intelligent interactive experience is achieved.

WO2025195000A1PCT designated stage Publication Date: 2025-09-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/072848
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-18
Filing Date
2025-01-16
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Instant messaging software has single functions and cannot meet the diverse needs of users.

Method used

Intelligent interaction is achieved by collecting audio to be processed, identifying target audio and displaying target interaction information associated with it.

Benefits of technology

It improves the intelligent interaction effect and enhances the information display capability in the interactive interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072848_25092025_PF_FP_ABST
    Figure CN2025072848_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are an information exchange method and apparatus, and an electronic device and a storage medium. The information exchange method comprises: in response to a trigger operation of a user for a first control, collecting audio to be processed; when said audio comprises target audio, receiving target association information corresponding to the target audio; and displaying in an interaction interface target exchange information comprising the target association information. By means of the technical solution provided in the embodiments of the present disclosure, target exchange information associated with target audio can be displayed when it is determined that audio to be processed comprises the target audio, thereby achieving the effect of intelligent interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Information interaction method, device, electronic device and storage medium

[0001] This application claims priority to Chinese Patent Application No. 202410310046.7 filed on March 18, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The embodiments of the present disclosure relate to an information interaction method, device, electronic device, and storage medium. Background Art

[0003] With the development of the Internet, more and more users interact with each other through instant messaging software installed on smart terminals.

[0004] The main function of instant messaging software is to provide communication between users. It may have a single function and fail to meet the diverse needs of users. Summary of the Invention

[0005] The present disclosure provides an information interaction method, device, electronic device and storage medium, so that in an interactive scenario, when the collected audio to be processed includes target audio, target interaction information associated with the target audio can be displayed, thereby achieving an intelligent interactive effect.

[0006] In a first aspect, an embodiment of the present disclosure provides an information interaction method, the method comprising:

[0007] In response to a user triggering operation on the first control, collecting audio to be processed;

[0008] In a case where the audio to be processed includes target audio, receiving target association information corresponding to the target audio;

[0009] The target interaction information including the target association information is displayed in the interaction interface.

[0010] In a second aspect, an embodiment of the present disclosure further provides an information interaction device, the device comprising:

[0011] An audio acquisition module, configured to acquire audio to be processed in response to a user triggering operation on the first control;

[0012] an information determining module, configured to receive target-associated information corresponding to the target audio when the audio to be processed includes the target audio;

[0013] The interaction information display module is used to display the target interaction information including the target association information in the interaction interface.

[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0015] one or more processors;

[0016] a storage device for storing one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the information interaction method as described in any one of the embodiments of the present disclosure.

[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the information interaction method as described in any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0020] FIG1 is a schematic diagram of the main interface of an application program provided in an embodiment of the present disclosure;

[0021] FIG2 is a schematic diagram of a conversation interaction interface provided by an embodiment of the present disclosure;

[0022] FIG3 is a flow chart of an information interaction method provided by an embodiment of the present disclosure;

[0023] FIG4 is a flow chart of an information interaction method provided by an embodiment of the present disclosure;

[0024] FIG5 is a flow chart of an information interaction method provided by an embodiment of the present disclosure;

[0025] FIG6 is a schematic structural diagram of an information interaction device provided by an embodiment of the present disclosure; and

[0026] FIG7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0033] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0034] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0035] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0036] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0037] Before introducing the technical solutions provided by the embodiments of the present disclosure, an example description of the application scenarios can be given. The device for executing the information interaction method provided by the embodiments of the present disclosure can be integrated into an application software that supports interactive functions, and the software can be installed in an electronic device. Optionally, the electronic device can be a mobile terminal or a PC, etc. The application software can be software with interactive functions, and the interactive functions can include conversation functions. The specific application software will not be described here one by one, as long as the interactive functions can be realized. It can also be a specially developed application program to process the collected audio to be processed in a conversation scenario, and display the target interactive information including the processing results in the conversation interface.

[0038] It should also be noted that the technical solution provided by the embodiment of the present disclosure is to analyze and process the collected audio to be processed under the condition of user authorization.

[0039] If a corresponding application is developed based on the method provided in the embodiment of the present disclosure, or if the method provided in the embodiment of the present disclosure is integrated as a functional component into any application with interactive functions, the application can be installed in a mobile terminal or a PC. When it is detected that the application is triggered, the main interface corresponding to the application can be displayed. The main interface can only include controls corresponding to the function, or it can include multiple controls, one of which corresponds to the function control corresponding to the embodiment of the present disclosure. When it is detected that the function control is triggered, the session interaction interface can be entered to display the corresponding target interaction information in the session interaction interface.

[0040] It should also be noted that the solution provided by the embodiment of the present disclosure can be applied in any interactive scenario. Optionally, the interactive scenario can be a multi-person video scenario or a multi-person conversation scenario. Here, the interactive scenario is a conversation scenario as an example for explanation.

[0041] Exemplarily, referring to FIG1 , the main interface of the application may include at least one control to be triggered. The control to be triggered displays corresponding text content and / or image content, and the text content and / or image content is used to indicate which aspect of the interface corresponding to the content is displayed after the control to be triggered is triggered. It can be understood that the control to be processed corresponds to a column, and each column displays different interfaces and content after being triggered to achieve different interactions. In this embodiment, each column can correspond to a project. Accordingly, the project corresponding to the control to be processed that is triggered is used as the target project.

[0042] In this embodiment, the method further includes: displaying a conversation interaction interface corresponding to the target item in response to a user's trigger operation, so as to collect the audio to be processed when a trigger operation on a first control on the conversation interaction interface is detected.

[0043] It can be understood that a user can trigger the corresponding to-be-triggered control in any column of the main interface. The currently triggered column corresponding to the solution provided by the embodiments of this disclosure is designated as the target item. Upon detecting the control that triggers the display of the target item, a display interface corresponding to the target item can be displayed. This display interface can be a conversation interaction interface. The conversation interaction interface includes at least an information display area for displaying the conversation content and a control for triggering audio capture.

[0044] For an exemplary diagram of a conversation interaction interface, see Figure 2. The conversation interaction interface includes two information display areas and a control for triggering audio capture. When the conversation interaction interface is first displayed, a text description guiding interaction can be displayed in the information display area, guiding the user on how to interact within the conversation interaction interface based on the text description. The text description is presented from the perspective of the other user. That is, the conversation content of at least two users can be displayed in the conversation interaction interface, and the text description can be used as text sent by one of the users in the conversation.

[0045] Further, referring to Figure 2, taking the target project as a music radio column as an example, the conversation interaction interface can display a text description of the interaction guidance sent by one of the conversation users. The text description can be: You can listen to music here, welcome to the music radio, you can find the corresponding song according to your description. At this time, the user can trigger the audio acquisition control to collect the audio to be processed based on the triggering operation of the audio acquisition control. The audio acquisition control includes at least two types, one is a first control, and the other is a control for simulating a voice call. After detecting that any of the above controls is triggered, the system can receive an instruction to collect audio, based on which the audio acquisition module of the smart terminal, PC or application can be called to collect the audio to be processed. The audio acquisition module can be a microphone array.

[0046] The technical solution provided by the embodiment of the present disclosure can collect audio to be processed in a conversation scenario, analyze and process the audio to be processed, thereby retrieving target association information corresponding to the audio to be processed, and then generating target interaction information corresponding to the audio to be processed, thereby achieving the effect of intelligent interaction.

[0047] Figure 3 is a flow chart of an information interaction method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to any application scenario that has a conversation function and can collect audio to be processed. On the basis of the above embodiment, when the audio to be processed is collected and it is determined that the audio to be processed includes the target audio, target association information corresponding to the target audio can be received, and target interaction information including the target association information can be displayed in the conversation interaction interface. The specific implementation method can be found in the detailed description of the embodiment of the present disclosure, wherein the technical terms that are the same as or corresponding to the above embodiment are not repeated in this embodiment.

[0048] As shown in FIG3 , the method includes:

[0049] S110 : In response to a user triggering operation on a first control, collecting audio to be processed.

[0050] The conversation interaction interface may include a first control that, when triggered, calls the microphone array to collect corresponding audio data. When the microphone array is in operation, it can collect any audio data that can be picked up. This audio data may include voice data emitted by the user and ambient audio data of the user's environment. Ambient audio data may include music, movies, etc. played in the user's environment. In other words, the audio to be processed can be any audio data that can be collected by the microphone array.

[0051] For example, when the triggering of the first control 1 shown in FIG. 2 is detected, an audio collection instruction may pop up to indicate that the user can make a corresponding voice, or to indicate that the microphone array is currently capable of collecting audio. Alternatively, when the triggering of the first control 2 shown in FIG. 2 is detected, a virtual call interface may be entered to indicate that the microphone array is currently capable of collecting audio to be processed. In this case, the audio data collected by the microphone array may be used as the audio to be processed.

[0052] It should also be noted that the audio to be processed may include only the voice information sent by the user, or may include ambient audio, or may include both the voice information sent by the user and the ambient audio.

[0053] S120: When the audio to be processed includes target audio, receive target association information corresponding to the target audio.

[0054] The target audio refers to audio with specific characteristics. These specific characteristics may be associated with a pre-defined recognition event. For example, if the recognition event is a music recognition event, the specific characteristics may include at least rhythm characteristics, beat characteristics, and / or pitch characteristics. Target-related information refers to information associated with the target audio.

[0055] It can be understood that: when the audio to be processed includes the target audio, a processing result of the audio to be processed can be received, and the processing result mainly refers to target association information related to the target audio.

[0056] S130: Display target interaction information including the target association information in an interaction interface.

[0057] The target interaction information refers to the information obtained by encapsulating the target association information according to a preset template or format and is used to display it in the interaction interface. Optionally, the target interaction information can be displayed in the session interaction interface.

[0058] Specifically, after receiving the target association information, the target association information may be further processed to obtain target interaction information, which is used as the response information of the virtual user to the user corresponding to the audio to be processed.

[0059] Specifically, within the conversational interaction interface, the audio to be processed corresponding to the client user can be collected based on the voice conversation. If the audio to be processed includes target audio, target-related information associated with the target audio can be received. Functions integrated within the application can further encapsulate this target-related information to obtain target interaction information that can be displayed within the conversational interaction interface, providing feedback based on the voice conversation.

[0060] The technical solution provided by the embodiment of the present disclosure can collect the audio to be processed and send the audio to be processed to the target server when a trigger operation of the user on the first control is detected, so that the target server can determine whether the audio to be processed includes the target audio. Under the condition that the server determines that the audio to be processed includes the target audio, the target association information associated with the target audio can be determined, and the target association information can be sent to the client to encapsulate the target association information based on the client to obtain the target interaction information that can be displayed in the display interface, and display it, thereby effectively identifying the audio to be processed and extracting the target interaction information associated with the target audio, thereby improving the effect of intelligent interaction.

[0061] FIG4 is a flow chart of an information interaction method provided by an embodiment of the present disclosure. Based on the aforementioned embodiment, it can be known that, under the condition that the target audio is included in the audio to be processed, target association information corresponding to the target audio can be received. Based on this, an application or a server associated with the application can determine whether the audio to be processed includes the target audio, and then determine the target association information corresponding to the target audio. The specific implementation method can be found in the detailed description of the embodiment of the present disclosure. Among them, the technical terms that are the same as or corresponding to the above embodiments are not repeated here.

[0062] S210: In response to a user triggering operation on a first control, collecting audio to be processed.

[0063] S220: Send the audio to be processed to a target server, so that the target server determines whether the audio to be processed includes target audio.

[0064] The target server can be understood as the server corresponding to the application, or a server capable of analyzing and processing the audio to be processed. Optionally, the target server can be a server integrated with an event judgment function, mainly used to determine whether the audio to be processed includes the target audio.

[0065] In this embodiment, determining whether the target audio is included in the audio to be processed can be: analyzing and processing the audio to be processed based on a pre-trained event judgment model, and outputting an event judgment result of whether the target audio is included; based on the event judgment result, determining whether the target audio is included in the audio to be processed.

[0066] Among them, the event judgment model can be understood as a pre-trained model for determining whether the target audio is included in the audio to be processed. The event judgment model can be a two-classification model, and accordingly, the event judgment result can be a result of including the target audio or not including the target audio. The event judgment model can be a model for determining whether the duration of the target audio in the audio to be processed reaches a preset duration threshold, and accordingly, the event judgment result can be whether the target audio is included in the audio to be processed or not. Of course, the event judgment model can also be a model for judging the probability of at least one event occurring in the audio to be processed, and accordingly, the event judgment result can be the probability corresponding to at least one event pre-set in the audio to be processed, so as to determine whether the target audio is included in the audio to be processed based on the probability.

[0067] Exemplarily, the target audio is audio that includes music features. The event judgment model can analyze the audio to be processed to determine whether the audio to be processed includes music and the duration corresponding to the music. If the output result includes music (target audio) and the duration exceeds a preset duration threshold, optionally, the preset duration threshold is 6S, and it is determined that the audio to be processed includes the target audio. The event judgment model can also be obtained by training based on a plurality of pre-set events and the labels corresponding to each event. Then, after the audio to be processed is input into the event judgment model, the event judgment model can output the probability of each pre-set event occurring. Based on each probability value, it can be determined whether the audio to be processed includes the target audio corresponding to the music features.

[0068] It should be noted that the at least one pre-set event may be a music event, a story event, a news event, and the like.

[0069] It should also be noted that the event judgment models listed above are merely exemplary. As long as the judgment of a specific event can be achieved, it will suffice. In this embodiment, the specific event may be a music event.

[0070] S230: If the audio to be processed includes target audio, determine target association information corresponding to the target audio and feed it back.

[0071] It should be noted that the server that determines whether the audio to be processed includes the target audio and the server that determines the target association information corresponding to the target audio can be the same server or different server. Whether they are the same or not is not limited in this embodiment, as long as they can implement the same function.

[0072] Optionally, the target audio may be audio corresponding to the rhythm of the music, and the target-related information may include information associated with the target music track, such as the target music track's title, album, release date, artist, composer, album cover, playable target music tracks, etc. The content that can be displayed in the target-related information is pre-set and can be configured by backend personnel based on actual needs, but at least includes the playable target music tracks corresponding to the target audio.

[0073] After the target association information is determined, the target association information may be fed back to the application program so that the application program processes the target association information to obtain target interaction information.

[0074] In this embodiment, determining the target association information corresponding to the target audio may be: determining the target association information corresponding to the target audio according to a preset priority of at least one recognition type.

[0075] If the target audio corresponds to music, there may be a scenario where the music is performed by multiple users, that is, a music may be performed by the original singer, a cover version, or a humming version. Accordingly, the target-related information found may include multiple pieces of information. In this embodiment, all target-related information may be fed back, or only one piece of target-related information may be displayed.

[0076] Of course, in order to constrain the target association information of the feedback, a method for determining the target association information can be set. Optionally, the target association information corresponding to the target audio is determined based on the priority corresponding to the recognition type, so that once the target association information is determined, it can be fed back.

[0077] It can be understood that after determining that the audio to be processed includes the target audio, the audio to be processed can be sent to the target server (song recognition server) to determine the target associated information. The target server can determine the target associated information corresponding to the music recognition type mentioned above.

[0078] The recognition type may be a music recognition type. Optionally, the recognition type may include at least one of a fingerprint recognition type, a cover recognition type, and a humming recognition type. Target association information may be determined based on the aforementioned recognition type methods. Alternatively, the target association information may be determined based on the priority corresponding to each of the aforementioned recognition types. When priorities are set, the priority order from highest to lowest is fingerprint recognition type, cover recognition type, and humming recognition type.

[0079] Specifically, the target association information that matches the target audio in the audio to be processed can be searched from the music library based on the fingerprint recognition type. If the target association information is found, it can be fed back. If not, the target association information that matches the target audio in the audio to be processed can be searched from the cover music library based on the cover recognition type. Similarly, if the target association information is found, it can be fed back. If the target association information is not found, the cover song recognition type can be used to search the music library to see if the target audio is included in the cover song.

[0080] It should be noted that, whether it is a cover song, a hummed song, or a song corresponding to the original version, they can all be stored in the same music library, or they can be stored in music libraries corresponding to different recognition types to improve search efficiency.

[0081] S240: Receive target association information corresponding to the target audio.

[0082] The target association information may include the song to be played, the singer, the album name corresponding to the target music track corresponding to the target audio, and the playing time corresponding to the target audio in the target music track, etc.

[0083] After receiving the target association information, the method further includes: acquiring historical conversation content in the conversation interaction interface; and determining the target interaction information based on the historical conversation content, the target association information, and a preset feedback template.

[0084] This can be understood as obtaining historical conversation content from the conversation interaction interface. The target-related information is then processed based on the historical conversation content and a pre-defined feedback template to obtain target interaction information. Optionally, the target interaction information may include the album cover of the target music track, the artist and release date of the target music track, the time the target audio track was played within the target music track, and whether you are searching for a text description of the song.

[0085] Of course, if the system can find target association information under different identification types, the target interaction information can also include whether to display other information found that matches the target music track, or it can also display descriptive information to guide the user to continue the operation, such as clicking on the album cover to play the description of the target music track.

[0086] S250: Display the target interaction information in the information display area of ​​the interaction interface.

[0087] It can be understood that after the target interaction information is determined, the target interaction information can be displayed in the information display area.

[0088] It should also be noted that the target interaction information is the conversation information fed back by the virtual user.

[0089] In this embodiment, the information display area in the conversation interaction interface is used to display the conversation content corresponding to at least one conversation user. For example, upon first entering the conversation interaction interface, a conversation message may be displayed in the conversation interaction interface. This conversation message contains descriptive text guiding the conversation interaction. The conversation user corresponding to this conversation message may be a virtual conversation user, primarily corresponding to an artificial intelligence. In other words, the virtual conversation user primarily provides feedback on target interaction information found by the artificial intelligence.

[0090] Of course, when there are multiple users participating in the conversation, the audio data to be processed can also be analyzed and processed to display the target interaction information.

[0091] The technical solution provided by the embodiment of the present disclosure can collect the audio to be processed and send the audio to be processed to the target server when a trigger operation of the user on the first control is detected, so that the target server can determine whether the audio to be processed includes the target audio. Under the condition that the server determines that the audio to be processed includes the target audio, the target association information associated with the target audio can be determined, and the target association information can be sent to the client, so as to encapsulate the target association information based on the client, obtain the target interaction information that can be displayed in the display interface, and display it, which solves the problem that in the conversation scenario, it can only be a conversation reply between at least two users, and does not involve the analysis and processing of the collected audio to be processed, so as to determine that the analysis result is that the audio to be processed includes a specific event, and determine the target association information corresponding to the event to be processed, thereby effectively identifying the audio to be processed and extracting the target interaction information associated with the target audio, thereby improving the effect of intelligent interaction.

[0092] As an optional embodiment of the above embodiment, the target audio can be taken as music audio for example. Figure 5 is a flow chart of an information interaction method provided by the embodiment of the present disclosure. The technical terms that are the same as or corresponding to the above embodiment will not be repeated in this embodiment.

[0093] As shown in Figure 5, the user can trigger the control corresponding to the target item (music station) displayed in the main interface of the application to display the conversation interaction interface, as shown in Figure 2. After detecting that the first control is triggered, the user's voice information or the ambient audio including the user's voice information can be collected. After obtaining the audio to be processed, the audio to be processed can be transparently transmitted to the music artificial intelligence (the server that determines the audio).

[0094] Music AI can use an event judgment model to determine whether music audio (target audio) exists. The event judgment model can be processed as follows: if the duration of the music audio in the audio to be processed exceeds a preset duration threshold, the judgment result is that the target audio is included; or, the event judgment model can analyze the audio to be processed and output the recognition probability value corresponding to each event. If the probability value of the presence of a music event is greater than the preset probability threshold, it means that the audio to be processed includes the target audio.

[0095] If it is determined that the audio to be processed includes the target audio, the audio to be processed can be sent to the music recognition server (target server). The target server can determine the target music track corresponding to the target audio and the target association information associated with the target music track based on the pre-set recognition priority. The recognition priority can be fingerprint recognition, cover recognition, and humming recognition in descending order.

[0096] The target association information can be used as target recognition result information, and the target recognition result can be fed back to the application. After receiving the target association information, it can be processed based on the audio to be processed, historical conversation content, etc. to obtain target interaction information. The target interaction information is displayed in the conversation interaction interface as feedback messages from the virtual conversation user on the audio to be processed. In other words, the target interaction information is the conversation message feedback from the virtual conversation user corresponding to the audio to be processed.

[0097] The technical solution provided by the embodiment of the present disclosure can collect the audio to be processed and send the audio to be processed to the target server when a trigger operation of the user on the first control is detected, so that the target server can determine whether the audio to be processed includes the target audio. Under the condition that the server determines that the audio to be processed includes the target audio, the target association information associated with the target audio can be determined, and the target association information can be sent to the client, so as to encapsulate the target association information based on the client, obtain the target interaction information that can be displayed in the display interface, and display it, which solves the problem that in the conversation scenario, it can only be a conversation reply between at least two users, and does not involve the analysis and processing of the collected audio to be processed, so as to determine that the analysis result is that the audio to be processed includes a specific event, and determine the target association information corresponding to the event to be processed, thereby effectively identifying the audio to be processed and extracting the target interaction information associated with the target audio, thereby improving the effect of intelligent interaction.

[0098] FIG6 is a schematic diagram of the structure of an information interaction device provided by an embodiment of the present disclosure. As shown in FIG6 , the device includes: an audio acquisition module 310 , an information determination module 320 , and an interactive information display module 330 .

[0099] Among them, the audio acquisition module 310 is used to collect the audio to be processed in response to the user's trigger operation on the first control; the information determination module 320 is used to receive target association information corresponding to the target audio when the audio to be processed includes the target audio; the interaction information display module 330 is used to display the target interaction information including the target association information in the interaction interface.

[0100] Based on the above technical solution, the device also includes: a main interface display module, which is used to display an interactive interface corresponding to the target item in response to a user's trigger operation, so as to collect the audio to be processed when a trigger operation of the first control on the interactive interface is detected.

[0101] Based on the above technical solution, the device also includes: an audio sending module, which is used to send the audio to be processed to the target server, so that the target server can determine the target association information corresponding to the target audio and feedback it when determining that the audio to be processed includes the target audio.

[0102] On the basis of the above technical solution, the device further includes:

[0103] An event judgment module is used to analyze and process the audio to be processed based on a pre-trained event judgment model, and output an event judgment result indicating whether the audio includes the target audio;

[0104] The audio determination module is configured to determine whether the audio to be processed includes the target audio based on the event determination result.

[0105] On the basis of the above technical solution, the device further includes: an information processing module, configured to determine target association information corresponding to the target audio according to a preset priority of at least one recognition type.

[0106] On the basis of the above technical solution, the target audio is audio corresponding to the rhythm and beat of the music, and the target association information includes information associated with the target music track.

[0107] On the basis of the above technical solution, the recognition type includes at least one of a fingerprint recognition type, a cover song recognition type and a humming recognition type.

[0108] On the basis of the above technical solution, the device further includes:

[0109] A session content acquisition module, configured to acquire historical session content in the session interaction interface;

[0110] The interaction information generating module is used to determine the target interaction information according to the historical conversation content, target association information and a preset feedback template.

[0111] On the basis of the above technical solution, the device further includes: an interaction information display module, configured to display the target interaction information in an information display area in the conversation interaction interface.

[0112] On the basis of the above technical solution, the device further includes: the conversation interaction interface includes at least two conversation users, the at least two conversation users include a virtual conversation user, and the virtual conversation user is used to feed back the target interaction information.

[0113] The technical solution provided by the embodiment of the present disclosure can collect the audio to be processed and send the audio to be processed to the target server when a trigger operation of the user on the first control is detected, so that the target server can determine whether the audio to be processed includes the target audio. Under the condition that the server determines that the audio to be processed includes the target audio, the target association information associated with the target audio can be determined, and the target association information can be sent to the client, so as to encapsulate the target association information based on the client, obtain the target interaction information that can be displayed in the display interface, and display it, which solves the problem that in the conversation scenario, it can only be a conversation reply between at least two users, and does not involve the analysis and processing of the collected audio to be processed, so as to determine that the analysis result is that the audio to be processed includes a specific event, and determine the target association information corresponding to the event to be processed, thereby effectively identifying the audio to be processed and extracting the target interaction information associated with the target audio, thereby improving the effect of intelligent interaction.

[0114] The task processing device provided by the embodiments of the present disclosure can execute the information interaction method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0115] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0116] FIG7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG7 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG7 ) 400 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG7 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0117] As shown in FIG7 , the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An edit / output (I / O) interface 405 is also connected to the bus 404.

[0118] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 illustrates the electronic device 400 with various devices, it should be understood that not all of the illustrated devices are required to be implemented or present. More or fewer devices may alternatively be implemented or present.

[0119] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0120] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0121] The electronic device provided by the embodiment of the present disclosure and the information interaction method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0122] An embodiment of the present disclosure provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the information interaction method provided by the above embodiment is implemented.

[0123] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0124] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0125] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0126] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0127] In response to a user triggering operation on the first control, collecting audio to be processed;

[0128] In a case where the audio to be processed includes target audio, receiving target association information corresponding to the target audio;

[0129] The target interaction information including the target association information is displayed in the interaction interface.

[0130] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0132] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0133] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0134] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0135] According to one or more embodiments of the present disclosure, Example 1 provides an information interaction method, which is applied to a robot. The method includes:

[0136] In response to a user triggering operation on the first control, collecting audio to be processed;

[0137] In a case where the audio to be processed includes target audio, receiving target association information corresponding to the target audio;

[0138] The target interaction information including the target association information is displayed in the interaction interface.

[0139] According to one or more embodiments of the present disclosure, Example 2 provides an information interaction method, the method further comprising:

[0140] Optionally, in response to a trigger operation of the user, a conversation interaction interface corresponding to the target item is displayed, so that when a trigger operation of a first control on the conversation interaction interface is detected, audio to be processed is collected.

[0141] According to one or more embodiments of the present disclosure, Example 3 provides an information interaction method, the method further comprising:

[0142] Optionally, the audio to be processed is sent to a target server, so that when the target server determines that the audio to be processed includes target audio, it determines target association information corresponding to the target audio and feeds back the information.

[0143] According to one or more embodiments of the present disclosure, Example 4 provides an information interaction method, the method further comprising:

[0144] Optionally, determining whether the audio to be processed includes target audio includes:

[0145] Analyze and process the audio to be processed based on a pre-trained event judgment model, and output an event judgment result indicating whether the audio includes the target audio;

[0146] Based on the event judgment result, it is determined whether the audio to be processed includes the target audio.

[0147] According to one or more embodiments of the present disclosure, Example 5 provides an information interaction method, the method further comprising:

[0148] Optionally, determining target association information corresponding to the target audio includes:

[0149] Target association information corresponding to the target audio is determined according to a preset priority of at least one recognition type.

[0150] According to one or more embodiments of the present disclosure, Example 6 provides an information interaction method, the method further comprising:

[0151] Optionally, the target audio is audio corresponding to the rhythm and beat of the music, and the target association information includes information associated with the target music track.

[0152] According to one or more embodiments of the present disclosure, Example 7 provides an information interaction method, the method further comprising:

[0153] Optionally, the recognition type includes at least one of a fingerprint recognition type, a cover recognition type, and a humming recognition type.

[0154] According to one or more embodiments of the present disclosure, Example 8 provides an information interaction method, the method further comprising:

[0155] Optionally, after receiving target association information corresponding to the target audio, the method further includes:

[0156] Obtaining historical content in the interactive interface;

[0157] The target interaction information is determined based on the historical content, target association information and a preset feedback template.

[0158] According to one or more embodiments of the present disclosure, Example 9 provides an information interaction method, the method further comprising:

[0159] Optionally, displaying the target interaction information including the target association information in the session interaction interface includes:

[0160] The target interaction information is displayed in an information display area in the session interaction interface.

[0161] According to one or more embodiments of the present disclosure, Example 10 provides an information interaction method, the method further comprising:

[0162] Optionally, the interaction interface is a conversation interaction interface, the conversation interaction interface includes at least two conversation users, the at least two conversation users include a virtual conversation user, and the virtual conversation user is used to feed back the target interaction information.

[0163] According to one or more embodiments of the present disclosure, Example 11 provides an information interaction device, the device including:

[0164] An audio acquisition module, configured to acquire audio to be processed in response to a user triggering operation on the first control;

[0165] an information determining module, configured to receive target-associated information corresponding to the target audio when the audio to be processed includes the target audio;

[0166] The interaction information display module is used to display the target interaction information including the target association information in the conversation interaction interface.

[0167] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0168] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0169] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An information interaction method, comprising: In response to a user triggering operation on the first control, collecting audio to be processed; In a case where the audio to be processed includes target audio, receiving target association information corresponding to the target audio; The target interaction information including the target association information is displayed in the interaction interface.

2. The method according to claim 1, further comprising: In response to a trigger operation by the user, an interactive interface corresponding to the target item is displayed, so that when a trigger operation on a first control on the interactive interface is detected, audio to be processed is collected.

3. The method according to claim 1, further comprising: The audio to be processed is sent to a target server, so that the target server, when determining that the audio to be processed includes the target audio, determines target association information corresponding to the target audio and feeds back the information.

4. The method according to claim 1, wherein Determining whether the audio to be processed includes the target audio includes: Analyze and process the audio to be processed based on a pre-trained event judgment model, and output an event judgment result indicating whether the audio includes the target audio; Based on the event judgment result, it is determined whether the audio to be processed includes the target audio.

5. The method according to claim 3, wherein The determining target association information corresponding to the target audio includes: Target association information corresponding to the target audio is determined according to a preset priority of at least one recognition type.

6. The method according to any one of claims 1 to 5, wherein: The target audio is audio corresponding to the rhythm of the music, and the target association information includes information associated with the target music track, and the target music track corresponds to the target audio.

7. The method according to claim 5, wherein: The recognition type includes at least one of a fingerprint recognition type, a cover recognition type, and a humming recognition type.

8. The method according to any one of claims 1 to 7, wherein: After receiving the target association information corresponding to the target audio, the method further includes: Obtaining historical interaction content in the interaction interface; The target interaction information is determined according to the historical interaction content, target association information and a preset feedback template.

9. The method according to any one of claims 1 to 8, wherein: The interaction interface includes a conversation interaction interface, the conversation interaction interface includes conversation messages corresponding to at least one conversation user, the conversation users include virtual conversation users, and the virtual conversation users are used to feed back the target interaction information.

10. An information interaction device, comprising: an audio collection module, configured to collect audio to be processed in response to a user triggering operation on the first control; an information determining module, configured to receive target-associated information corresponding to the target audio when the audio to be processed includes the target audio; The interaction information display module is configured to display target interaction information including the target association information in the interaction interface.

11. An electronic device comprising: one or more processors; A storage device for storing one or more programs, wherein: When the one or more programs are executed by the one or more processors, the one or more processors implement the information interaction method according to any one of claims 1 to 9.

12. A storage medium containing computer-executable instructions, wherein: When the computer executable instructions are executed by a computer processor, they are used to execute the information interaction method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Information interaction method and device, electronic equipment and storage medium

    CN120670615A

  • Method and apparatus for pushing music information based on instant messaging

    CN106559469A

  • Song processing method

    CN111404808A

  • Song recognition method and device

    CN112148754A

  • Voice interactive system and voice interactive method

    JP2018091911A