Voice application recognition method, device, equipment, medium and product
In device interconnection scenarios, the source end listens for and sends audio focus and media playback status information, while the target end displays the corresponding sound identifier. This solves the problem that users cannot distinguish the current sound application of the other end device, achieving fast and accurate sound application identification, and improving user experience and system operating efficiency.
Patent Information
- Application Number
- CN202610509102.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-16
- Publication Date
- 2026-07-21
AI Technical Summary
In interconnected device scenarios, users cannot accurately distinguish the application currently making a sound from the other device, which affects the user experience.
After the target end and the source end establish an interconnection, the source end listens for and sends audio focus information and media playback status information. The target end then displays the corresponding sound identifier in the display interface based on this information to represent the sound status of the application.
It enables rapid and accurate identification of voice-generating applications in device interconnection scenarios, improving user experience and intuitive interaction, reducing processor resource consumption on the target end, and ensuring the real-time performance and accuracy of status recognition.
Smart Images

Figure CN122435930A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of device interconnection technology, specifically to sound application identification methods, devices, equipment, media, and products. Background Technology
[0002] With the widespread adoption of Android devices, the number of applications on these devices is also increasing. In device interconnection scenarios, Android devices need to compete with each other for speaker resources. Users often cannot distinguish which application is currently making noise on the other device, affecting the user experience. Summary of the Invention
[0003] In view of this, this disclosure provides a method, apparatus, device, medium and product for identifying sound-emitting applications, in order to solve the problem in related technologies where users often cannot distinguish the current sound-emitting application of the other end device in device interconnection scenarios, which affects the user experience.
[0004] In a first aspect, this disclosure provides a method for identifying voice-generating applications, the method comprising:
[0005] In response to establishing an interconnection with the source, at least one application from the source is displayed in the display interface; In response to receiving first audio focus information and first media playback status information of the first application sent by the source end, a first sound identifier is displayed in the display interface based on the first audio focus information and the first media playback status information. The first application is any one of the at least one application, and the first sound identifier is used to characterize the sound status of the first application.
[0006] This embodiment of the disclosure establishes an interconnection between the target end and the source end, displays the source end application, and displays a first sound identifier on the display interface based on the first audio focus information and the first media playback status information sent by the source end. This allows for a clear representation of the sound status of the first application on the source end directly on the target end interface. Compared to traditional methods that require users to make their own judgments or additionally detect audio data, this solution eliminates the need for the target end to sample and analyze the audio stream, avoiding the consumption of target end processor resources. Furthermore, relying on the dual information of audio focus and media playback status ensures the accuracy of sound status recognition, enabling users to quickly and intuitively distinguish sound-emitting applications in device interconnection scenarios, thus improving the interconnection user experience and intuitiveness of interaction.
[0007] In one optional implementation, displaying the first sound identifier in the display interface based on the first audio focus information and the first media playback status information includes: In response to the first audio focus information indicating that the first application has focus, and the first media playback status information indicating that the first application is in playback state, a first form of the first sound identifier is displayed in the display interface, the first form being used to indicate that the first application is emitting sound; and / or, In response to the first audio focus information indicating that the first application has no focus, and / or the first media playback status information indicating that the first application is in a non-playback state, a second form of the first sound identifier is displayed in the display interface, the second form being used to indicate that the first application is not making a sound.
[0008] In this embodiment, the first and second forms of the first sound identifier are displayed based on the audio focus and media playback status, respectively. By distinguishing between the sound-emitting and non-sound-emitting states through a rule that the sound-emitting state corresponds to being in focus and playing, and the non-sound-emitting state corresponds to being out of focus or not playing, a visual distinction is achieved. This setting allows users to quickly determine whether the application is emitting sound based on the identifier form, without additional operation or judgment, simplifying the user identification process. Simultaneously, a standardized status determination logic is adopted to adapt to various application sound-emitting scenarios, ensuring consistency in status display and improving the clarity and ease of use of the target terminal interface.
[0009] In one optional implementation, the first form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in the display interface using an animated display method, and the first form includes the animated display method.
[0010] In this embodiment, the first form of the sound indicator adopts a dynamic display method. The dynamic visual effect is more prominent and can quickly attract the user's attention on the target device's display interface, allowing the user to immediately perceive that the application is in a sound state and avoid missing the sound application prompt. Furthermore, the dynamic display enhances the status reminder effect, adapts to the user's habit of quickly browsing the interface in the device interconnection scenario, and can identify the sound application without prolonged staring, improving the effectiveness and recognizability of the status prompt. At the same time, the implementation of the dynamic effect does not require a large amount of additional system resources, balancing the prompt effect and device performance.
[0011] In one optional implementation, the second form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in a static display mode in the display interface, and the second form includes a static display mode.
[0012] In this embodiment, the second-form identifier is displayed statically when no sound is emitted. The static display style is simple and stable, forming a sharp visual contrast with the animated sound identifier, further enhancing the distinction between the sound and non-sound states. Static display eliminates the need for continuous animation rendering, reducing the resource consumption of the target interface and improving interface smoothness. At the same time, the static identifier is clear and stable, without causing visual interference to the user, adapting to long-term display scenarios, ensuring the simplicity and readability of the non-sound state prompt, and making the interface state display more reasonable and comfortable.
[0013] In one optional implementation, the first form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed at a first preset position on the display interface, and the first form includes the first preset position.
[0014] In this embodiment, the sound identifier is displayed in a first preset position on the display interface. This fixed position helps users develop a consistent visual habit, eliminating the need to search for the sound identifier and allowing for quick location of the application's sound status. The preset position layout conforms to user visual browsing patterns, improving status recognition efficiency. Simultaneously, the fixed position ensures a neat and uniform interface layout, avoiding recognition confusion caused by random changes in identifier position. It adapts to scenarios where multiple applications are displayed simultaneously, accurately corresponding to the sound status of the first application, and enhancing the standardization and convenience of interface interaction.
[0015] In one optional implementation, the second form of displaying the first sound identifier in the display interface includes: The first sounding identifier is displayed at a second preset position on the display interface, and the second form includes the second preset position.
[0016] In this embodiment, the non-voiced identifier is displayed in a second preset position. By differentiating itself from the voiced identifier by this preset position, the visual difference between the voiced and non-voiced states is further enhanced, avoiding confusion. The dual preset positions make the interface display more organized, allowing users to quickly determine the application's voicing status and simplifying the identification process. Simultaneously, the position differentiation requires no additional interface elements, does not increase interface complexity, ensures a clean display, and adapts to the simultaneous display needs of multiple applications in interconnected device scenarios, improving the accuracy and intuitiveness of the status display.
[0017] In one optional implementation, the first form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in the display interface as a first interface display element, and the first form includes the first interface display element.
[0018] In this embodiment, the sound status is displayed using a first interface element. This dedicated interface element highlights the sound status, providing a clear visual distinction from the silent state and allowing users to quickly identify when the application is speaking. The first interface element can be adapted to the target device's interface style, ensuring overall aesthetics. Simultaneously, the dedicated element is highly recognizable, unaffected by interference from other interface content, and improves the accuracy of the sound status prompt. This setting requires no complex interface modifications, resulting in low implementation costs, while simultaneously enhancing the intuitiveness of the status display and optimizing the user's online experience.
[0019] In one optional implementation, the second form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in the display interface as a second interface display element. The second form includes a second interface display element, which is different from the first interface display element.
[0020] In this embodiment, the non-voice state uses a second interface display element that differs from the first interface element. This dual-interface design accurately distinguishes between the voice-emitting and non-voice-emitting states, avoiding confusion. The differentiated design of the two interface elements creates a strong visual contrast, allowing users to quickly determine the application status and improving recognition efficiency. Simultaneously, the differentiated elements are compatible with the target device's interface design specifications, ensuring overall interface consistency. Clear status display is achieved without additional system resources, resolving the issue of ambiguous recognition of voice-emitting applications in interconnected device scenarios and enhancing the interactive experience.
[0021] In an optional implementation, the method further includes: In response to receiving the second audio focus information and the second media playback status information of the first application sent by the source end, the first sound identifier is updated and displayed in the display interface based on the second audio focus information and the second media playback status information.
[0022] This embodiment of the disclosure achieves real-time synchronization and dynamic updating of the sound status by adding a step of updating the sound identifier based on the updated audio focus and media playback status information. When the sound status of the source application changes, the target device can immediately update the identifier display, ensuring the real-time and accurate display of the status and avoiding recognition errors caused by status lag. Furthermore, this update mechanism does not require manual triggering by the user; it automatically completes status synchronization, adapting to dynamic scenarios such as application playback, pause, and focus switching, maintaining consistency between the sound status and the actual situation throughout, and improving the smoothness and reliability of device interconnection.
[0023] Secondly, this disclosure provides a method for identifying sound-generating applications, applied at the source end, the method comprising: In response to establishing an interconnection with the target end, the first audio focus information and the first media playback status information of the first application are detected, wherein the first application is the application of the source end and the first application is displayed in the display interface of the target end; The first audio focus information and the first media playback status information of the first application are sent to the target terminal, so that the target terminal displays a first sound identifier in the display interface based on the first audio focus information and the first media playback status information. The first sound identifier is used to characterize the sound status of the first application.
[0024] This embodiment detects the application's audio focus and media playback status after the source and target ends are interconnected and sends them to the target end. The source end is responsible for status acquisition, eliminating the need for audio detection and analysis on the target end, thus reducing the computational burden on the target end. The source end directly acquires system-level audio focus and media session information, ensuring high accuracy and avoiding errors and resource consumption in audio sampling judgment. Simultaneously, it proactively sends status information, ensuring timely information transmission and supporting the target end to quickly display the sound emission indicator. This achieves synchronization of sound emission status between the source and target ends, resolving the issue of unrecognizable sound applications in device interconnection scenarios and improving the overall efficiency and accuracy of the solution.
[0025] In an optional implementation, the method further includes: In response to detecting an update to the first audio focus information and / or the first media playback status information of the first application, the updated second audio focus information and the second media playback status information of the first application are sent to the target terminal, so that the target terminal updates the display of the first sound identifier in the display interface based on the second audio focus information and the second media playback status information.
[0026] In this embodiment, the source device immediately sends the latest information upon detecting a state update, achieving real-time synchronous transmission of the sound state. When the application audio focus or playback state changes, the source device triggers information transmission instantly, ensuring that the information received by the target device is not delayed compared to the actual state of the source device, thus avoiding display errors caused by state lag. This dynamic update mechanism does not require continuous polling and detection; it only transmits data when the state changes, reducing data transmission volume and system resource consumption. At the same time, it ensures accurate and timely state synchronization, adapts to various dynamic sound scenarios, and improves the smoothness and stability of device interconnection state synchronization.
[0027] Thirdly, this disclosure provides a voice application recognition system, the system comprising: a source end and a target end, wherein the source end and the target end are interconnected, and at least one application of the source end is displayed in the display interface of the target end; The source end listens to the first audio focus information and the first media playback status information of the first application, and sends the first audio focus information and the first media playback status information of the first application to the target end. The first application can be any one of the at least one application. Based on the first audio focus information and the first media playback status information, the target device displays a first sound identifier in the display interface. The first sound identifier is used to characterize the sound status of the first application.
[0028] This disclosed embodiment utilizes a system comprised of a source and a target end to achieve a collaborative working mode of source-end data acquisition and target-end display, covering device interconnection scenarios. The source end listens for and sends status information, while the target end displays a sound identifier based on the information. The two have clearly defined roles, requiring no additional hardware or complex algorithms, resulting in low system implementation costs. This system relies on both audio focus and media playback status information to ensure accurate recognition. Simultaneously, the visual identifier allows users to quickly distinguish the sound application, solving the problems of high resource consumption, poor real-time performance, and ambiguous recognition inherent in traditional solutions, thereby improving user experience and system operating efficiency in device interconnection scenarios.
[0029] Fourthly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the voice application recognition method of the first aspect or any corresponding embodiment described above, or to perform the voice application recognition method of the second aspect or any corresponding embodiment described above.
[0030] Fifthly, this disclosure provides a computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the voice application recognition method of the first aspect or any corresponding embodiment described above, or to execute the voice application recognition method of the second aspect or any corresponding embodiment described above.
[0031] In a sixth aspect, this disclosure provides a computer program product, including computer instructions, which are used to cause a computer to execute the voice application recognition method of the first aspect or any corresponding embodiment described above, or to execute the voice application recognition method of the second aspect or any corresponding embodiment described above. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of the structure of a voice recognition system according to an embodiment of the present disclosure; Figure 2 This is a schematic flowchart of a voice application recognition method according to an embodiment of the present disclosure; Figure 3 This is a flowchart illustrating another sound application recognition method according to an embodiment of the present disclosure; Figure 4 This is an example diagram illustrating the display of a sound-emitting symbol in a display interface according to an embodiment of the present disclosure; Figure 5 This is another example diagram showing a sound-emitting symbol displayed in a display interface according to an embodiment of the present disclosure; Figure 6 This is a flowchart illustrating another speech application recognition method according to an embodiment of the present disclosure; Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0035] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0036] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0037] With the widespread adoption of Android devices, the number of applications on these devices is also increasing. In device-to-device communication scenarios, Android devices need to compete with each other for speaker resources. Although audio focus is managed by the system since Android 12, forcing the audio playback in the other application to fade out when another application requests audio focus, and also muting audio playback when receiving an incoming call, improper use of audio playback by developers can lead to users being unable to distinguish which application's sound is coming from, sometimes even causing audio mixing and negatively impacting the user experience.
[0038] To address the aforementioned problems, according to one aspect of the present disclosure, a method for identifying voice-generating applications is provided. Optionally, in this embodiment, the above-described voice-generating application identification method can be applied to applications such as... Figure 1 The sound recognition application system shown. For example... Figure 1 As shown, the sound application identification system includes a target terminal 101 and a source terminal 102. The target terminal 101 is interconnected with the source terminal 102, and at least one application of the source terminal 102 is displayed on the display interface of the target terminal 101. The source terminal 102 listens to the first audio focus information and the first media playback status information of the first application, and sends the first audio focus information and the first media playback status information of the first application to the target terminal 101. The first application can be any one of the at least one application. Based on the first audio focus information and the first media playback status information, the target terminal 101 displays a first sound identifier on the display interface. The first sound identifier is used to characterize the sound status of the first application.
[0039] In this example, taking a mobile phone as the source device 102 and a vehicle infotainment system as the target device 101, the interconnection between the target device 101 and the source device 102 can be achieved by projecting the mobile phone screen onto the vehicle infotainment system screen, using the vehicle infotainment system screen as a secondary screen or mirror of the mobile phone. In this case, the vehicle infotainment system screen acts as an extended display area for the mobile phone, and the operation of the projected application displayed on the vehicle infotainment system screen is still completed through the mobile phone. Alternatively, an application on the mobile phone can be "pinned" or "tethered" to the desktop of the vehicle infotainment system (or other large-screen device) using a PIN application method, allowing it to run independently like a native application, and the user can directly operate the application on the vehicle infotainment system screen. It should be noted that this embodiment uses a vehicle infotainment system as the target device and a mobile phone as the source device. In practical applications, the target device 101 and the source device 102 can also be other interconnected devices, such as a television as the target device 101 and a tablet computer as the source device 102. This embodiment is not limited to this.
[0040] Specifically, the further working process of the target end 101 and the source end 102 is detailed in the relevant description of the method embodiment below, and will not be repeated here.
[0041] This disclosed embodiment utilizes a system comprised of a source and a target end to achieve a collaborative working mode of source-end data acquisition and target-end display, covering device interconnection scenarios. The source end listens for and sends status information, while the target end displays a sound identifier based on the information. The two have clearly defined roles, requiring no additional hardware or complex algorithms, resulting in low system implementation costs. This system relies on both audio focus and media playback status information to ensure accurate recognition. Simultaneously, the visual identifier allows users to quickly distinguish the sound application, solving the problems of high resource consumption, poor real-time performance, and ambiguous recognition inherent in traditional solutions, thereby improving user experience and system operating efficiency in device interconnection scenarios.
[0042] According to an embodiment of this disclosure, a method for identifying voice applications is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0043] This embodiment provides a method for recognizing voice-generating applications, which can be used for, for example Figure 1 The sound recognition system shown is an example of a sound emission application. Figure 2 As shown, Figure 2 This is a flowchart of a voice application recognition method according to an embodiment of the present disclosure, wherein the target end is used to perform steps S201 and S202, and the source end is used to perform steps S301 and S302.
[0044] In step S201, in response to establishing an interconnection with the source, at least one application from the source is displayed in the display interface.
[0045] This disclosure uses mobile phone-vehicle interconnection as an example for illustration. The source device is a mobile phone, and the target device is the vehicle's infotainment system. The vehicle's infotainment system acts as a secondary screen for the mobile phone, and mobile phone applications are fixedly displayed on the vehicle's desktop using a PIN, enabling application projection and independent operation. Once the mobile phone and vehicle's infotainment system establish a communication connection via wireless or wired means, the vehicle's infotainment system responds to the interconnection establishment event, loading and displaying one or more application interfaces from the mobile phone on the vehicle's display interface. These applications are presented as secondary screens or PIN-enabled applications, displayed and running independently like native applications on the vehicle's infotainment system. Users can directly view the application list and running status of the source mobile phone on the vehicle's infotainment system interface.
[0046] For example, after the mobile phone and the vehicle's infotainment system are connected, the vehicle's screen displays the icons and interfaces of the mobile phone's music app, video app, navigation app, and other applications, which the user can view and operate directly on the vehicle's infotainment system.
[0047] Step S301: In response to establishing an interconnection with the target terminal, detect the first audio focus information and the first media playback status information of the first application.
[0048] The first application is the source application, and the first application is displayed in the target application's display interface.
[0049] Specifically, after the mobile phone establishes a connection with the vehicle's infotainment system, it immediately performs system-level monitoring on any application displayed on the system (i.e., the first application), obtaining the application's audio focus information and media playback status information in real time. The audio focus information indicates whether the application has obtained system audio output permissions, while the media playback status information indicates whether the application is currently playing, paused, or stopped.
[0050] For example, if the mobile phone detects that the music app is currently receiving system audio focus and the media is playing, it will obtain the corresponding audio focus information and media playback status information.
[0051] Step S302: Send the first audio focus information and the first media playback status information of the first application to the target terminal.
[0052] Specifically, the source phone will send the audio focus information and media playback status information of the first application it listens to to the target vehicle system in real time through the established interconnection channel, so that the vehicle system can obtain real and accurate sound-related status data of the source application without the vehicle system collecting or analyzing audio data itself.
[0053] For example, the mobile phone sends the information that "the music app has audio focus and is in playback mode" to the vehicle's infotainment system via the interconnection link.
[0054] This embodiment detects the application's audio focus and media playback status after the source and target ends are interconnected and sends them to the target end. The source end is responsible for status acquisition, eliminating the need for audio detection and analysis on the target end, thus reducing the computational burden on the target end. The source end directly acquires system-level audio focus and media session information, ensuring high accuracy and avoiding errors and resource consumption in audio sampling judgment. Simultaneously, it proactively sends status information, ensuring timely information transmission and supporting the target end to quickly display the sound emission indicator. This achieves synchronization of sound emission status between the source and target ends, resolving the issue of unrecognizable sound applications in device interconnection scenarios and improving the overall efficiency and accuracy of the solution.
[0055] In step S202, in response to receiving the first audio focus information and the first media playback status information of the first application sent by the source end, the first sound identifier is displayed in the display interface based on the first audio focus information and the first media playback status information.
[0056] Wherein, the first application is any one of at least one application, and the first sound identifier is used to characterize the sound state of the first application.
[0057] Specifically, after receiving the audio focus information and media playback status information sent by the mobile phone, the vehicle's infotainment system determines whether the application is in a sound-producing state based on the combination of this information, and displays the first sound-producing icon in the corresponding application's display interface. This icon is used to intuitively indicate whether the current application is producing sound.
[0058] For example, after receiving the message "audio focus is on + playing" from a music app, the vehicle's infotainment system displays a sound indicator at a preset location in the app window (such as the top of a flexible desktop); if the music app is paused or loses audio focus, the infotainment system displays a no-sound indicator.
[0059] This embodiment of the disclosure establishes an interconnection between the target end and the source end, displays the source end application, and displays a first sound identifier on the display interface based on the first audio focus information and the first media playback status information sent by the source end. This allows for a clear representation of the sound status of the first application on the source end directly on the target end interface. Compared to traditional methods that require users to make their own judgments or additionally detect audio data, this solution eliminates the need for the target end to sample and analyze the audio stream, avoiding the consumption of target end processor resources. Furthermore, relying on the dual information of audio focus and media playback status ensures the accuracy of sound status recognition, enabling users to quickly and intuitively distinguish sound-emitting applications in device interconnection scenarios, thus improving the interconnection user experience and intuitiveness of interaction.
[0060] This embodiment provides a method for recognizing voice-generating applications, which can be used for, for example Figure 1 The sound recognition system shown is an example of a sound emission application. Figure 3 As shown, Figure 3 This is a flowchart of a voice application recognition method according to an embodiment of the present disclosure, wherein the target end is used to perform steps S401 and S402, and the source end is used to perform steps S501 and S502.
[0061] Step S401: In response to establishing an interconnection with the source, at least one application from the source is displayed in the display interface. See below for details. Figure 2 The relevant descriptions of step S201 shown will not be repeated here.
[0062] Step S501: In response to establishing an interconnection with the target end, the first audio focus information and the first media playback status information of the first application are detected. The first application is the application of the source end, and the first application is displayed in the display interface of the target end. See details below. Figure 2 The relevant descriptions of step S301 shown will not be repeated here.
[0063] Step S502: Send the first audio focus information and the first media playback status information of the first application to the target terminal. See details below. Figure 2 The relevant description of step S302 shown will not be repeated here.
[0064] Step S402: In response to receiving the first audio focus information and the first media playback status information of the first application sent by the receiving source, a first sound identifier is displayed in the display interface based on the first audio focus information and the first media playback status information. The first application is any one of at least one application, and the first sound identifier is used to characterize the sound status of the first application.
[0065] Specifically, step S402 includes: In step S4021, in response to the first audio focus information indicating that the first application has focus and the first media playback status information indicating that the first application is in playback status, the first form of the first sound identifier is displayed in the display interface.
[0066] The first form is used to characterize the sound produced by the first application.
[0067] Specifically, when the target terminal determines that the first audio focus information indicates that the first application currently has audio focus, and the first media playback status information indicates that the first application is currently in playback state, the target terminal displays the first form of the first sound identifier in the display interface of the corresponding application to clearly indicate that the application is currently making a sound.
[0068] Furthermore, the first form of displaying the first sound identifier in the display interface includes one or more combinations of the following steps a1, a2, and a3.
[0069] Step a1: Display the first sound identifier in the display interface using an animated display method. The first form includes the animated display method.
[0070] Specifically, the target device renders the first sound indicator in the display interface using dynamic effects, such as wave effects, flashing effects, pulse diffusion effects, and volume ripple effects, to give the indicator a continuously changing visual effect, highlighting that the application is in a sound-producing state. For example, in a car-to-vehicle connectivity scenario, the vehicle's infotainment system displays the sound indicator corresponding to the music app using a volume ripple effect, intuitively prompting the user that the application is producing sound.
[0071] In this embodiment, the first form of the sound indicator adopts a dynamic display method. The dynamic visual effect is more prominent and can quickly attract the user's attention on the target device's display interface, allowing the user to immediately perceive that the application is in a sound state and avoid missing the sound application prompt. Furthermore, the dynamic display enhances the status reminder effect, adapts to the user's habit of quickly browsing the interface in the device interconnection scenario, and can identify the sound application without prolonged staring, improving the effectiveness and recognizability of the status prompt. At the same time, the implementation of the dynamic effect does not require a large amount of additional system resources, balancing the prompt effect and device performance.
[0072] Step a2: Display the first sound identifier at the first preset position on the display interface. The first form includes the first preset position.
[0073] Specifically, the target device displays the first sound identifier at a first preset position on the display interface. This position is an area on the interface that is easy for the user to observe and can form a clear correspondence with the application interface.
[0074] For example, the first preset position is set to the upper right corner of the application window, the flexible desktop menu area, or one side of the application title bar, so that the sound state corresponds one-to-one with the application.
[0075] In this embodiment, the sound identifier is displayed in a first preset position on the display interface. This fixed position helps users develop a consistent visual habit, eliminating the need to search for the sound identifier and allowing for quick location of the application's sound status. The preset position layout conforms to user visual browsing patterns, improving status recognition efficiency. Simultaneously, the fixed position ensures a neat and uniform interface layout, avoiding recognition confusion caused by random changes in identifier position. It adapts to scenarios where multiple applications are displayed simultaneously, accurately corresponding to the sound status of the first application, and enhancing the standardization and convenience of interface interaction.
[0076] Step a3: Display the first sound identifier in the display interface using the first interface display elements. The first form includes the first interface display elements.
[0077] Specifically, the target end uses a first interface display element to present the first sound identifier. This element is a graphical, icon, color block, symbol and other interface components specifically used to represent the sound state, with distinctive visual features.
[0078] For example, the first interface displays elements such as a green speaker icon, a filled volume icon, and a highlighted dot to clearly indicate that the application is making a sound. For instance, when a music application from a mobile phone is projected onto the car's infotainment system, and the music application is playing a song, such as... Figure 4 As shown, a speaker icon representing the sound output is displayed in the upper left corner of the in-vehicle infotainment system to indicate to the user that the music application is currently producing sound.
[0079] The following are examples of several different combinations of the first form of the first sound identifier displayed in the display interface: Example 1: In a scenario where a mobile phone connects to a car's infotainment system, the phone projects a music app onto the car's screen and displays it as a PIN application. The car's infotainment system receives this information, indicating that the music app currently has audio focus and is playing media. It then continuously displays a sound indicator with a volume ripple effect next to the title bar in the upper right corner of the application window (the first preset position). The effect cycles through the playback status, maintaining a fixed and prominent position, allowing users to intuitively and quickly identify that the music app is currently playing audio.
[0080] Example 2: In a scenario where a mobile phone and vehicle infotainment system are interconnected, a video app on the phone is projected onto the vehicle's secondary screen. The vehicle's infotainment system receives this information, indicating that the video app has audio focus and is playing. It then uses a green-filled speaker icon (the first element displayed on the screen) as the primary sound indicator and applies a pulsating flashing effect to the speaker icon. This combination of a clear interface element and dynamic effects enhances the sound status indication, allowing users to quickly distinguish the currently playing application within the vehicle's infotainment system.
[0081] Example 3: In a scenario where a mobile phone and a vehicle's infotainment system are interconnected, the mobile phone and the vehicle's infotainment system establish a device connection. Music apps on the mobile phone are displayed as PIN applications on the vehicle's secondary screen. After the vehicle's infotainment system receives information that the music app has audio focus and is playing, a solid green speaker icon (the first interface display element) is displayed as the first sound indicator in the three-dot menu area on the right side of the top flexible desktop (the first preset position). A continuous volume ripple diffusion effect is then displayed on this speaker icon (animation display method). By combining a fixed position, a dedicated interface element, and dynamic effects, the system clearly, prominently, and intuitively indicates that the music app is currently playing sound, allowing users to quickly and accurately identify the app playing the sound.
[0082] In this embodiment, the sound status is displayed using a first interface element. This dedicated interface element highlights the sound status, providing a clear visual distinction from the silent state and allowing users to quickly identify when the application is speaking. The first interface element can be adapted to the target device's interface style, ensuring overall aesthetics. Simultaneously, the dedicated element is highly recognizable, unaffected by interference from other interface content, and improves the accuracy of the sound status prompt. This setting requires no complex interface modifications, resulting in low implementation costs, while simultaneously enhancing the intuitiveness of the status display and optimizing the user's online experience.
[0083] And / or, in step S4022, in response to the first audio focus information indicating that the first application has no focus, and / or the first media playback status information indicating that the first application is in a non-playback state, the second form of the first sound identifier is displayed in the display interface.
[0084] The second form is used to characterize the first application as not producing sound.
[0085] Specifically, when the target terminal determines that the first audio focus information indicates that the first application has no audio focus, and / or the first media playback status information indicates that the first application is in a non-playback state such as paused, stopped, or muted, the target terminal displays the second form of the first sound identifier in the display interface to clearly indicate that the application is not currently making a sound.
[0086] In this embodiment, the first and second forms of the first sound identifier are displayed based on the audio focus and media playback status, respectively. By distinguishing between the sound-emitting and non-sound-emitting states through a rule that the sound-emitting state corresponds to being in focus and playing, and the non-sound-emitting state corresponds to being out of focus or not playing, a visual distinction is achieved. This setting allows users to quickly determine whether the application is emitting sound based on the identifier form, without additional operation or judgment, simplifying the user identification process. Simultaneously, a standardized status determination logic is adopted to adapt to various application sound-emitting scenarios, ensuring consistency in status display and improving the clarity and ease of use of the target terminal interface.
[0087] Furthermore, the second form of displaying the first sound identifier in the display interface includes one or more combinations of the following steps b1, b2 and b3.
[0088] Step b1: Display the first sound identifier in a static display mode on the display interface. The second form includes the static display mode.
[0089] Specifically, the target device displays the first sound icon in a static, non-animated manner on the display interface. The icon maintains a fixed shape, does not change, and does not flash, thus clearly distinguishing it from the first shape with animation.
[0090] For example, the in-vehicle infotainment system displays the sound output indicator of the music app as a static icon without any animation effects, indicating that the app is not currently producing sound.
[0091] In this embodiment, the second-form identifier is displayed statically when no sound is emitted. The static display style is simple and stable, forming a sharp visual contrast with the animated sound identifier, further enhancing the distinction between the sound and non-sound states. Static display eliminates the need for continuous animation rendering, reducing the resource consumption of the target interface and improving interface smoothness. At the same time, the static identifier is clear and stable, without causing visual interference to the user, adapting to long-term display scenarios, ensuring the simplicity and readability of the non-sound state prompt, and making the interface state display more reasonable and comfortable.
[0092] Step b2: Display the first sound indicator at the second preset position on the display interface. The second form includes the second preset position.
[0093] Specifically, the target device displays the first sound indicator at a second preset position on the display interface. The second preset position is different from the first preset position, and the position difference further distinguishes between the sounding and non-sounding states.
[0094] For example, the second preset position is the lower right corner of the application window, a gray placeholder area, or a semi-hidden corner position, which is different from the first preset position when the sound is emitted.
[0095] In this embodiment, the non-voiced identifier is displayed in a second preset position. By differentiating itself from the voiced identifier by this preset position, the visual difference between the voiced and non-voiced states is further enhanced, avoiding confusion. The dual preset positions make the interface display more organized, allowing users to quickly determine the application's voicing status and simplifying the identification process. Simultaneously, the position differentiation requires no additional interface elements, does not increase interface complexity, ensures a clean display, and adapts to the simultaneous display needs of multiple applications in interconnected device scenarios, improving the accuracy and intuitiveness of the status display.
[0096] Step b3: Display the first sound identifier in the display interface using the second interface display elements. The second form includes the second interface display elements, which are different from the first interface display elements.
[0097] Specifically, the target terminal uses a second interface display element to present the first sound identifier. The second interface display element is different from the first interface display element in terms of shape, color, style, and fill status to avoid user confusion.
[0098] For example, the second interface displays elements such as a gray speaker icon, a diagonally crossed prohibition symbol, a hollow dot, and a semi-transparent icon to clearly indicate that the application is not making a sound. For instance, when a music application on a mobile phone is projected onto the car's infotainment system, and the music application is playing a song, such as... Figure 5 As shown, the user is not notified that the music application is not playing music by displaying a speaker icon in the upper left corner of the in-vehicle infotainment system's display interface.
[0099] In this embodiment, the non-voice state uses a second interface display element that differs from the first interface element. This dual-interface design accurately distinguishes between the voice-emitting and non-voice-emitting states, avoiding confusion. The differentiated design of the two interface elements creates a strong visual contrast, allowing users to quickly determine the application status and improving recognition efficiency. Simultaneously, the differentiated elements are compatible with the target device's interface design specifications, ensuring overall interface consistency. Clear status display is achieved without additional system resources, resolving the issue of ambiguous recognition of voice-emitting applications in interconnected device scenarios and enhancing the interactive experience.
[0100] The following are examples of several different combinations of the first form of the first sound identifier displayed in the display interface: Example 1: In a scenario where a mobile phone and a car's infotainment system are connected, when the car's infotainment system receives a notification that a music app has lost audio focus or is paused, a static, non-animated sound indicator is displayed in the semi-transparent corner area at the bottom right of the application window (the second preset position). The indicator remains fixed and without any dynamic effects, clearly distinguishing it from the animation and fixed position of the sound, clearly indicating that the application is not currently producing sound.
[0101] Example 2: In a scenario where a mobile phone and a car's infotainment system are connected, if a video app loses audio focus or pauses due to an incoming call, the car's infotainment system uses a gray hollow horn icon with a dormant line (displayed on the second screen) as a sound indicator, displayed statically without flickering or diffusion. This element is clearly different from the solid green horn icon when the app is playing sound, intuitively and clearly indicating to the user that the application is not currently producing sound.
[0102] Example 3: In a scenario where a mobile phone and a car infotainment system are connected, a music app may switch from playing to paused, or become silent due to other applications taking over the audio focus. On the car infotainment system, a preset gray mute symbol (a second interface display element) is displayed as a sound indicator in the non-protruding area on the left side of the bottom flexible desktop (the second preset position), and displayed in a completely static manner without any animation. This triple distinction—position, element, and display method—creates a strong visual contrast with the sound state, allowing users to instantly determine that the application is silent, avoiding confusion and misjudgment.
[0103] This embodiment provides a method for recognizing voice-generating applications, which can be used for, for example Figure 1 The sound recognition system shown is an example of a sound emission application. Figure 6 As shown, Figure 6 This is a flowchart of a voice application recognition method according to an embodiment of the present disclosure, wherein the target end is used to perform steps S601 to S603, and the source end is used to perform steps S701 and S703.
[0104] Step S601: In response to establishing an interconnection with the source, at least one application from the source is displayed in the display interface. See below for details. Figure 3 The relevant descriptions of step S401 shown will not be repeated here.
[0105] Step S701: In response to establishing an interconnection with the target end, the first audio focus information and the first media playback status information of the first application are detected. The first application is the application of the source end, and the first application is displayed in the display interface of the target end. See details below. Figure 3 The relevant descriptions of step S501 shown will not be repeated here.
[0106] Step S702: Send the first audio focus information and the first media playback status information of the first application to the target terminal. See details below. Figure 3 The relevant descriptions of step S502 shown will not be repeated here.
[0107] Step S602: In response to receiving the first audio focus information and first media playback status information of the first application sent by the source end, based on the first audio focus information and first media playback status information, a first sound identifier is displayed on the display interface. The first application can be any one of at least one application, and the first sound identifier is used to characterize the sound status of the first application. See details below. Figure 3 The relevant description of step S402 shown will not be repeated here.
[0108] Step S703: In response to detecting an update to the first audio focus information and / or the first media playback status information of the first application, the updated second audio focus information and the second media playback status information of the first application are sent to the target terminal.
[0109] Specifically, the source continuously monitors the first application with established interconnection. When it detects a change in the first audio focus information and / or the first media playback status information of the first application, it immediately acquires the updated second audio focus information and the second media playback status information, and sends the updated second audio focus information and the second media playback status information to the target end in real time through the established device interconnection channel. Status updates include situations such as audio focus changing from focused to unfocused, from unfocused to focused, and media playback status switching between play, pause, and stop.
[0110] For example, in a mobile phone and vehicle interconnection scenario, when a music app on a mobile phone switches from playing to paused, or when the music app loses audio focus due to an incoming call, the mobile phone immediately detects the status update, generates and sends the updated audio focus information and media playback status information to the vehicle's infotainment system.
[0111] In this embodiment, the source device immediately sends the latest information upon detecting a state update, achieving real-time synchronous transmission of the sound state. When the application audio focus or playback state changes, the source device triggers information transmission instantly, ensuring that the information received by the target device is not delayed compared to the actual state of the source device, thus avoiding display errors caused by state lag. This dynamic update mechanism does not require continuous polling and detection; it only transmits data when the state changes, reducing data transmission volume and system resource consumption. At the same time, it ensures accurate and timely state synchronization, adapts to various dynamic sound scenarios, and improves the smoothness and stability of device interconnection state synchronization.
[0112] Step S603: In response to receiving the second audio focus information and the second media playback status information of the first application sent by the source end, update the display of the first sound identifier in the display interface based on the second audio focus information and the second media playback status information.
[0113] Specifically, after receiving the updated second audio focus information and second media playback status information from the source end, the target end re-determines the current sound output status of the application based on the updated information. Based on the latest determination, it updates the displayed first sound output indicator in real time on the display interface, ensuring the indicator's form matches the actual sound output status of the application. This update includes switching the first form to the second form, switching the second form back to the first form, or synchronously updating animations, positions, and interface elements within the same form.
[0114] For example, the car's infotainment system originally displayed a music app playing audio with a dynamic, bright speaker icon. After receiving the updated message "No audio focus, in paused state," it immediately switched the audio indicator to a static, gray mute icon, completing the status update display.
[0115] This embodiment of the disclosure achieves real-time synchronization and dynamic updating of the sound status by adding a step of updating the sound identifier based on the updated audio focus and media playback status information. When the sound status of the source application changes, the target device can immediately update the identifier display, ensuring the real-time and accurate display of the status and avoiding recognition errors caused by status lag. Furthermore, this update mechanism does not require manual triggering by the user; it automatically completes status synchronization, adapting to dynamic scenarios such as application playback, pause, and focus switching, maintaining consistency between the sound status and the actual situation throughout, and improving the smoothness and reliability of device interconnection.
[0116] In a specific application scenario, after the mobile phone and the vehicle's infotainment system establish a connection, the mobile phone application is fixed to the vehicle's desktop using a PIN. The mobile phone monitors the audio focus information (MediaSession) and media playback status information (AudioFocus) of the music application and sends the monitored MediaSession and AudioFocus data to the vehicle's infotainment system.
[0117] Based on the received information, the in-vehicle infotainment system marks the status of the music application's audio output on a window. For example, in a mirrored setup with the phone connected to the car's infotainment system, three dots at the top of the flexible desktop are used. If there is sound, an animated effect is displayed; otherwise, three static dots are shown. Specifically, for AudioFocus, having focus indicates sound; no focus indicates no sound. For MediaSession, a paused state indicates no sound, while a playing state indicates sound. Sound is only confirmed and the three dots are animated only if both AudioFocus and MediaSession indicate sound; otherwise, three static dots are displayed.
[0118] The embodiments disclosed herein can determine whether an application is making a sound at a low cost (no additional CPU resources are required to detect whether it is muted, and the judgment result is relatively accurate).
[0119] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0120] The following is a detailed reference. Figure 7 This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 701, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 702 or a program loaded from memory 703 into random access memory (RAM) 703. RAM 703 also stores various programs and data required for the operation of the electronic device. The processor 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0121] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 703 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 7 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0122] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 709, or installed from a memory 703, or installed from a ROM 702. When the computer program is executed by the processor 701, it performs the functions defined in the voice application recognition method of embodiments of this disclosure.
[0123] Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0124] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor electronic devices, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the voice application identification method shown in the above embodiments is implemented.
[0125] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0126] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for identifying a sound generating application, applied to a target end, comprising: receiving a sound generating application identifier from a source end; and identifying the sound generating application according to the sound generating application identifier. The method includes: In response to establishing an interconnection with the source, at least one application from the source is displayed in the display interface; In response to receiving first audio focus information and first media playback status information of the first application sent by the source end, a first sound identifier is displayed in the display interface based on the first audio focus information and the first media playback status information. The first application is any one of the at least one application, and the first sound identifier is used to characterize the sound status of the first application.
2. The method according to claim 1, characterized in that, The step of displaying the first sound identifier in the display interface based on the first audio focus information and the first media playback status information includes: In response to the first audio focus information indicating that the first application has focus, and the first media playback status information indicating that the first application is in playback state, a first form of the first sound identifier is displayed in the display interface, the first form being used to indicate that the first application is emitting sound; and / or, In response to the first audio focus information indicating that the first application has no focus, and / or the first media playback status information indicating that the first application is in a non-playback state, a second form of the first sound identifier is displayed in the display interface, the second form being used to indicate that the first application is not making a sound.
3. The method according to claim 2, characterized in that, The first form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in the display interface using an animated display method, and the first form includes the animated display method.
4. The method according to claim 3, characterized in that, The second form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in a static display mode in the display interface, and the second form includes a static display mode.
5. The method according to claim 2 or 3, characterized in that, The first form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed at a first preset position on the display interface, and the first form includes the first preset position.
6. The method according to claim 2 or 4, characterized in that, The second form of displaying the first sound identifier in the display interface includes: The first sounding identifier is displayed at a second preset position on the display interface, and the second form includes the second preset position.
7. The method according to claim 2, characterized in that, The first form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in the display interface as a first interface display element, and the first form includes the first interface display element.
8. The method according to claim 7, characterized in that, The second form of displaying the first sound identifier in the display interface includes: The first sound identifier is displayed in the display interface as a second interface display element. The second form includes a second interface display element, which is different from the first interface display element.
9. The method according to claim 1, characterized in that, The method further includes: In response to receiving the second audio focus information and the second media playback status information of the first application sent by the source end, the first sound identifier is updated and displayed in the display interface based on the second audio focus information and the second media playback status information.
10. A method for identifying sound sources, characterized in that, The method includes: In response to establishing an interconnection with the target end, the first audio focus information and the first media playback status information of the first application are detected, wherein the first application is the application of the source end and the first application is displayed in the display interface of the target end; The first audio focus information and the first media playback status information of the first application are sent to the target terminal, so that the target terminal displays a first sound identifier in the display interface based on the first audio focus information and the first media playback status information. The first sound identifier is used to characterize the sound status of the first application.
11. The method according to claim 10, characterized in that, The method further includes: In response to detecting an update to the first audio focus information and / or the first media playback status information of the first application, the updated second audio focus information and the second media playback status information of the first application are sent to the target terminal, so that the target terminal updates the display of the first sound identifier in the display interface based on the second audio focus information and the second media playback status information.
12. A voice recognition system, characterized in that, The system includes: a source end and a target end, wherein, The source end establishes an interconnection with the target end, and displays at least one application of the source end in the display interface of the target end; The source end listens to the first audio focus information and the first media playback status information of the first application, and sends the first audio focus information and the first media playback status information of the first application to the target end. The first application can be any one of the at least one application. Based on the first audio focus information and the first media playback status information, the target device displays a first sound identifier in the display interface. The first sound identifier is used to characterize the sound status of the first application.
13. An electronic device, characterized in that, The electronic device includes a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the sound application recognition method according to any one of claims 1 to 9, or to perform the sound application recognition method according to any one of claims 10 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the voice application recognition method according to any one of claims 1 to 9, or to perform the voice application recognition method according to any one of claims 10 to 11.
15. A computer program product, characterized in that, The method includes computer instructions for causing a computer to perform the voice application recognition method according to any one of claims 1 to 9, or to perform the voice application recognition method according to any one of claims 10 to 11.