Audio recognition method, electronic device, storage medium, and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-08-04
AI Technical Summary
[0005]本申请实施例提供了一种音频识别方法、电子设备、存储介质以及程序产品,能够解决音频识别失败或误识别的问题
Smart Images

Figure CN122511291A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio data processing technology, and in particular to an audio recognition method, electronic device, storage medium, and program product. Background Technology
[0002] Song recognition, as an important application in the field of song audio information retrieval, has now become a standard feature of various music streaming platforms and smart terminals.
[0003] In related technologies, audio acquisition typically relies on a mobile phone microphone. An audio clip is recorded from the environment, its audio fingerprint features are extracted, and then these features are matched against the audio fingerprint features of songs in a database. The song with the highest match is then returned to the user. Under ideal quiet environments and with the sound source relatively close (near field), high recognition accuracy and fast recognition speed can already be achieved.
[0004] However, when users are in noisy public places or when the phone is far from the sound source (far field), the phone's built-in omnidirectional microphone will indiscriminately collect the target sound source and a large amount of environmental noise, resulting in a serious decrease in the signal-to-noise ratio of the collected audio signal. This directly affects the accuracy of subsequent audio fingerprint feature extraction, thus causing audio recognition failure or misidentification. Summary of the Invention
[0005] This application provides an audio recognition method, electronic device, storage medium, and program product, which can solve the problem of audio recognition failure or misrecognition. The technical solution is as follows: On the one hand, an audio recognition method is provided, the method comprising: During the sound acquisition process, a viewfinder window is displayed, in which the acquired image is displayed and a sound source location selection box is displayed; In response to an audio recognition trigger operation, audio recognition is performed on the target audio signal generated by the target sound source in the acquired mixed audio signal, and the recognition result is displayed; the target sound source is the sound source selected by the sound source location selection box.
[0006] In one possible implementation, the audio recognition of the target audio signal generated by the target sound source in the acquired mixed audio signal includes: Determine the direction of the target sound source relative to the mobile terminal; Based on the stated direction, beamforming processing is performed on the mixed audio signal to obtain the target audio signal generated by the target sound source; Audio recognition is performed on the target audio signal.
[0007] In another possible implementation, the audio recognition trigger operation includes: For the two-finger touch zoom operation of the image, the two-finger touch zoom operation includes a touch operation in which two touch points slide in a direction that moves away from each other; In the magnified image, the sound source selected by the sound source location selection box is determined as the target sound source. In another possible implementation, the beamforming process on the mixed audio signal based on the direction includes: The magnification factor of the image is determined, and the gain factor is determined based on the magnification factor, wherein the magnification factor and the gain factor are positively correlated; Beamforming processing of the mixed audio signal is performed based on the stated direction and the stated gain factor. In another possible implementation, the beamforming processing of the mixed audio signal based on the stated direction includes: The magnification factor of the image is determined, and the beam angle in the direction is determined based on the magnification factor, wherein the magnification factor is negatively correlated with the beam angle; Beamforming processing is performed on the mixed audio signal based on the direction and the beam angle.
[0008] In another possible implementation, the method further includes: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. If the audio recognition process fails, the system will redirect to a failure message interface, which includes the entry control for the viewfinder window. The viewfinder display window includes: In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0009] In another possible implementation, the method further includes: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. During the audio recognition process, the entry control of the viewfinder window is displayed on the audio recognition interface; The viewfinder display window includes: In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0010] In another possible implementation, the method further includes: In response to a trigger operation of the audio recognition control, audio recognition processing is performed based on the acquired mixed audio signal, and an audio recognition interface is displayed, the audio recognition interface including a function option control; During the audio recognition process, in response to the triggering operation of the function option control, the entry control of the viewfinder window is displayed; The viewfinder display window includes: In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0011] In another possible implementation, the position of the sound source location selection box in the viewfinder is fixed.
[0012] In another possible implementation, the display sound source location selection box includes: In response to a click operation on the image, a sound source location selection box is displayed at the location corresponding to the click operation.
[0013] On the other hand, an audio recognition device is provided, the device comprising: The display module is configured to display a viewfinder window during sound acquisition, in which the acquired image is displayed and a sound source location selection box is displayed. The recognition module is configured to, in response to an audio recognition trigger operation, perform audio recognition on the target audio signal generated by the target sound source in the acquired mixed audio signal and display the recognition result; the target sound source is the sound source selected by the sound source location selection box.
[0014] In one possible implementation, the identification module is used to: Determine the direction of the target sound source relative to the mobile terminal; Based on the stated direction, beamforming processing is performed on the mixed audio signal to obtain the target audio signal generated by the target sound source; Audio recognition is performed on the target audio signal.
[0015] In another possible implementation, the audio recognition trigger operation includes: For the two-finger touch zoom operation of the image, the two-finger touch zoom operation includes a touch operation in which two touch points slide in a direction that moves away from each other; The recognition module is further configured to: identify the sound source selected by the sound source location selection box as the target sound source in the magnified image.
[0016] In another possible implementation, the identification module is further configured to: The magnification factor of the image is determined, and the gain factor is determined based on the magnification factor, wherein the magnification factor and the gain factor are positively correlated; Beamforming processing is performed on the mixed audio signal based on the direction and the gain factor.
[0017] In another possible implementation, the identification module is further configured to: The magnification factor of the image is determined, and the beam angle in the direction is determined based on the magnification factor, wherein the magnification factor is negatively correlated with the beam angle; Beamforming processing is performed on the mixed audio signal based on the direction and the beam angle.
[0018] In another possible implementation, the display module is used for: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. If the audio recognition process fails, the system will redirect to a failure message interface, which includes the entry control for the viewfinder window. In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0019] In another possible implementation, the display module is further configured to: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. During the audio recognition process, the entry control of the viewfinder window is displayed on the audio recognition interface; In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0020] In another possible implementation, the display module is further configured to: In response to a trigger operation of the audio recognition control, audio recognition processing is performed based on the acquired mixed audio signal, and an audio recognition interface is displayed, the audio recognition interface including a function option control; During the audio recognition process, in response to the triggering operation of the function option control, the entry control of the viewfinder window is displayed; In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0021] In another possible implementation, the position of the sound source location selection box in the viewfinder is fixed.
[0022] In another possible implementation, the display module is used for: In response to a click operation on the image, a sound source location selection box is displayed at the location corresponding to the click operation.
[0023] On the other hand, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the method described in any of the above.
[0024] On the other hand, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described in any of the preceding claims.
[0025] On the other hand, a computer program product is provided, including computer program instructions that, when run on a computer, cause the computer to perform the method described in any of the preceding claims.
[0026] The beneficial effects of the technical solution provided in this application are: when a target sound source is selected by the sound source location selection box, the mixed audio signal can be directionally processed based on the spatial location information of the target sound source, thereby enabling audio recognition of the target audio signal generated by the target sound source in the mixed audio signal, effectively suppressing environmental noise and other interfering sound sources, significantly improving the signal-to-noise ratio of the target audio signal, and enhancing the accuracy of audio recognition in noisy and far-field scenarios. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application; Figure 2 This is a flowchart of the audio recognition method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the audio recognition method provided in an embodiment of this application; Figure 4 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 5 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 6 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 7 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 8 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 9 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 10 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 11 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 12 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 13 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 14 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 15 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 16 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 17 This is a schematic diagram of an audio recognition method provided in another embodiment of this application; Figure 18 This is a schematic diagram of the structure of the audio recognition device provided in the embodiments of this application; Figure 19 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0030] This application provides an audio recognition method that can be implemented by a mobile terminal, such as a mobile phone or tablet computer. The mobile terminal is equipped with an image acquisition device and an audio acquisition device. The image acquisition device can be a camera for capturing images, and the audio acquisition device can be a microphone for capturing audio signals from a sound source.
[0031] From a hardware perspective, the structure of a mobile terminal can be as follows: Figure 1 As shown, it includes a processor 110, a memory 120, and a display unit 130.
[0032] The processor 110 can be a central processing unit (CPU) or a system on chip (SoC), etc. The processor 110 can be used for beamforming processing of audio signals, etc.
[0033] The memory 120 may include various volatile or non-volatile memories, such as solid-state disks (SSDs) and dynamic random access memory (DRAM). The memory 120 can be used to store pre-stored data, intermediate data, and result data, such as images acquired by an image acquisition device.
[0034] The display component 130 can be a standalone screen or a screen integrated with the mobile terminal body. The screen is a touch screen, and the display component 130 is used to display the viewfinder, the sound source position selection box, etc.
[0035] In addition to processors, memory, and display components, a terminal may also include communication components, such as wired network connectors, wireless fidelity (WiFi) modules, Bluetooth modules, cellular communication modules, etc. These communication components can be used to transmit data with other devices, which may be servers or other terminals.
[0036] This application provides an audio recognition method, such as... Figure 2 As shown, in some embodiments, the method includes: S201. During the sound acquisition process, a viewfinder window is displayed, in which the acquired image is displayed and a sound source location selection box is displayed.
[0037] In specific implementation, when the mobile terminal performs audio recognition processing solely based on the mixed audio signal collected by the audio acquisition device (such as a composite audio signal in a real environment where the target sound source audio is simultaneously collected by the microphone and various environmental noises and interference sounds are superimposed), and the audio recognition processing fails, or when the user opens an application on the mobile terminal, the entry control of the viewfinder window is displayed. When the entry control is detected to be triggered, the viewfinder window 320 is displayed in the mobile terminal 310 (e.g., ...). Figure 3 As shown in the image (see diagram), this window displays the image captured by the image acquisition device (i.e., the camera) in real time. Simultaneously, a sound source location selection box 330 is overlaid within the viewfinder, allowing the user to select the sound source location. The user visually specifies the spatial location of the target sound source 340 (such as a speaker or human voice) in the image using the sound source location selection box, thereby converting the visual direction information of the target sound source into spatial coordinates, providing directional guidance for subsequent beamforming processing.
[0038] S202. In response to the audio recognition trigger operation, perform audio recognition on the target audio signal generated by the target sound source in the acquired mixed audio signal, and display the recognition result; the target sound source is the sound source selected by the sound source location selection box.
[0039] In specific implementation, when a user performs an audio recognition trigger operation, such as clicking a control or performing a two-finger touch zoom operation on the image (a touch operation where two touch points slide away from each other), the sound source selected by the sound source location selection box in the zoomed-in image is identified as the target sound source. The direction of the target sound source relative to the mobile terminal is determined, i.e., the pixel coordinates of the sound source location selection box in the image are obtained. Then, based on the camera intrinsic parameter matrix of the mobile terminal (including focal length, principal point coordinates, distortion coefficients, etc.) and the physical dimensions of the image acquisition device, the two-dimensional pixel coordinates are converted into a ray direction vector in three-dimensional space through a perspective projection model. This direction vector represents the direction from the optical center of the camera towards the object space region (the region where the target sound source is located) corresponding to the sound source location selection box. This direction has a defined pitch angle, yaw angle, and distance information relative to the coordinate system of the mobile terminal.
[0040] Then, based on the stated direction, beamforming processing is performed on the mixed audio signal to obtain the target audio signal generated by the target sound source. For example... Figure 4 As shown, the core of beamforming processing is to utilize an array of multiple microphones on a mobile terminal. By accurately calculating the time difference of sound waves arriving at each microphone and performing corresponding delay compensation and weighted summation on the signals of each channel, constructive interference is generated to form the main beam. This enhances the target sound source in a specified direction (i.e., the direction of the object space region relative to the mobile terminal) and within a specified beam angle (for example, the specified beam angle can be set to 60°). At the same time, destructive interference is generated in other directions to suppress environmental noise and interference, thereby obtaining an audio signal with enhanced gain in that direction.
[0041] After obtaining the target audio signal with a significantly improved signal-to-noise ratio in the target direction through beamforming processing, audio recognition processing is performed. This can begin by preprocessing the enhanced target audio signal, including pre-emphasis to compensate for high-frequency components and frame-by-frame windowing (such as Hamming windows) to smooth out truncation effects. Then, a short-time Fourier transform is used to convert the signal from the time domain to the frequency domain, obtaining a time-spectrum containing amplitude and phase information. Based on this, features characterizing the uniqueness of the audio are extracted, such as Mel-frequency cepstral coefficients, which constitute the "fingerprint" of the audio segment. Subsequently, this audio fingerprint feature is matched with a massive number of reference fingerprints in a pre-established song database using high-speed similarity matching (usually employing optimization algorithms such as approximate nearest neighbor search). The matching degree or distance score with each candidate fingerprint in the database is calculated, and one or more candidate results with the highest matching degree are selected. Finally, the most likely song recognition result (such as song name, singer, etc.) with a matching score exceeding a preset threshold is returned and presented to the user, completing the entire song recognition process.
[0042] In this embodiment, the sound source location is intuitively specified through image interaction, providing a precise guiding vector for beamforming and avoiding the problem of traditional sound source localization algorithms easily failing in noisy environments. Beamforming processing can form a gain main lobe in the target direction, significantly improving the output signal-to-noise ratio, thereby ensuring the quality of subsequent audio fingerprint feature extraction. It effectively mitigates interference from reverberation and other issues in far-field environments, improving the recognition success rate of mobile terminals in noisy scenarios.
[0043] In some embodiments, beamforming processing of the mixed audio signal based on the direction includes: The magnification factor of the image is determined, and the gain factor is determined based on the magnification factor, wherein the magnification factor and the gain factor are positively correlated; Beamforming processing is performed on the mixed audio signal based on the direction and the gain factor.
[0044] In practice, when a user zooms in on the image in the viewfinder using gestures (such as two-finger zoom), the magnification factor is first determined, and then a positively correlated audio gain factor is determined based on a preset mapping relationship (such as a linear function, exponential function, or lookup table method). The acoustic principle is that image magnification usually corresponds to an increase in the visual size of the subject (i.e., the target sound source) in the image, indicating that the target sound source (object spatial area) may be farther away or require more precise directional sound pickup, thus requiring higher gain to compensate for the attenuation of sound waves propagating in the air. Subsequently, the beamforming algorithm simultaneously uses the direction and gain factor calculated from the sound source location selection box to perform spatial filtering on the multi-channel audio signals collected by the microphone array. Specifically, when calculating the delay and weighting coefficients of each channel signal, this gain factor is incorporated as a multiplication factor into the weight vector, significantly enhancing the audio signal of the target sound source while further suppressing noise interference in the sidelobe direction. The relationship between the magnification factor and the gain factor is shown in Table 1. Table 1
[0045] Table 1 illustrates the relationship between magnification and gain; the higher the magnification, the greater the gain. When the user zooms out of the image in the viewfinder using gestures, the corresponding gain decreases.
[0046] In this embodiment, by mapping the positive correlation between image magnification and gain, the system can sense the user's intention (to zoom in on the visual focus) and infer that the sound source may attenuate more due to increased distance. This allows for higher gain compensation, avoiding the lag and operational burden of manual adjustment. Beamforming algorithms create a narrower main lobe width and a higher sidelobe suppression ratio in a specified direction, ensuring the clarity of the sound quality of the far-field target sound source. By using lookup tables or function mapping (such as linear or exponential relationships), image operations are directly converted into acoustic parameters. This reduces the user's interaction cost of multiple trials and avoids the problem of inaccurate distance estimation in pure audio detection algorithms at low signal-to-noise ratios. Even in complex environments (such as shooting a distant stage in a noisy square), the system can quickly locate the sound source and output usable audio, significantly improving the success rate and robustness of audio fingerprint feature extraction.
[0047] In some embodiments, beamforming processing of the mixed audio signal based on the direction includes: The magnification factor of the image is determined, and the beam angle in the direction is determined based on the magnification factor, wherein the magnification factor is negatively correlated with the beam angle; Beamforming processing is performed on the mixed audio signal based on the direction and the beam angle.
[0048] In practice, when a user zooms in on the image in the viewfinder using gestures (such as two-finger zoom), the magnification factor is first determined, and a negatively correlated beam angle (i.e., beamwidth) is calculated based on a preset mapping rule (such as an inverse proportional function or a lookup table). The physical principle is that an increase in image magnification usually indicates that the user wants to focus more precisely on a distant target sound source. In this case, a narrower beam angle is needed to improve spatial resolution and directional accuracy, thereby more effectively suppressing interference noise outside the main lobe. The beamforming algorithm simultaneously uses the direction and beam angle determined by the sound source location selection box to perform spatial filtering on the multi-channel audio signals collected by the microphone array. Specifically, when calculating the delay and weighting vector of each channel signal, the beam angle parameter is used to constrain the array's directional response vector, resulting in a narrower main lobe width in the final beam in the specified direction. This allows for precise pickup of the target sound source signal while significantly suppressing noise and reverberation interference in the side lobe direction. The relationship between the magnification factor and the beam angle is shown in Table 2. Table 2
[0049] Table 2 is used to illustrate the relationship between magnification and beam angle. The larger the magnification, the smaller the beam angle.
[0050] In this embodiment, a narrower beam angle significantly improves spatial filtering performance. This is achieved by reducing the main lobe width to decrease the pickup of side lobe interference and reverberation, and by enhancing spatial selectivity to suppress nearby interfering sound sources in the same direction, thereby ensuring the purity of far-field audio fingerprint features. By transforming complex beamforming parameter adjustments into an intuitive visual scaling operation, the challenge of beamwidth adaptation in pure audio localization algorithms at low signal-to-noise ratios is solved. In dynamic scenes, it automatically maintains optimal directivity balance (avoiding excessively narrow beams that could cause sound source loss), ultimately improving the robustness and success rate of audio recognition.
[0051] In some embodiments, after beamforming processing of the mixed audio signal based on the direction and the beam angle, an audio signal with enhanced gain in the direction is obtained. Then, based on the audio signal with enhanced gain in the direction, audio recognition processing is performed, that is, the audio fingerprint features corresponding to the audio signal are matched with the audio fingerprints of songs in a pre-established song database. That is, the matching degree score between the audio fingerprint and the audio fingerprint of each song in the database is calculated (the matching degree score is obtained by calculating cosine similarity or Euclidean distance, etc.). Songs with matching degree scores exceeding a preset threshold (for example, the preset threshold can be set to 0.9) are returned and presented to the user. If the maximum matching degree score of the selected songs is still lower than the preset threshold, the matching degree of the song is determined. The difference between the maximum score and a preset threshold (preset threshold minus the maximum matching score) is calculated. If this difference is less than or equal to a preset difference (an example preset difference can be set to 0.1) and the duration reaches a preset duration (e.g., 0.5 seconds), the image is reduced in size, and the beam angle is enlarged to obtain the target beam angle. An adjustment coefficient can be determined based on this difference (the difference and the adjustment coefficient are positively correlated). The product of the adjustment coefficient and the beam angle is then used as the target beam angle. The adjustment coefficient Z = 1 + C, where C represents the difference. For example, when the difference is 0.08, the adjustment coefficient is 1.08, and the beam angle will be enlarged by 1.08 times to obtain the target beam angle. Based on the target beam angle, beamforming processing is performed on the audio signal acquired by the audio acquisition device, and then the above audio recognition processing process is repeated. Since users typically zoom in on the image to its maximum, the beam angle is at its smallest and the directionality is strongest. Users need to accurately point their mobile devices at the target spatial area (the area where the target sound source is located). However, in actual operation, there are often deviations, which prevent users from accurately pointing their mobile devices at the target spatial area. In this case, the audio signal of the target sound source cannot be completely obtained. Therefore, appropriately increasing the beam angle can avoid deviations caused by user operation, obtain the complete audio signal, and improve the accuracy of audio recognition and processing.
[0052] In this embodiment, when the matching score is close to but does not reach the preset threshold, the beam angle is appropriately widened to the target beam angle by calculating the difference and generating an adjustment coefficient. This avoids missing sound source components that may be missed due to the initial beam angle being too narrow (caused by the user over-magnifying the image). It also gradually expands the acoustic coverage through iterative processing, ensuring that the key fingerprint features of the target sound source are completely captured. This significantly reduces the dependence on the accuracy of user operation (e.g., the phone can still effectively pick up sound even if it is slightly offset by 5°), and is especially suitable for robust recognition in far-field or high-noise environments.
[0053] In some embodiments, the position of the sound source location selection box in the viewfinder is fixed.
[0054] In practice, since the sound source location selection box is fixed in the viewfinder, such as being located in the center of the viewfinder (e.g., ...), ... Figure 5 As shown), the user adjusts the phone's shooting angle to position the target sound source within the sound source location selection box. When the image is subsequently zoomed in, if the target sound source deviates from the position within the sound source location selection box (e.g., ...), the user will be affected. Figure 6 As shown), it is still possible to adjust the phone's shooting angle by fine-tuning (e.g., Figure 7 As shown), place the target sound source in the sound source location selection box (e.g., Figure 8 As shown), the direction of the object space region (i.e., the region where the target sound source is located) in the final sound source location selection box is the direction facing the mobile terminal.
[0055] In this embodiment, the sound source location selection box is fixed at the center of the viewfinder, reducing the complexity of user operation and ensuring the stability and accuracy of the directional parameters required for beamforming. Users only need to adjust the phone angle to place the target sound source within the sound source location selection box to quickly calculate the direction vector (pitch and yaw angles approximately zero degrees) facing the mobile terminal. This direction coincides with the optical axis of the phone's camera, allowing the subsequent beamforming algorithm to directly perform precise delay compensation and weighted summation based on the mobile terminal's coordinate system. This results in the formation of the strongest beam directly in front, avoiding positioning errors that might be introduced by manually dragging the selection box. This significantly improves the efficiency of sound pickup in far-field or noisy environments, ultimately ensuring the quality and success rate of audio fingerprint feature extraction.
[0056] In some embodiments, the display sound source location selection box includes: In response to a click operation on the image, a sound source location selection box is displayed at the location corresponding to the click operation.
[0057] In practice, the sound source location selection box can be non-fixed. The user clicks on the location of the sound source (target sound source) to be identified in the image viewfinder, and the sound source location selection box will appear at the clicked location, allowing the user to select the target sound source (e.g., ...). Figure 9 As shown), when the image is subsequently magnified, the target sound source may deviate from the position of the sound source location selection box (e.g. Figure 10 As shown), the sound source location selection box gradually moves as the image is magnified (e.g., ...). Figure 11 As shown), ensure that the target sound source can still be selected (e.g. Figure 12 As shown in the figure, this avoids the loss of the target sound source or the drift of the selected position after the image is enlarged.
[0058] In this embodiment, the user's intent is transformed into a continuous and accurate spatial orientation through the collaborative interaction of click positioning and zoom tracking. This ensures that the beamforming algorithm always locks onto the correct target in complex multi-sound-source or dynamic scenes. The moment the user clicks on the target sound source, not only are the two-dimensional pixel coordinates converted into three-dimensional spatial direction vectors, but more importantly, a dynamic binding relationship is established between the sound source and the selection box. When the image is magnified, causing the apparent position of the target to change, the selection box will perform real-time position compensation based on the initially calibrated spatial coordinates and camera zoom parameters, keeping it always anchored at the center of the target sound source. This avoids the operational cost of repeatedly adjusting the phone angle, and is especially suitable for noisy environments with multiple sound sources. Users can accurately specify a specific sound source rather than the nearest sound source. Through the automatic tracking mechanism during the zooming process, the problem of selection box drift caused by lens distortion and changes in viewing angle is effectively overcome. This ensures that the beamforming process can still enhance the audio signal in the correct direction after image magnification, thereby maintaining the stability of high signal-to-noise ratio audio acquisition in dynamic usage scenarios and ultimately improving the robustness and recognition success rate of audio fingerprint feature extraction.
[0059] In some embodiments, the method further includes: In response to a trigger operation of the audio recognition control, audio recognition processing is performed based on the acquired mixed audio signal, and an audio recognition interface is displayed; if the audio recognition processing fails, the user is redirected to a failure prompt interface, which includes the entry control of the viewfinder window.
[0060] The viewfinder window display includes: displaying the viewfinder window in response to a trigger operation on the entry control.
[0061] In specific implementation, when the user clicks the audio recognition control 1310 in the interface (such as...), Figure 13 As shown), this triggers the audio recognition operation. At this time, the audio acquisition device is activated to capture mixed audio signals in the environment, and audio recognition processing is performed based on the mixed audio signals. Simultaneously, the audio recognition interface is displayed on the mobile terminal (e.g., ...). Figure 14 As shown), this interface can provide real-time feedback on the recognition process using visual elements such as dynamic waveforms, progress prompts, or progress bars; if audio recognition processing fails, it automatically redirects to a failure prompt interface (e.g., ...). Figure 15 As shown, the interface can display failure messages in the form of text or icons, as well as an entry control 1510 for the viewfinder window. When the user subsequently triggers the entry control, the viewfinder window is displayed.
[0062] This embodiment effectively solves the problem of users lacking effective remedial measures when recognition fails, significantly improving the consistency of user experience and the convenience of operation. It organically combines the originally independent audio acquisition with visual interaction, allowing users to proactively provide beamforming-based directional audio enhancement solutions without manually searching for or guessing the auxiliary function entry point. This significantly improves the success rate of audio recognition in complex acoustic environments such as noisy or far-field conditions.
[0063] In some embodiments, the method further includes: In response to a trigger operation of the audio recognition control, audio recognition processing is performed based on the acquired mixed audio signal, and an audio recognition interface is displayed; during the audio recognition processing, the entry control of the viewfinder window is displayed on the audio recognition interface.
[0064] The viewfinder window display includes: displaying the viewfinder window in response to a trigger operation on the entry control.
[0065] In practice, when a user triggers an audio recognition operation by clicking the audio recognition control on the interface, the audio acquisition device is activated to capture mixed audio signals in the environment. Audio recognition processing is then performed based on these mixed audio signals, and the audio recognition interface is simultaneously displayed on the mobile terminal. This interface can provide real-time feedback on the recognition process using visual elements such as dynamic waveforms, progress prompts, or progress bars. While the audio recognition processing is ongoing, the entry control 1610 of the viewfinder window (e.g., ...) is simultaneously displayed on the audio recognition interface. Figure 16 As shown, this entry control can be presented in the form of text, icons, etc. Users do not need to wait for the recognition result; they can directly trigger the entry control to display the viewfinder window.
[0066] This embodiment significantly improves recognition efficiency and user operation flexibility by simultaneously displaying the viewfinder entry control during audio recognition processing. Furthermore, it fully considers the uncertainties of pure audio recognition in complex acoustic environments, enabling users to proactively decide whether to activate visual assistance functions in advance based on real-time environmental conditions. This allows the mobile terminal to actively intervene and guide the target sound source through beamforming technology, quickly obtaining high signal-to-noise ratio audio signals in noisy or far-field scenarios, significantly shortening the overall recognition time and improving the success rate.
[0067] In some embodiments, the method further includes: In response to a trigger operation of the audio recognition control, audio recognition processing is performed based on the acquired mixed audio signal, and an audio recognition interface is displayed, the audio recognition interface including a function option control; during the audio recognition processing, in response to a trigger operation of the function option control, the entry control of the viewfinder window is displayed.
[0068] The viewfinder window display includes: displaying the viewfinder window in response to a trigger operation on the entry control.
[0069] In practice, when a user triggers an audio recognition operation by clicking the audio recognition control on the interface, the audio acquisition device is activated to capture mixed audio signals in the environment. Audio recognition processing is then performed based on these mixed audio signals, and the audio recognition interface is displayed on the mobile terminal. This interface can provide real-time feedback on the recognition process using visual elements such as dynamic waveforms, progress prompts, or progress bars, and also includes at least one function option control 1710 (such as...). Figure 17 As shown); during the ongoing audio recognition process, if the user triggers this function option control, the entry control 1720 of the viewfinder window will be displayed on the audio recognition interface in the form of a pop-up menu, drop-down list, or add button; when the user's trigger operation on the entry control is detected, the viewfinder window will be displayed.
[0070] This embodiment achieves a balance between simplifying the audio recognition interface and enhancing its functionality by incorporating the viewfinder's entry controls into a hierarchical menu beneath the function option controls. Simultaneously, users can control the mobile terminal to directionally enhance the target sound source using beamforming technology by clicking the entry controls, quickly obtaining high signal-to-noise ratio audio signals in noisy or far-field scenarios, significantly reducing overall recognition time and improving the success rate.
[0071] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0072] Based on the same inventive concept, and corresponding to the audio recognition method provided in the embodiments of this application, this application also provides an audio recognition device.
[0073] refer to Figure 18 The audio recognition device includes: Display module 1801 is configured to display a viewfinder window during sound acquisition, wherein the acquired image is displayed and a sound source location selection box is displayed in the viewfinder window; The recognition module 1802 is configured to, in response to an audio recognition trigger operation, perform audio recognition on the target audio signal generated by the target sound source in the acquired mixed audio signal and display the recognition result; the target sound source is the sound source selected by the sound source location selection box.
[0074] In one possible implementation, the identification module 1802 is used for: Determine the direction of the target sound source relative to the mobile terminal; Based on the stated direction, beamforming processing is performed on the mixed audio signal to obtain the target audio signal generated by the target sound source; Audio recognition is performed on the target audio signal.
[0075] In another possible implementation, the audio recognition trigger operation includes: For the two-finger touch zoom operation of the image, the two-finger touch zoom operation includes a touch operation in which two touch points slide in a direction that moves away from each other; The recognition module 1802 is further configured to: determine the sound source selected by the sound source location selection box as the target sound source in the magnified image.
[0076] In another possible implementation, the identification module 1802 is further configured to: The magnification factor of the image is determined, and the gain factor is determined based on the magnification factor, wherein the magnification factor and the gain factor are positively correlated; Beamforming processing is performed on the mixed audio signal based on the direction and the gain factor.
[0077] In another possible implementation, the identification module 1802 is further configured to: The magnification factor of the image is determined, and the beam angle in the direction is determined based on the magnification factor, wherein the magnification factor is negatively correlated with the beam angle; Beamforming processing is performed on the mixed audio signal based on the direction and the beam angle.
[0078] In another possible implementation, the display module 1801 is used for: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. If the audio recognition process fails, the system will redirect to a failure message interface, which includes the entry control for the viewfinder window. In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0079] In another possible implementation, the display module 1801 is further configured to: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. During the audio recognition process, the entry control of the viewfinder window is displayed on the audio recognition interface; In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0080] In another possible implementation, the display module 1801 is further configured to: In response to a trigger operation of the audio recognition control, audio recognition processing is performed based on the acquired mixed audio signal, and an audio recognition interface is displayed, the audio recognition interface including a function option control; During the audio recognition process, in response to the triggering operation of the function option control, the entry control of the viewfinder window is displayed; In response to a trigger operation on the entry control, a viewfinder window is displayed.
[0081] In another possible implementation, the position of the sound source location selection box in the viewfinder is fixed.
[0082] In another possible implementation, the display module 1801 is used for: In response to a click operation on the image, a sound source location selection box is displayed at the location corresponding to the click operation.
[0083] It should be noted that the audio recognition device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio recognition device and the audio recognition method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0084] Based on the same inventive concept, corresponding to the audio recognition method provided in the embodiments of this application, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the audio recognition method described in the above embodiments.
[0085] Figure 19A structural block diagram of an electronic device 1900 provided in an exemplary embodiment of this application is shown. The electronic device 1900 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The electronic device 1900 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0086] Typically, electronic device 1900 includes a processor 1901 and a memory 1902.
[0087] Processor 1901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1901 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0088] The memory 1902 may include one or more computer-readable storage media, which may be non-transitory. The memory 1902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1902 is used to store at least one instruction, which is executed by the processor 1901 to implement the audio recognition method provided in the method embodiments of this application.
[0089] In some embodiments, the electronic device 1900 may optionally include a peripheral device interface 1903 and at least one peripheral device. The processor 1901, memory 1902, and peripheral device interface 1903 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1903 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1904, a display screen 1905, a camera assembly 1906, an audio circuit 1907, a positioning assembly 1908, and a power supply 1909.
[0090] Peripheral interface 1903 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1901 and memory 1902. In some embodiments, processor 1901, memory 1902 and peripheral interface 1903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1901, memory 1902 and peripheral interface 1903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0091] The radio frequency (RF) circuit 1904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1904 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1904 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1904 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0092] Display screen 1905 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1901 for processing. In this case, display screen 1905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1905, disposed on the front panel of electronic device 1900; in other embodiments, there may be at least two display screens, disposed on different surfaces of electronic device 1900 or in a folded design; in still other embodiments, display screen 1905 may be a flexible display screen, disposed on a curved or folded surface of electronic device 1900. Furthermore, display screen 1905 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1905 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0093] The camera assembly 1906 is used to acquire images or videos. Optionally, the camera assembly 1906 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1906 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0094] The audio circuit 1907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 1901 for processing, or to the radio frequency circuit 1904 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the electronic device 1900. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1901 or the radio frequency circuit 1904 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1907 may also include a headphone jack.
[0095] Positioning component 1908 is used to locate the current geographic location of electronic device 1900 for navigation or LBS (Location Based Service). Positioning component 1908 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.
[0096] Power supply 1909 is used to supply power to various components in electronic device 1900. Power supply 1909 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1909 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0097] Those skilled in the art will understand that Figure 19 The structure shown does not constitute a limitation on the electronic device 1900, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0098] The electronic devices described above are used to implement the corresponding audio recognition methods in the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0099] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to perform the audio recognition method described above. This computer-readable storage medium can be non-transitory. For example, the computer-readable storage medium can be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices, etc.
[0100] In an exemplary embodiment, a computer program product is also provided, including computer program instructions that, when executed on a computer, cause the computer to perform the audio recognition method described above.
[0101] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0102] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0103] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0104] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An audio recognition method, characterized in that, The method is applied to a mobile terminal; the mobile terminal; the method includes: During the sound acquisition process, a viewfinder window is displayed, in which the acquired image is displayed and a sound source location selection box is displayed; In response to an audio recognition trigger operation, audio recognition is performed on the target audio signal generated by the target sound source in the acquired mixed audio signal, and the recognition result is displayed; the target sound source is the sound source selected by the sound source location selection box.
2. The audio recognition method according to claim 1, characterized in that, The audio recognition of the target audio signal generated by the target sound source in the acquired mixed audio signal includes: Determine the direction of the target sound source relative to the mobile terminal; Based on the stated direction, beamforming processing is performed on the mixed audio signal to obtain the target audio signal generated by the target sound source; Audio recognition is performed on the target audio signal.
3. The audio recognition method according to claim 2, characterized in that, The audio recognition trigger operation includes: For the two-finger touch zoom operation of the image, the two-finger touch zoom operation includes a touch operation in which two touch points slide in a direction that moves away from each other; In the magnified image, the sound source selected by the sound source location selection box is determined as the target sound source.
4. The audio recognition method according to claim 3, characterized in that, The beamforming process applied to the mixed audio signal based on the stated direction includes: The magnification factor of the image is determined, and the gain factor is determined based on the magnification factor, wherein the magnification factor and the gain factor are positively correlated; Beamforming processing is performed on the mixed audio signal based on the direction and the gain factor.
5. The audio recognition method according to claim 3, characterized in that, The beamforming process applied to the mixed audio signal based on the stated direction includes: The magnification factor of the image is determined, and the beam angle in the direction is determined based on the magnification factor, wherein the magnification factor is negatively correlated with the beam angle; Beamforming processing is performed on the mixed audio signal based on the direction and the beam angle.
6. The audio recognition method according to claim 1, characterized in that, The method further includes: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. If the audio recognition process fails, the system will redirect to a failure message interface, which includes the entry control for the viewfinder window. The viewfinder display window includes: In response to a trigger operation on the entry control, a viewfinder window is displayed.
7. The audio recognition method according to claim 1, characterized in that, The method further includes: In response to the trigger operation of the audio recognition control, the system performs audio recognition processing based on the acquired mixed audio signal and displays the audio recognition interface. During the audio recognition process, the entry control of the viewfinder window is displayed on the audio recognition interface; The viewfinder display window includes: In response to a trigger operation on the entry control, a viewfinder window is displayed.
8. The audio recognition method according to claim 1, characterized in that, The method further includes: In response to a trigger operation of the audio recognition control, audio recognition processing is performed based on the acquired mixed audio signal, and an audio recognition interface is displayed, the audio recognition interface including a function option control; During the audio recognition process, in response to the triggering operation of the function option control, the entry control of the viewfinder window is displayed; The viewfinder display window includes: In response to a trigger operation on the entry control, a viewfinder window is displayed.
9. The audio recognition method according to claim 1, characterized in that, The position of the sound source location selection box in the viewfinder is fixed.
10. The audio recognition method according to claim 1, characterized in that, The sound source location selection box includes: In response to a click operation on the image, a sound source location selection box is displayed at the location corresponding to the click operation.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 10.
12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method described in any one of claims 1 to 10.
13. A computer program product comprising computer program instructions, characterized in that, When the computer program instructions are executed on a computer, the computer causes the computer to perform the method as described in any one of claims 1 to 10.