Audio processing method, audio processing device, electronic equipment, computer readable storage medium and computer program product

By identifying volume jump points in audio signals, splitting sub-audio signals and generating new audio signals, the problems of single sound source type and low efficiency in audio production in existing technologies are solved, and audio generation and richness improvement across sound source types are achieved.

CN120690168APending Publication Date: 2025-09-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410340885.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the existing technology, the sound source type of the recorded materials in the audio production process is the same, making it difficult to generate other music, and the accuracy of the sound effect parameters collected by the auxiliary equipment is not high, which affects the user experience and efficiency.

Method used

By collecting the volume jump points in the audio signal, dividing multiple sub-audio signals, and generating new audio signals based on the pitch heights of these sub-audio signals, audio generation across different sound source types can be achieved.

Benefits of technology

It improves the granularity and efficiency of audio generation, and is able to generate new audio signals that are different from the original audio signals in terms of sound source types, thus enriching the generated richness of audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690168A_ABST
    Figure CN120690168A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and device, electronic equipment, a computer program product and a computer readable storage medium. The method comprises the steps that an audio processing interface is displayed, and the audio processing interface comprises a first recording control; in response to a trigger operation for the first recording control, collecting a first audio signal; and outputting a second audio signal in response to the first audio signal comprising a plurality of volume jump points, the second audio signal being generated based on tone heights respectively corresponding to a plurality of first sub-audio signals, and the plurality of first sub-audio signals being segmented from the first audio signal based on the plurality of volume jump points. According to the invention, the audio production efficiency and richness can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular to an audio processing method, an audio processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Audio types include music, sound effects, and voice. In related technologies, audio production can be performed by recording real-world musical instruments or by collecting sound effect parameters using auxiliary devices such as vibration sensors and cameras connected to the terminal device. However, these recordings can only generate audio of the same source type as the recorded material, making it difficult to generate other types of music. Furthermore, the accuracy of sound effect parameters collected using auxiliary devices is low, impacting the user's audio production experience.

[0003] In the related technology, there is currently no better way to improve the efficiency and richness of audio production. Summary of the Invention

[0004] Embodiments of the present application provide an audio processing method, an audio processing device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the efficiency and richness of audio production.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] This embodiment of the present application provides an audio processing method, the method comprising:

[0007] Displaying an audio processing interface, wherein the audio processing interface includes a first recording control;

[0008] In response to a trigger operation on the first recording control, collecting a first audio signal;

[0009] In response to the first audio signal including multiple volume jump points, a second audio signal is output, wherein the second audio signal is generated based on the pitch heights corresponding to multiple first sub-audio signals, and the multiple first sub-audio signals are separated from the first audio signal based on the multiple volume jump points.

[0010] This embodiment of the present application provides an audio processing method, the method comprising:

[0011] displaying an audio processing interface, wherein the audio processing interface includes a list of audio signals, including at least one pre-recorded audio signal;

[0012] In response to a selection operation on the audio signal list, the selected audio signal is used as the first audio signal.

[0013] In response to the first audio signal including multiple volume jump points, a second audio signal is output, wherein the second audio signal is generated based on the pitch heights corresponding to multiple first sub-audio signals, and the multiple first sub-audio signals are separated from the first audio signal based on the multiple volume jump points.

[0014] This embodiment of the present application provides an audio processing method, the method comprising:

[0015] Acquire a first audio signal;

[0016] performing volume recognition processing on the first audio signal to obtain a plurality of volume jump points;

[0017] dividing the first audio signal into a plurality of first sub-audio signals based on the plurality of volume jump points;

[0018] Obtaining pitch heights corresponding to the plurality of first sub-audio signals respectively;

[0019] A second audio signal is generated based on each of the pitches.

[0020] An embodiment of the present application provides an audio processing device, comprising:

[0021] A display module, configured to display an audio processing interface, wherein the audio processing interface includes a first recording control;

[0022] an acquisition module, configured to acquire a first audio signal in response to a trigger operation on the first recording control;

[0023] an output module, configured to output a second audio signal in response to the first audio signal including a plurality of volume jump points, wherein the second audio signal is generated based on pitches respectively corresponding to a plurality of first sub-audio signals, and the plurality of first sub-audio signals are segmented from the first audio signal based on the plurality of volume jump points.

[0024] An embodiment of the present application provides an audio processing device, comprising:

[0025] a display module, configured to display an audio processing interface, wherein the audio processing interface includes a list of audio signals, including at least one pre-recorded audio signal;

[0026] The collecting module is configured to, in response to a selection operation on the audio signal list, use the selected audio signal as the first audio signal.

[0027] an output module, configured to output a second audio signal in response to the first audio signal including a plurality of volume jump points, wherein the second audio signal is generated based on pitches respectively corresponding to a plurality of first sub-audio signals, and the plurality of first sub-audio signals are segmented from the first audio signal based on the plurality of volume jump points.

[0028] An embodiment of the present application provides an audio processing device, comprising:

[0029] An acquisition module, configured to acquire a first audio signal;

[0030] an acquisition module, configured to perform volume recognition processing on the first audio signal to obtain a plurality of volume jump points;

[0031] an acquisition module, configured to separate the first audio signal into a plurality of first sub-audio signals based on the plurality of volume jump points;

[0032] an output module, configured to obtain pitch heights corresponding to the plurality of first sub-audio signals;

[0033] An output module is configured to generate a second audio signal based on each of the pitches.

[0034] An embodiment of the present application provides an electronic device, comprising:

[0035] a memory for storing computer-executable instructions or computer programs;

[0036] The processor is configured to implement the audio processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer program stored in the memory.

[0037] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program, which is used to implement the audio processing method provided in the embodiment of the present application when executed by a processor.

[0038] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the audio processing method provided in the embodiment of the present application is implemented.

[0039] The embodiments of the present application have the following beneficial effects:

[0040] Audio processing is performed based on the sub-audio signal with jumps in the collected first audio signal. Compared with the solution of performing full processing based on the collected signal, the second audio signal is generated based on the sub-audio signal with sudden volume changes in the first audio signal in a targeted manner, and the audio generation processing is more granular. After the first audio signal is collected, the second audio signal generated based on the first sub-audio signal can be played. Other audio signals can be generated based on the first audio signal in real time, thereby improving the efficiency of audio signal generation. The second audio signal is generated based on the pitch height corresponding to the sub-audio signal obtained by dividing the volume jump point in the first audio signal. Compared with the solution in the related art that can only output the audio signal edited based on the first audio signal, the first audio signal and the second audio signal are related by pitch height, and can cross the sound source type, thereby improving the richness of the generated audio signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Schematic diagram of an application mode of the audio processing method provided in an embodiment of the present application;

[0042] Figure 2A This is a schematic diagram of the structure of the terminal device provided in an embodiment of the present application;

[0043] Figure 2B This is a schematic diagram of the structure of the server provided in the embodiment of the present application;

[0044] Figure 3A This is a schematic diagram of a first flow chart of the audio processing method provided in an embodiment of the present application;

[0045] Figure 3B This is a second flow chart of the audio processing method provided in an embodiment of the present application;

[0046] Figure 3C 3 is a schematic diagram of a third flow chart of the audio processing method provided in an embodiment of the present application;

[0047] Figure 3D 4 is a schematic diagram of a fourth flow chart of the audio processing method provided in an embodiment of the present application;

[0048] Figure 3E 5 is a schematic diagram of a fifth flow chart of the audio processing method provided in an embodiment of the present application;

[0049] Figure 3F 6 is a schematic diagram of a sixth flow chart of the audio processing method provided in an embodiment of the present application;

[0050] Figure 4A is a schematic diagram of an audio signal provided in an embodiment of the present application;

[0051] Figure 4B Schematic diagram of the training principle of the neural network model provided in the embodiment of the present application;

[0052] Figure 4C It is a structural diagram of the neural network model provided in the embodiment of the present application;

[0053] Figure 4D This is a first principle schematic diagram of synthesizing a second audio signal provided by an embodiment of the present application;

[0054] Figure 4E is a second principle schematic diagram of synthesizing a second audio signal provided by an embodiment of the present application;

[0055] Figure 4F 3 is a schematic diagram of a third principle of synthesizing a second audio signal provided in an embodiment of the present application;

[0056] Figure 4G is a first schematic diagram of a volume curve provided in an embodiment of the present application;

[0057] Figure 5A This is a schematic diagram of the first interface of the audio processing method provided in an embodiment of the present application;

[0058] Figure 5B This is a schematic diagram of the second interface of the audio processing method provided in an embodiment of the present application;

[0059] Figure 5C This is a schematic diagram of the third interface of the audio processing method provided in an embodiment of the present application;

[0060] Figure 5D This is a schematic diagram of the fourth interface of the audio processing method provided in an embodiment of the present application;

[0061] Figure 5E This is a schematic diagram of the fifth interface of the audio processing method provided in an embodiment of the present application;

[0062] Figure 6A This is a seventh flow chart of the audio processing method provided in an embodiment of the present application;

[0063] Figure 6B This is an eighth flow chart of the audio processing method provided in an embodiment of the present application;

[0064] Figure 7 This is a model structure diagram of the audio processing method provided in the embodiment of the present application;

[0065] Figure 8 This is a schematic diagram of the sixth interface of the audio processing method provided in an embodiment of the present application;

[0066] Figure 9 This is a second schematic diagram of a volume curve provided in an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0068] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0069] If similar descriptions of "first / second" appear in the application documents, the following explanation is added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0070] It should be pointed out that the collection and processing of relevant data (e.g., audio recording data) in this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in practice, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0071] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0073] In the embodiments of the present application, "at least one" refers to one or more situations, and "multiple" is equivalent to "at least two", which refers to two or more situations.

[0074] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0075] 1) Sound effects: These are the effects created by processing sound signals during audio production and post-processing to achieve a specific expression or enhance the audio content. Sound effects typically include natural sounds, synthetic sounds, and those produced through digital signal processing. Common sound effects include explosions, percussion, rain, wind, and animal calls. Sound effects also play a crucial role in film, television, and gaming, adding vividness and emotional depth to the visuals.

[0076] 2) Volume jump point: a sampling point in the time domain of an audio signal. The volume of the volume jump point is greater than the volume threshold, and the absolute value of the difference between the volume of the volume jump point and other adjacent sampling points is greater than the difference threshold. The time span between the volume jump point and other adjacent sampling points is determined by the sampling frequency of the recording device. The time span is specifically the ratio between the sampling unit time and the sampling frequency. For example: for a recording device with a sampling rate of 8K, the sampling unit time is 1 second. 1 second divided by 8000 gives a duration of 0.125 milliseconds (ms) corresponding to the signal obtained for each sampling. The time span between the moments corresponding to each sampling point is 0.125 milliseconds, and the time domain span between the volume jump point and other adjacent sampling points is 0.125 milliseconds. Volume jump points are, for example: the starting point of a volume increase or the end point of a volume decrease.

[0077] 3) Scale step: This is a term in music theory used to describe each independent pitch in the musical system. In the musical system, all notes can be called scale steps. Scale steps can be divided into two categories: basic scale steps and altered scale steps. Basic scale steps refer to the seven independent pitches with fixed names in the musical system. These pitches are named C, D, E, F, G, A, and B. They usually correspond to the white keys on the piano keyboard. Alternated scale steps are different pitches obtained by raising or lowering the basic scale steps. The naming of these scale steps may involve different notation systems, such as 1, 2, 3, 4, 5, 6, and 7 in simplified notation, or do, re, mi, sol, la, and xi in solfège.

[0078] 4) Convolutional Neural Networks (CNNs) are a type of feedforward neural network (FNN) with a deep structure that incorporates convolutional computations. They are a representative algorithm for deep learning. CNNs possess representation learning capabilities and can perform shift-invariant classification on input images based on their hierarchical structure.

[0079] 5) Musical elements. Basic musical elements refer to the various elements that make up music, including pitch, duration, strength, and timbre. These basic elements combine to form the commonly used "formal elements" of music, such as rhythm, melody, harmony, dynamics, tempo, mode, form, texture, mood, style, rhythm, and notes. The musical elements in the embodiments of this application are specifically formal elements.

[0080] 6) Musical note, short for musical notation, is used to record the progression of notes of varying lengths. Whole notes, half notes, quarter notes, eighth notes, and sixteenth notes are the most common musical notes. They are the most important elements of musical notation. In the embodiments of this application, musical notes are used.

[0081] 7) Timing, that is, the order in which audio signals are played in the time domain. Timing refers to multiple audio signals. For example, an audio signal with a time duration of 10 seconds is divided into sub-audio signal 1 and sub-audio signal 2, which are consecutive in the time domain. Sub-audio signal 1 corresponds to 0 to 5 seconds in the time domain, and sub-audio signal 2 corresponds to 5 to 10 seconds in the time domain. Then the timing of sub-audio signal 1 precedes the timing of sub-audio signal 2.

[0082] 8) Time domain features describe signal characteristics in the time domain. They are the most basic and intuitive form of signal expression. Time domain features are functions of time and are used to analyze the signal's values ​​at different times and how they change.

[0083] 9) Frequency domain features refer to the characteristics obtained by analyzing the frequency variation and distribution of a signal. Frequency domain features include the signal’s transfer function, input function, output function, and their Laplace transform and inverse transform.

[0084] Embodiments of the present application provide an audio processing method, an audio processing device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the efficiency and richness of audio production.

[0085] The following describes exemplary applications of electronic devices provided by embodiments of the present application. The electronic devices provided by embodiments of the present application can implement terminal devices, such as laptop computers, tablet computers, desktop computers, set-top boxes, smart TVs, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), vehicle-mounted terminals, virtual reality (VR) devices, augmented reality (AR) devices, and other types of user terminals, and can also be implemented as servers. Below, exemplary applications when the electronic device is implemented as a terminal device or a server will be described.

[0086] refer to Figure 1 , Figure 1 This is a schematic diagram of an application mode of the audio processing method provided in an embodiment of the present application; for example, Figure 1 The server 200, the network 300, the terminal device 400 and the database 500 are involved. The terminal device 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0087] For example, the first audio signal can be the sound of clapping, knocking on a table, or door emitted by a user, and the second audio signal can be the sound of an instrument such as a piano. The terminal device 400 collects the first audio signal and converts the first audio signal into audio data of the first audio signal. The audio data is sent to the server 200 via the network 300. The server 200 extracts a large amount of music symbols, music data, etc. from the database 500. The server 200 generates audio data of the second audio signal based on the audio data of the first audio signal and sends the audio data of the second audio signal to the terminal device 400. The terminal device 400 plays the second audio signal to the user based on the audio data of the second audio signal. The second audio signal can be a signal generated by a sound source completely different from the first audio signal.

[0088] In some embodiments, the audio processing methods of the embodiments of this application can also be applied in the following application scenarios: 1. Music creation: Based on the audio processing methods provided in the embodiments of this application, non-instrumental sounds such as clapping or knocking on a wooden board can be converted into instrumental sounds, allowing users to create music. 2. Game control: Based on the audio processing methods provided in the embodiments of this application, the user's clapping or knocking on a table can be collected to generate corresponding sound effects in the virtual game scene, and virtual objects can be controlled to perform corresponding operations.

[0089] The embodiments of the present application can be implemented using database technology. A database, in short, can be considered an electronic filing cabinet that stores electronic files, allowing users to add, query, update, and delete data in these files. A "database" is a collection of data that is stored together in a specific manner, can be shared by multiple users, has minimal redundancy, and is independent of applications.

[0090] A database management system (DBMS) is a computer software system designed for managing databases, typically providing basic functions such as storage, retrieval, security, and backup. DBMSs can be categorized by the database model they support, such as relational or XML (Extensible Markup Language); by the type of computer they support, such as server clusters or mobile phones; by the query language they use, such as SQL or XQuery; by performance priorities, such as maximum scale or maximum speed; or by other classification methods. Regardless of the classification method used, some DBMSs are cross-category, for example, supporting multiple query languages ​​simultaneously.

[0091] The embodiments of the present application can also be implemented through cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, as well as the promotion of search services, social networks, mobile commerce and open collaboration, each item may have its own hash code identification mark in the future, and all of them need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0092] In some embodiments, the embodiments of the present application may be implemented through artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI is the study of the design principles and implementation methods of various intelligent machines, enabling them to have the capabilities of perception, reasoning, and decision-making. Machine Learning (ML) is a multidisciplinary interdisciplinary field that involves probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of AI and the fundamental way to make computers intelligent. Its applications are widespread in all areas of AI. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. The pre-trained model is the latest development in deep learning and incorporates the above technologies. In this embodiment of the application, a deep learning neural network model is called based on the collected first audio signal to predict the pitch corresponding to the first sub-audio signal in the first audio signal, and a second audio signal is generated based on each pitch.

[0093] In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal device and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiments of the present application.

[0094] In one implementation scenario, when the audio processing method provided by the embodiment of the present application is applied to the field of gaming, the embodiment of the present application is suitable for some application modes that completely rely on the graphics processing hardware computing power of the terminal device 400 to complete the relevant data calculation of the virtual scene, such as stand-alone / offline mode games, and complete the output of the virtual scene through various types of terminal devices 400 such as smart phones, tablets, and virtual reality / augmented reality devices.

[0095] As an example, types of graphics processing hardware include a central processing unit (CPU) and a graphics processing unit (GPU).

[0096] When forming visual perception of a virtual scene, the terminal device 400 calculates the data required for display through graphics computing hardware, and completes the loading, parsing and rendering of the display data, and outputs video frames that can form visual perception of the virtual scene on the graphics output hardware, for example, presenting two-dimensional video frames on the display screen of a smartphone, or projecting video frames to achieve a three-dimensional display effect on the lenses of augmented reality / virtual reality glasses; in addition, in order to enrich the perception effect, the terminal device 400 can also use different hardware to form one or more of auditory perception, tactile perception, motion perception and taste perception.

[0097] As an example, a client (such as a stand-alone game application) is running on the terminal device 400, and during the operation of the client, a virtual scene including role-playing is output. The virtual scene can be an environment for game characters to interact, such as plains, streets, valleys, etc. for game characters to fight; the virtual object can be a game character controlled by the user, that is, the virtual object is controlled by the real user, and will move in the virtual scene in response to the real user's operation on the controller (such as a touch screen, voice-controlled switch, keyboard, mouse and joystick, etc.). For example, when the real user moves the joystick to the right, the virtual object will move to the right in the virtual scene. It can also remain still, jump, and control the virtual object to perform shooting operations, etc.

[0098] In some embodiments, the terminal device and the server jointly implement the game mode. The game mode implemented by the terminal device and the server in collaboration mainly involves two game modes, namely local game mode and cloud game mode. Among them, the local game mode refers to the terminal device and the server jointly running the game processing logic. The operation instructions entered by the player in the terminal device are partially processed by the terminal device running the game logic, and the other part is processed by the server running the game logic. In addition, the game logic processing run by the server is often more complex and requires more computing power; the cloud game mode refers to the game logic processing run entirely by the server, and the cloud server renders the game scene data into an audio and video stream, and transmits it to the terminal device for display through the network. The terminal device only needs to have basic streaming media playback capabilities and the ability to obtain the player's operation instructions and send them to the server.

[0099] In another implementation scenario, see Figure 1 , Figure 1This is a schematic diagram of the application mode of the audio processing method provided in an embodiment of the present application, which is applied to the terminal device 400 and the server 200, and is suitable for an application mode that relies on the computing power of the server 200 to complete the virtual scene calculation and output the virtual scene on the terminal device 400.

[0100] Taking the visual perception of a virtual scene as an example, the server 200 calculates virtual scene-related display data (such as scene data) and sends it to the terminal device 400 through the network 300. The terminal device 400 relies on graphics computing hardware to complete the loading, parsing and rendering of the display data, and relies on graphics output hardware to output the virtual scene to form visual perception. For example, a two-dimensional video frame can be presented on the display screen of a smartphone, or a video frame with a three-dimensional display effect can be projected on the lenses of augmented reality / virtual reality glasses. As for the perception of the form of the virtual scene, it can be understood that the corresponding hardware output of the terminal device 400 can be used, such as using a microphone to form auditory perception, using a vibrator to form tactile perception, and so on.

[0101] As an example, a client (such as an online version of a game application) is running on the terminal device 400, and during the operation of the client, a virtual scene including role-playing is output. The virtual scene can be an environment for game characters to interact, such as plains, streets, valleys, etc. for game characters to fight; the virtual object can be a game character controlled by the user, that is, the virtual object is controlled by the real user, and will move in the virtual scene in response to the real user's operation on the controller (such as a touch screen, voice-controlled switch, keyboard, mouse and joystick, etc.). For example, when the real user moves the joystick to the right, the virtual object will move to the right in the virtual scene. It can also remain still, jump, and control the virtual object to perform shooting operations, etc.

[0102] In some embodiments, a terminal device or a server can implement the audio processing method provided in an embodiment of the present application by running a computer program. For example, the computer executable instructions can be microprogram-level commands, machine instructions, or software instructions. The computer program can be a native program or software module in the operating system; it can be a native application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a game APP, a music APP, or an instant messaging APP; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module, or plug-in in any form.

[0103] See also Figure 2A , Figure 2Ais a structural diagram of a terminal device provided in an embodiment of the present application. The electronic device may be a terminal device 400. Figure 2A The terminal device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2A Various buses are labeled as bus system 440 .

[0104] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0105] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0106] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0107] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0108] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0109] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0110] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).

[0111] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0112] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0113] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2A An audio processing device 455 stored in the memory 450 is shown, which can be software in the form of a program and plug-in, including the following software modules: a display module 4551, an acquisition module 4552 and an output module 4553. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.

[0114] See also Figure 2B , Figure 2B 2 is a schematic diagram of the structure of the server provided in the embodiment of the present application. The electronic device may be the server 200. Figure 2B The server 200 shown includes: at least one processor 210, a memory 250, and at least one network interface 220. The various components in the server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 2B Various buses are labeled as bus system 240 .

[0115] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0116] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.

[0117] The memory 250 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0118] In some embodiments, the memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0119] Operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0120] A network communication module 252 for reaching other electronic devices via one or more (wired or wireless) network interfaces 220 , exemplary network interfaces 220 including Bluetooth, WiFi, and USB;

[0121] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2B An audio processing device 255 stored in the memory 250 is shown, which can be software in the form of a program and a plug-in, including the following software modules: an acquisition module 2551 and an output module 2552. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.

[0122] The audio processing method provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the terminal device provided in the embodiment of the present application.

[0123] The following describes the audio processing method provided by the embodiment of the present application. As mentioned above, the electronic device that implements the audio processing method of the embodiment of the present application can be a terminal device or a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.

[0124] It should be noted that the audio processing examples below are explained using music production as an example. Those skilled in the art, based on their understanding of the following, can apply the audio processing method provided in the embodiments of this application to the processing of other types of audio.

[0125] See also Figure 3A , Figure 3A This is a flowchart of the audio processing method provided by the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained. Figure 3A The execution subject of the steps is the terminal device.

[0126] In step 301, an audio processing interface is displayed.

[0127] For example, the audio processing interface includes a first recording control; the audio processing interface can be an interface displayed in an audio application of a terminal device. The audio processing interface includes controls for triggering functions such as recording and playing audio on the terminal device. The recording control is used to trigger the recording process of the terminal device when triggered.

[0128] In step 302 , in response to a trigger operation on a first recording control, a first audio signal is collected.

[0129] For example, the triggering operation may be a click, double-click, long press, drag, or gesture command on the first recording control. The terminal device collects sound signals in the environment through a microphone. The first audio signal is an audio signal collected from the environment in which the terminal device is located. The first audio signal may include the following: voice emitted by the user, sound emitted by the user through other objects, sound played by the user through other terminal devices, noise in the environment, or other sounds.

[0130] refer to Figure 5C , Figure 5C This is a schematic diagram of the third interface of the audio processing method provided in an embodiment of the present application. Audio processing interface 501C includes a recording control 502C and a sound effect generation control 503C. In response to recording control 502C being triggered, audio acquisition processing begins; in response to sound effect generation control 503C being triggered, an audio signal generated based on the acquired audio signal is output.

[0131] The audio processing method provided in the embodiment of the present application can be applied to scenarios where users create audio. The sound emitted by the user through an object can be collected by the terminal device through an internal microphone, and multiple volume jump points can be identified from the first audio signal. Based on the multiple volume jump points, multiple first sub-audio signals can be segmented from the first audio signal.

[0132] For example, a jump refers to a sudden change in the process of a signal or parameter on the time axis or frequency axis. In the embodiment of the present application, a volume jump refers to a change rate of the sound represented by the audio signal reaching a threshold and the volume being greater than the volume threshold. The volume jump point can be the starting point where the volume starts to rise or the end point where the volume drops. Figure 4A , Figure 4A is a schematic diagram of an audio signal provided in an embodiment of the present application, Figure 4A In the audio signal diagram, the horizontal axis represents time, and the vertical axis represents the amplitude of the sound wave. First audio signal 401A includes multiple sub-audio signals (e.g., first sub-audio signal 402A, first sub-audio signal 404A, and other sub-audio signal 403A). The starting point of first sub-audio signal 402A is a volume transition point, and the ending point is also a volume transition point. Furthermore, the volume of the first sub-audio signal is different from that of the other audio signals.

[0133] The first sub audio signal is a sound emitted by a user through an object. The types of the first sub audio signal include:

[0134] Type 1: Audio signal of the sound produced by an object being knocked; for example, the sound produced by a user knocking on a door panel or a user knocking on a table.

[0135] Type 2: Audio signal of the sound produced by objects being rubbed; for example, the sound produced by the friction of pieces of paper.

[0136] Type 3: Audio signals of the sound produced by object deformation. For example, object deformation, such as being squeezed or torn, corresponds to an audio signal such as the sound produced by a rubber toy being squeezed.

[0137] In some embodiments, step 302 may be implemented by: in response to a third trigger operation on the first recording control, starting the audio acquisition process; in response to a fourth trigger operation on the first recording control, using the audio signal acquired between the interval between the third trigger operation and the fourth trigger operation as the first audio signal. Figure 5C For explanation, when the user clicks the recording control 502C for the first time, recording starts and the first audio signal is collected. When the user stops making sound and clicks the recording control 502C again, recording ends and the audio signal collected between the two clicks is used as the first audio signal.

[0138] In some embodiments, after step 302, a time-domain volume curve of the first audio signal is determined, wherein different points in the time-domain volume curve represent volume values ​​of the first audio signal at different moments; points that meet a volume jump condition are extracted from the time-domain volume curve as volume jump points, wherein the volume jump condition includes: a volume value corresponding to the point is greater than a volume threshold, and an absolute value of a differential value of the volume corresponding to the point in the time-domain volume curve is greater than that of at least one adjacent point; and in response to the volume of a time period between two temporally adjacent volume jump points being greater than the volume threshold, the first sub-audio signal is segmented from the first audio signal according to the moments corresponding to the two temporally adjacent volume jump points to obtain the first sub-audio signal.

[0139] For example, the volume corresponding to each moment in the time domain of the first audio signal is counted to obtain the time domain volume curve corresponding to the first audio signal. The differential value can be used to characterize the slope of a certain point in a function, and can also be used to describe the rate of change of the function at a certain point. The differential value of the volume refers to the rate of change of the volume in the time domain volume curve. The volume threshold is greater than the volume of the noise in the environment in which the first audio signal is collected. In a specific implementation, the volume threshold can be set according to the needs of the user. When the volume corresponding to the time period between two adjacent volume jump points is greater than the volume threshold, it means that in the first audio signal, the volume of the sub-audio signal corresponding to the time period is greater than the volume threshold, which can be used as the first sub-audio signal. Based on the moments corresponding to the two adjacent volume jump points, the first sub-audio signal is obtained by segmenting from the first audio signal in the time domain.

[0140] For ease of understanding, the following description is given with reference to the accompanying drawings. Figure 9 , Figure 9 This is a second schematic diagram of a volume curve provided in an embodiment of the present application. Figure 9 The time domain volume curve 901 is Figure 4AThe time-domain volume curve of the first audio signal 401A is shown in Figure 901. The horizontal axis of time-domain volume curve 901 represents time, and the vertical axis represents volume, with the unit of volume being decibels (dB). In time-domain volume curve 901, points 905 and 906 are volume transition points. The differential value of the volume corresponding to point 905 increases compared to adjacent points, and the volume corresponding to point 905 is greater than the volume threshold. Points 906 and 905 are adjacent volume transition points, and the volume corresponding to curve segment 902 between points 906 and 905 is greater than the volume threshold. Based on the time instants corresponding to points 906 and 905, the first audio signal 401A is segmented in the time domain to obtain first sub-audio signal 402A. Similarly, the endpoints of curve segment 903 are also volume transition points. Based on the time instants corresponding to the endpoints of curve segment 903, the first audio signal 401A is segmented in the time domain to obtain first sub-audio signal 404A.

[0141] In the embodiment of the present application, a first audio signal is segmented by volume transition points to obtain a sub-audio signal with a higher volume than other parts of the first audio signal. The segmented first sub-audio signal is then used as the basis for generating a second audio signal. Compared to a solution that generates audio based on the entire first audio signal, this saves computing resources required for generating the audio signal. Generating an audio signal based only on the sub-audio signal with a higher volume than other parts of the first audio signal also improves the accuracy of the audio generation process.

[0142] Continue to refer Figure 3A In step 303, in response to the first audio signal including a plurality of volume jump points, a second audio signal is output.

[0143] For example, the second audio signal is generated based on pitches corresponding to the plurality of first sub-audio signals, and the plurality of first sub-audio signals are separated from the first audio signal based on a plurality of volume jump points.

[0144] The sound source of the second audio signal can be completely different from the sound source of the first audio signal. For example, if the sound source of the first sub-audio signal in the first audio signal is the sound of a user knocking on a wooden board, the pitch corresponding to the sound of the user knocking on the wooden board is determined, and music or sound effects are generated based on the pitch and a specific sound source (such as a piano) to form a second audio signal, and the second audio signal is played. Figure 5C After the control 502C is triggered for the second time, in response to the control 503C for generating a sound effect being triggered, an audio signal generated based on the collected audio signal is output.

[0145] In some embodiments, reference Figure 3B , Figure 3B This is a second flow chart of the audio processing method provided in the embodiment of the present application; before step 303, execute Figure 3B Steps 3031A to 3033A in the embodiment of the present invention are used to obtain the second audio signal, which will be described in detail below.

[0146] In step 3031A, the pitch heights corresponding to the plurality of first sub-audio signals are obtained.

[0147] For example, the pitch heights respectively corresponding to the first sub-audio signals may have a preset mapping relationship with the first sub-audio signals, or may be randomly generated.

[0148] In some embodiments, there is a mapping relationship between the first sub-audio signal and the corresponding pitch height. Obtaining the pitch heights corresponding to multiple first sub-audio signals can be achieved through a neural network model. Taking a deep learning network model as an example, the deep learning network model has pre-learned the relationship between sample audio signals with the same characteristics as the first sub-audio signal and the pitch height.

[0149] Step 3031A can be implemented in the following manner: calling a deep learning network model based on each first sub-audio signal, and performing the following processing: performing feature extraction processing on the first sub-audio signal to obtain time domain features and frequency domain features; performing prediction processing based on the time domain features and frequency domain features corresponding to the first sub-audio signal to obtain the predicted probability of mapping the first sub-audio signal to different pitch heights; and taking the pitch height with the maximum predicted probability as the pitch height corresponding to the first sub-audio signal.

[0150] For example, the deep learning network model can be pre-trained. The training process will be explained in detail below. Here we only explain the application process of the application model. Figure 4C In the embodiment of the present application, the neural network model is taken as an example to illustrate that it is a deep learning network model. The deep learning network model 401C includes a feature extraction layer 402C and a feature classification layer 403C. The feature extraction layer 402C is composed of multiple convolutional layers, and the feature classification layer 403C includes a gated unit (GRU), a fully connected unit (FC), and a normalization layer (sigmoid). The feature extraction layer 402C is used to perform Fourier transform on the first sub-audio signal to obtain the original power spectrum of the first sub-audio signal, perform feature extraction processing on the original power spectrum, and obtain time domain features and frequency domain features. The time domain features and frequency domain features can be used to more accurately segment the first sub-audio signal from the first audio signal based on the jump point. The feature classification layer predicts the predicted probability of mapping the first sub-audio signal to different pitch heights based on the time domain features and the frequency domain features, and uses the pitch height with the maximum predicted probability as the pitch height corresponding to the first sub-audio signal.

[0151] In some embodiments, there is no clear mapping relationship between the pitch height and the first sub-audio signal, and the pitch height can be randomly generated. Step 3031A can be implemented in the following way: obtain a mapping relationship table between the pitch height and the label value, wherein each pitch height in the mapping relationship table corresponds one-to-one to each label value; perform the following processing for each first sub-audio signal: perform random number generation processing based on the first sub-audio signal to obtain a random value; in response to the first value being equal to the label value, use the pitch height corresponding to the same label value as the pitch height corresponding to the first sub-audio signal.

[0152] For example, based on the different sounds produced by different sound sources at the same pitch, the total number of random values ​​corresponding to each pitch is the product of the pitch and the sound source type. Furthermore, in the mapping table, the same pitch corresponding to different sound source signals is considered a type, and each type corresponds to a random value. For example, a sound source type piano at pitch A corresponds to random number 1, and a sound source type violin at pitch A corresponds to random number 2.

[0153] For example, a random number generation process is performed based on the first sub-audio signal C, where the random range of the random number generation process is the number of pitch types in the mapping relationship table. A random value N is obtained, where N is a positive integer, and the pitch corresponding to the random value N in the mapping relationship table is obtained. The random number generation process can be implemented using the Java Math.random() function. Assuming that the number of pitch types is M, where M is a positive integer, calling the function (int)(Math.random()*M)+1 can generate a random positive integer greater than or equal to 1 and less than or equal to M.

[0154] In the embodiment of the present application, the pitch corresponding to the first sub-audio signal is determined based on the determination method of the random number generation process, which can improve the richness of the generated audio signal.

[0155] In some embodiments, step 3031A can also be implemented as follows: 1. Obtaining the average volume of the first sub-audio signal, obtaining a mapping relationship table between different sub-volume intervals within the volume interval and different note signals, determining the note signal mapped to the sub-volume interval to which the average volume belongs, and using the mapped note signal as the pitch corresponding to the first sub-audio signal. 2. Obtaining actual audio features corresponding to the first sub-audio signal and reference audio features for each reference audio signal, where the reference audio signals are generated based on pitch, obtaining similarities between the actual audio features and the reference audio features, selecting a target reference audio signal corresponding to the highest similarity, and using the pitch of the generated target reference audio signal to generate the second sub-audio signal.

[0156] In an embodiment of the present application, the pitch height corresponding to the first sub-audio signal is determined based on the parameters of the first sub-audio signal and the mapping relationship. The correlation between the first sub-audio signal and the generated audio signal is ensured through the mapping relationship, thereby improving the accuracy of the audio generation process.

[0157] Continue to refer Figure 3B In step 3032A, a plurality of second sub audio signals are generated based on the pitches corresponding to the plurality of first sub audio signals.

[0158] Here, the pitch corresponding to a first sub-audio signal is used to generate a second sub-audio signal, the first sub-audio signal and the corresponding second sub-audio signal have the same or different durations, and the timing of the multiple second sub-audio signals is the same as the timing of the first sub-audio signals corresponding to the multiple second sub-audio signals.

[0159] For example, the first sub-audio signal corresponding to the second sub-audio signal refers to the first sub-audio signal corresponding to the pitch used to generate the second sub-audio signal. The same timing specifically means that the playback order of the second sub-audio signal in the second audio signal is the same as the playback order of the first sub-audio signal corresponding to the second sub-audio signal in the first audio signal. For example, if the playback order of first sub-audio signal A among all first sub-audio signals in the first audio signal is the third, then the playback order of the second sub-audio signal corresponding to first sub-audio signal A in the second audio signal is also the third.

[0160] In some embodiments, step 3032A may be implemented as follows:

[0161] The following processing is performed on each first sub-audio signal: obtaining a first duration of the first sub-audio signal; determining a sound source type pre-associated with the pitch; determining a second duration of the second sub-audio signal based on the first duration of the first sub-audio signal; and generating an audio signal of a second duration based on the second duration, the sound source type, and the pitch as the second sub-audio signal.

[0162] For example, the first duration of the first sub-audio signal refers to the duration between the corresponding start and end times of the first sub-audio signal in the time domain. The sound source type determines the timbre of a sound. Different sound source types produce different tones corresponding to the same note. For example, if sound source type 1 is a piano and sound source type 2 is a violin, the sound of the musical symbol do on the piano scale will sound different from the sound of the violin.

[0163] The second duration of the second sub audio signal may be the same as the first duration of the first sub audio signal, or the second duration is obtained by multiplying the first duration by a preset coefficient, and the second duration is positively correlated with the first duration.

[0164] For example, the sound source types corresponding to the pitch height include: different types of musical instruments, pre-configured sound effects. The method of generating the second sub-audio signal can be achieved in the following way: obtain the initial audio signal of the sound corresponding to the pitch height emitted by the sound source type that has been stored in the database, adjust the duration of the initial audio signal to the second duration, and use the adjusted audio signal as the second sub-audio signal. For another example: if the corresponding initial music signal does not exist in the database, the corresponding audio generation model can be called to generate the corresponding initial audio signal based on artificial intelligence technology, and the duration of the initial audio signal can be adjusted to the second duration to obtain the second sub-audio signal.

[0165] In some embodiments, if the volume of the second sub audio signal does not meet user requirements, the user can adjust the volume of the generated second sub audio signal, and generate a second sub audio signal based on the adjusted second sub audio signal. The volume adjustment method can be to make the volume of all parts the same, or to insert key frames with different volumes into the second sub audio signal.

[0166] In the embodiment of the present application, through the above-mentioned scheme, the sound of musical instruments or sound effects can be generated based on the sound emitted by the deformation of objects and the sound of knocking, thereby realizing audio generation processing across sound source types, and being able to meet the user's various needs for audio production. Compared with the traditional scheme that can only obtain new audio based on recorded audio clips, it can generate audio signals of different sound source types, improve the audio processing efficiency, and when applied in music creation scenarios, it can enhance the user's music creation experience.

[0167] In step 3033A, a second audio signal is synthesized based on the plurality of second sub audio signals according to the time sequence of the plurality of second sub audio signals.

[0168] For example, a method of synthesizing the second audio signal based on the multiple second sub-audio signals includes:

[0169] Method 1: sort each second sub audio signal in sequence according to time sequence to form a second audio signal.

[0170] For easier understanding, refer to Figure 4D , Figure 4D1 is a schematic diagram of the first principle of synthesizing a second audio signal provided by an embodiment of the present application. The audio track of the first sub-audio signal contains the collected first audio signal, including multiple first sub-audio signals (401D, 402D, and 403D) and other audio signal 404D. A corresponding second sub-audio signal 405D is generated based on the first sub-audio signal 401D, and a corresponding second sub-audio signal 406D is generated based on the first sub-audio signal 402D. The timing of the first sub-audio signal 401D and the second sub-audio signal 405D are identical, both being first in the playback order of their respective signals.

[0171] Method 2: sort each second sub audio signal in sequence according to time sequence, and insert a corresponding transition signal between each second sub audio signal to form a second audio signal.

[0172] For easier understanding, refer to Figure 4E , Figure 4E 4 is a schematic diagram of a second principle for synthesizing a second audio signal according to an embodiment of the present application. The audio signals in the audio track of the second sub-audio signal form a second audio signal, which sequentially includes multiple audio signals, including a second sub-audio signal 405D, a transition audio signal 404E, and a second sub-audio signal 406D.

[0173] In some embodiments, step 3033A can be implemented by: synthesizing a transition audio signal corresponding to each two second sub-audio signals based on each two second sub-audio signals that are adjacent in time sequence; arranging each second sub-audio signal according to the time sequence of the multiple second sub-audio signals, and inserting a corresponding transition audio signal between each two second sub-audio signals that are adjacent in time sequence to form the second audio signal.

[0174] For example, the corresponding transition audio signal refers to an inserted transition audio signal generated based on every two second sub-audio signals.

[0175] In some embodiments, the transition audio signal corresponding to every two second sub-audio signals can be obtained by performing the following processing on every two second sub-audio signals that are adjacent in time sequence: using the first second sub-audio signal as the first signal to be processed and the second second sub-audio signal as the second signal to be processed; obtaining a tail segment of the first signal to be processed and a head segment of the second signal to be processed; superimposing the tail segment and the head segment to obtain a superimposed signal; and performing volume adjustment processing on the superimposed signal to obtain the transition audio signal corresponding to the two second sub-audio signals that are adjacent in time sequence.

[0176] For example, during the superposition process, the tail of the tail segment and the head of the head segment at least partially overlap, and the duration of the superposition signal is much shorter than the duration of the two second sub-audio signals adjacent in time sequence; the head segment and the tail segment are determined from the time domain perspective, and are intercepted from the tail / head according to a preset ratio of duration. For ease of understanding, continue to refer to Figure 4E After generating multiple second sub-audio signals based on multiple first sub-audio signals, taking second sub-audio signals 405D and 406D as examples, a tail segment 401E is cut from the tail of second sub-audio signal 405D, and a head segment 402E is cut from the head of second sub-audio signal 406D. Tail segment 401E and head segment 402E are placed on different audio tracks and superimposed by merging the tracks to generate a superimposed signal 403E. The volume of superimposed signal 403E is adjusted to generate a transition audio signal 404E.

[0177] For example, after volume adjustment, the volume curve of the transition audio signal is characterized by gradually decreasing from an initial volume and then gradually increasing back to the initial volume, so that when the second audio signal is played, the user can perceive a smooth transition between adjacent second sub-audio signals. The initial volume refers to the original volume at the starting position of the transition audio signal, and the initial volume is the same as the volume at the ending position of the second sub-audio signal before the transition audio signal.

[0178] refer to Figure 4G , Figure 4G This is a first schematic diagram of a volume curve provided in an embodiment of the present application. The horizontal axis of the volume curve represents time, and the vertical axis represents volume. The first curve 401G is the volume curve of the transition audio signal before volume adjustment, and the second curve 402G is the volume curve of the transition audio signal after volume adjustment. The third curve 403G is the volume curve of the two second sub-audio signals inserted into the transition audio signal. The volume corresponding to point 404G represents the volume at the end position of the second sub-audio signal. After volume adjustment, the volume of the transition audio signal in the second curve 402G decreases from the initial volume and then increases.

[0179] In some embodiments, the transition audio signal may be a preset audio signal, and the volume of the preset transition audio signal also conforms to gradually decreasing from the initial volume and then gradually increasing to the initial volume, so that the user can feel the smooth transition of the audio from the body sense, thereby improving the user experience.

[0180] In the embodiment of the present application, by inserting a transition audio signal between the generated second sub-audio signals, the smoothness of the generated second audio signal can be improved, thereby enhancing the user's experience of the second audio signal.

[0181] In some embodiments, reference Figure 3C , Figure 3C This is a third flow chart of the audio processing method provided in the embodiment of the present application; before step 303, execute Figure 3C Steps 3031B to 3033B in the process are described in detail below.

[0182] In step 3031B, the pitch heights corresponding to the plurality of first sub-audio signals are obtained.

[0183] For example, the principle of step 3032B is the same as that of step 3032A, and will not be repeated here.

[0184] In step 3032B, a plurality of second sub audio signals are generated based on the pitches corresponding to the plurality of first sub audio signals.

[0185] Here, the pitch corresponding to a first sub-audio signal is used to generate a second sub-audio signal. The principle of step 3032B is the same as that of step 3032A, and will not be repeated here.

[0186] In step 3033B, multiple first sub audio signals in the first audio signal are replaced with second sub audio signals corresponding to the first sub audio signals to form a second audio signal.

[0187] For example, multiple first sub-audio signals in a first audio signal are muted, that is, the first audio signal is multiplexed, and the signal portion corresponding to each first sub-audio signal in the first audio signal is eliminated so that the corresponding portion has a volume of zero. In the time domain, a second sub-audio signal corresponding to the first sub-audio signal is superimposed on the portion corresponding to the volume of zero to obtain a second audio signal, wherein the first sub-audio signal corresponding to the second sub-audio signal is the first sub-audio signal corresponding to the pitch height used to generate the second sub-audio signal. Figure 4F This is a third schematic diagram of the principle of synthesizing a second audio signal provided by an embodiment of the present application. Multiple first sub-audio signals (401D to 403D) in a first audio signal are muted. The audio track of the first sub-audio signal includes other audio signals 404D and a blank playback position 401F after muting. Taking the first sub-audio signal 401D as an example, a second sub-audio signal 405D generated based on the first sub-audio signal 401D is superimposed on the blank playback position 401F corresponding to the first sub-audio signal 401D to form the second audio signal. The audio file corresponding to the first audio signal is multiplexed.

[0188] The second audio signal formed in the above manner retains a portion of the first audio signal except for the first sub-audio signal. Eliminating the first sub-audio signal from the first audio signal may include performing a Fourier transform on the first audio signal to extract an original power spectrum, performing feature extraction on the original power spectrum to obtain features of the first sub-audio signal, and performing an inverse Fourier transform on a result of superimposing the features of the first sub-audio signal and the original power spectrum to obtain the first audio signal from which the first sub-audio signal has been eliminated.

[0189] In an embodiment of the present application, a second audio signal in the form of music or sound effects can be generated across sound source types through a first sub-audio signal that is not music or sound effects, thereby improving the efficiency and richness of audio processing and enhancing the user's audio production experience.

[0190] In some embodiments, the pitch heights corresponding to the plurality of first sub-audio signals are obtained through a pre-trained neural network model, and the audio processing interface includes a second recording control, a third recording control, and a model training control. Figure 3D , Figure 3D This is a fourth flow chart of the audio processing method provided in the embodiment of the present application, executing Figure 3D Steps 3021 to 3025 train the deep learning network model, as described in detail below.

[0191] In step 3021, a sample pitch associated with a sample reference signal to be recorded is obtained.

[0192] For example, there are two ways to obtain the pitch of a sample:

[0193] Method 1: The user selects the sound source type and the pitch corresponding to the sound source type.

[0194] In some embodiments, the audio processing interface also includes: a first option list, the first option list includes first options corresponding to multiple different sound source types; step 3021 can be implemented in the following manner: in response to a first selection operation for the first option, the sound source type corresponding to the first option selected in the first selection operation is used as a sample sound source type; at least one pitch height associated with the sample sound source type is displayed; in response to a second selection operation for the pitch height, the pitch height selected in the second selection operation is used as the sample pitch height of the sample reference signal.

[0195] refer to Figure 5A , Figure 5AThis is a schematic diagram of the first interface of the audio processing method provided by an embodiment of the present application; interface 501A includes an option list 503A, multiple controls 504A, and control 505A. When the user performs a first selection operation on the option list 503A, the sound source type corresponding to the first option selected in the first selection operation is used as the sample sound source type. Figure 5A The example of the sample sound source type being a violin is used to illustrate, and the notes C, D, etc. corresponding to the violin are displayed. Note C and note D represent different pitch heights respectively. In response to the user's second selection operation on the control 504C corresponding to note C, note C is used as the sample pitch height of the sample reference signal.

[0196] Method 2: Provide a default sound source type, and the user can select the corresponding note.

[0197] For example, the principle of selecting notes and the recording process in method 2 can refer to method 1.

[0198] In step 3022, in response to the first trigger operation on the second recording control, audio acquisition processing is started and the third recording control is displayed.

[0199] For example, when the third recording control is displayed, a prompt message "Recording" can also be displayed. The third recording control is used to trigger the terminal device to stop performing audio collection processing. Figure 5B ,refer to Figure 5B , Figure 5B This is a schematic diagram of the second interface of the audio processing method provided in the embodiment of the present application; Figure 5A When control 504A for any note in the audio recording is triggered, audio signal acquisition begins and the screen presented in interface 501B is displayed. Interface 501B includes prompt information 502B and control 503B; prompt information 502B and control 503B are displayed as a floating layer above the layers corresponding to other controls in interface 501B. Prompt information 502B reminds the user that the terminal device's microphone is turned on and recording is in progress. The user can produce different sounds by tapping, rubbing, or squeezing objects.

[0200] In step 3023, in response to the second trigger operation on the third recording control, the audio signal collected between the interval time of the first trigger operation and the second trigger operation is used as the third audio signal.

[0201] For example, continue to refer to Figure 5B When the user finishes making a sound through the object, in response to the triggering operation of the control 503B for the end of recording, the display Figure 5A The audio signal collected between the interval between the first trigger operation and the second trigger operation is used as the third audio signal.

[0202] In step 3024, the third audio signal is used as a sample reference signal corresponding to the sample pitch.

[0203] For example, a mapping relationship between the third audio signal and the sample pitch is established in the terminal device, and the mapping relationship can be used as data for training the model.

[0204] In step 3025, in response to a trigger operation for a model training control, the sample reference signal and the sample pitch height are combined into a sample pair, and the initialized neural network model is trained based on the sample pair to obtain a pre-trained neural network model.

[0205] For example, the pre-trained neural network model is used to determine the pitch heights corresponding to the plurality of first sub-audio signals. The pre-trained neural network model may be a deep learning network model, and further reference is made to Figure 5A In response to the triggering operation of the control 505A for starting training recognition in the interface 501A, the neural network model for recognizing audio signals is trained. When the model training is completed, a corresponding prompt message can be displayed in the interface 501A to inform the user that the training is completed.

[0206] For example, a neural network model initialized based on sample training is used to obtain a pre-trained neural network model, which can be achieved in the following way: calling a deep learning network model to perform feature extraction processing based on a sample reference signal to obtain sample time domain features and sample frequency domain features of the sample reference signal; performing prediction processing based on the sample time domain features and sample frequency domain features to obtain the predicted probability of mapping the sample reference signal to different pitch heights; using the pitch height with the maximum predicted probability as the predicted pitch height corresponding to the sample reference signal; determining the loss function of the deep learning network model based on the difference between the sample pitch height and the predicted pitch height; and updating the parameters of the deep learning network model based on the loss function to obtain a pre-trained neural network model.

[0207] For example, the difference between the sample pitch and the predicted pitch represents the difference between the types, which can be determined by the probability corresponding to the label value. The loss function can be a cross-entropy loss or an information entropy loss. Updating the parameters of the deep learning network model can be a backpropagation process.

[0208] In some embodiments, after step 301, in response to the first audio signal including at least one volume jump point and the first sub-audio signal included in the first audio signal not belonging to the learned audio type, the first audio signal is used as a new sample signal, wherein the learned audio type is a sample sound source type that has been learned by the deep learning network model; the deep learning network model is trained based on the new sample signal to obtain an updated neural network model.

[0209] For example, when the sound emitted by the user through an object does not belong to the sound corresponding to the learned sample sound source type, the data of the new sound signal is stored in the terminal device, or uploaded to the server through the terminal device. The terminal device or server trains and processes the deep learning network model based on the new sample signal to realize online update of the deep learning network model and improve the accuracy of audio generation and processing of the deep learning network model in practical applications.

[0210] In some embodiments, before step 303, a virtual scene is displayed, wherein the virtual scene includes a virtual object; when the second audio signal is output, in response to the volume of the first sub-audio signal being greater than a volume threshold, and the first sub-audio signal and the reference sound signal being sound signals emitted by objects of the same type, a screen of the virtual object performing a preset operation is displayed.

[0211] For example, the reference sound signal is pre-configured, and the types of preset operations include:

[0212] Type 1: Movement of virtual objects; for example, a virtual object moves by itself, or moves via a vehicle. Movement can be by flying, running, jumping, or walking.

[0213] Type 2: Interaction between virtual objects and virtual props. For example: virtual objects use virtual props, virtual objects replace virtual props. Figure 8 , Figure 8 8 is a schematic diagram of the sixth interface of the audio processing method provided in an embodiment of the present application. Virtual scene 801 includes virtual object 802 wearing component 803. In response to a user's clapping sound signal, virtual object 802 wearing component 805 is displayed in virtual scene 801, along with prompt 804 stating "Appearance Changed."

[0214] Type 3: Interactions between virtual objects. Examples include physical contact (e.g., hugs, handshakes) and non-physical contact (e.g., waving).

[0215] For example, the user pre-binds the relationship between the operation of the virtual object and the reference sound signal in the game. Then, the user can tap the surrounding objects as shortcut keys to trigger the corresponding operation of the virtual object, thereby improving the efficiency of the user's control of the virtual object and improving the efficiency of human-computer interaction.

[0216] The present application also provides an audio processing method. Figure 3E , Figure 3E This is a flowchart of the audio processing method provided by the embodiment of the present application, which will be combined with Figure 3E The steps shown are explained. Figure 3E The execution subject of the steps is the terminal device.

[0217] In step 311 , an audio processing interface is displayed.

[0218] For example, the audio processing interface includes an audio signal list including at least one pre-recorded audio signal. The pre-recorded audio signal can be a user-recorded audio signal or an audio signal corresponding to audio data obtained from other channels. For example, the second audio signal can be generated based on an existing recording file.

[0219] In step 312 , in response to a selection operation on the audio signal list, the selected audio signal is used as the first audio signal.

[0220] Example, reference Figure 5E , Figure 5E This is a schematic diagram of the fifth interface of the audio processing method provided in an embodiment of the present application. An audio signal list 502E is displayed in the audio processing interface 501E. The audio signal list 502E includes icons corresponding to multiple audio files. In response to a selection operation on any audio file, the audio signal of the selected audio file is used as the first audio signal.

[0221] In step 313 , in response to the first audio signal including a plurality of volume transition points, a second audio signal is output.

[0222] The second audio signal is generated based on the pitch heights corresponding to the plurality of first sub-audio signals, and the plurality of first sub-audio signals are separated from the first audio signal based on the plurality of volume jump points. The principle of step 313 is referred to in step 303 above and will not be described here. In some embodiments, continue to refer to Figure 5E , in response to a trigger operation on control 503E, a second audio signal is output.

[0223] The present application also provides an audio processing method. Figure 3F , Figure 3F This is a flowchart of the audio processing method provided by the embodiment of the present application, which will be combined with Figure 3F The steps shown are explained. Figure 3F The execution subject of the steps is the terminal device or server.

[0224] In step 321, a first audio signal is acquired.

[0225] For example, the principle of step 321 refers to step 301 or step 311 above, and will not be repeated here.

[0226] In step 322, volume recognition processing is performed on the first audio signal to obtain a plurality of volume jump points.

[0227] In step 323 , a plurality of first sub-audio signals are separated from the first audio signal based on the plurality of volume jump points.

[0228] For example, the principles of step 322 to step 323 refer to step 302 above and are not repeated here.

[0229] In step 324 , the pitches corresponding to the plurality of first sub-audio signals are obtained.

[0230] In step 325 , a second audio signal is generated based on each pitch.

[0231] For example, the principles of step 324 to step 325 refer to step 303 above, which will not be repeated here.

[0232] In an embodiment of the present application, audio processing is performed based on a sub-audio signal with jumps in the collected first audio signal. Compared with a solution that performs full processing based on the collected signal, the accuracy of the generated audio signal is improved and the computing resources required for audio processing are saved; the second audio signal is generated based on the pitch heights corresponding to the first sub-audio signals, so that the first audio signal and the second audio signal can cross the sound source type, thereby improving the richness of the generated audio signal.

[0233] Below, an exemplary application of the audio processing method provided in an embodiment of the present application in a practical application scenario will be described.

[0234] Related technologies include the following: 1. Collecting real-world sounds and editing them to create music or sound effects; 2. Hand motion image recognition: Users use sensors or cameras to collect and extract dynamic hand image data, analyze hand feature joints, such as finger movements and trajectories, and map finger movements to corresponding sound effects. For example, left hand pinky movement corresponds to scale 1 "do," and left hand ring finger movement corresponds to scale 2 "re." Simultaneously with the user's finger movements, the system recognizes and outputs a corresponding music signal, which can be played to a speaker or other device. 3. Vibration sensing: A customized vibration sensor is placed on the surface of a vibrating object. The user taps the surface to generate vibration signals of varying degrees. The sensor detects the vibration direction and intensity, maps them to different sound effects, and ultimately generates a music signal, which can be played to a speaker or other device. The vibration sensor used to collect vibrations is connected to a terminal device.

[0235] However, the methods for producing music or sound effects in related technologies have limitations. Methods based on audio editing captured from the real world can only generate music from the same source as the captured sound. Hand movement image recognition requires support from a camera image acquisition device, and the camera's shooting angle has strict requirements. Incorrect placement can easily lead to misidentification or omission due to occlusion of hand joints. Image recognition also has limited ability to detect the intensity of user movements, and can only indirectly judge based on movement speed. However, some subtle movements are difficult to detect and identify with ordinary cameras. For example, finger tapping has a very short movement range, small amplitude, and short duration. Ordinary cameras have a relatively low frame rate, making speed detection difficult, and even more difficult to identify the strength of hand movements. Consequently, the diversity and matching of sound effects generated are difficult to ensure, affecting the user's real experience. Vibration sensing detection requires the use of an external customized vibration sensor for vibration detection and analysis. However, the amount of detectable information is relatively limited, resulting in a poor user experience.

[0236] An embodiment of the present application proposes an audio processing method that can generate different sound effects and music based on the sound that a user can produce by knocking on any object. The user can knock on one or more objects according to the preset prompt steps to produce different types of sounds. The terminal device or server recognizes the collected knocking sound signals and matches them with the user-defined sound effects or musical instruments to generate music or sound effects based on the sounds of the customized sound effects or musical instruments.

[0237] The following is an explanation of the audio processing method provided by the embodiment of the present application with reference to the accompanying drawings. Figure 6A , Figure 6A This is a flow chart of the audio processing method provided in an embodiment of the present application; the following description is given using the example of a terminal device as the execution subject.

[0238] In step 601A, a mapping relationship between a pre-configured pitch and a sound signal is obtained, and a neural network model is trained.

[0239] The audio processing method provided in the embodiments of the present application can be applied to music creation scenarios. The mapping relationship between the pre-configured pitch and the sound signal can be customized by the user. The user can customize the mapping relationship between the pitch and the emitted sound signal. The emitted sound signal can be generated by knocking, rubbing, or squeezing an object. For example, it can be: knocking on a table, clapping, tearing paper, or regularly rubbing cloth.

[0240] The following explains the process of mapping the relationship between user-defined pitch height and sound signal, refer to Figure 5A , Figure 5AThis is a schematic diagram of the first interface of the audio processing method provided in an embodiment of the present application; the interface 501A includes an option list 503A, multiple controls 504A and a control 505A.

[0241] The user can select the musical instrument corresponding to the generated audio signal, such as piano, drums and violin, etc. When the user selects the violin option, a violin icon 502A is displayed. When the control 504A of any note is triggered, the audio signal is collected.

[0242] For example, the embodiments of the present application are only given as examples, and the multiple pitch heights can be musical instrument notes, or one of a specific sound effect or other pre-stored sounds.

[0243] For example, when interface 501A can be the application interface of music software, the user opens the application software and displays a custom pitch height matching page (interface 501A), which displays music symbols of different musical instruments, different sound effect categories or different scales. The user records a kind of object knocking sound at the selected pitch height, and is trained by the system to recognize the subsequent user's knocking sound on this object. When the system recognizes the object knocking sound, it will be mapped to the corresponding pitch height. For example, the user knocks on the wooden door corresponding to the big drum sound, and the clapping of the hands corresponds to the gong sound. The system adjusts the volume and pitch of the output pitch height according to the intensity and spectrum analysis of the knocking sound. The final user knocks on the object at his own rhythm, and the system will recognize the knocking object sound in real time or non-real time, and convert it into the corresponding pitch height. The user can select the pure pitch height to generate output (for example: in real time, directly generate audio based on all knocking sounds), or the system automatically eliminates the knocking sound and replaces it with the corresponding pitch height based on the original recording. For example, if a recording signal is acquired in non-real time and the original recording signal is composed of multiple percussive sounds, the percussive sounds are automatically removed and replaced with a sound signal of the corresponding pitch. Another example is if the generated music is pure music, the generated music retains the rest of the recording signal except for the percussive sounds in addition to the sound effects.

[0244] For example, the user can select the desired instrument or sound effect component through the drop-down menu (option list 503A). After selecting the component, the interface displays the selected content image and lists the various note elements corresponding to the component. Figure 5B , Figure 5B This is a schematic diagram of the second interface of the audio processing method provided in the embodiment of the present application; Figure 5AWhen the control 504A of any note in the interface is triggered, the audio signal is collected and the screen presented by the interface 501B is displayed. The interface 501B includes a prompt message 502B and a control 503B; the prompt message 502B and the control 503B are displayed in a floating manner above the layers corresponding to other controls in the interface 501B. The prompt message 502B is used to remind the user that the microphone of the terminal device has been turned on and is in the recording process. The user can make different sounds by knocking, rubbing, and squeezing objects. When the sound is finished, the interface 501A is displayed in response to the trigger operation of the control 503B for the end of recording. In response to the trigger operation of the control 505A for starting training recognition in the interface 501A, the neural network model for recognizing audio signals is trained.

[0245] For example, the user clicks "Record Audio" after each note element to record. Each time the "Record Audio" control is clicked, the interface displays a prompt message saying "Audio is being recorded". At this time, the user taps different objects corresponding to different notes to produce different sound signals. After the audio corresponding to each note element is recorded, the user clicks "End Recording" to start the sound recording of the next note element. When the user completes the audio recording corresponding to each note element, the user clicks "Start Training Recognition" and the system enters the process of training the neural network model.

[0246] refer to Figure 4B , Figure 4B This is a diagram of the training principle of the neural network model provided in an embodiment of the present application; the recording signal is converted into the spectral features of the power spectrum, and the spectral features are output through the neural network model to identify the result, and the identification result is a classification label.

[0247] For example, the recorded audio signal undergoes spectral analysis, and a deep learning network is trained based on the spectral analysis results. The goal of the training process is to enable the system to recognize and identify the differences in the user's different groups of recorded signals and their characteristics, and to correctly classify the input audio signals in subsequent applications. Based on the input audio signal, the system determines the musical note element associated with the audio signal. For example, if the user selects piano as the desired instrument sound, the note C is recorded as the sound of knocking on a door panel. The user knocks on the door panel a preset number of times (for example, 5 times), resulting in the preset number of knocking sounds. The preset number of sounds is extracted to ensure sufficient samples, improve training accuracy, and avoid recognition errors. The user selects the notes D / E / F / G in sequence, and uses the same method to record the sounds of knocking on a wooden barrel (D), knocking on iron sheets (E), clapping hands (F), and stepping on the floor (G) for each note. After recording is completed, the user clicks to start training recognition. The notes C / D / E / F / G correspond to different fundamental pitches. Fundamental pitches refer to the seven independent pitches with fixed names in the musical sound system.

[0248] Identifying the collected sound signals can be achieved in the following ways: converting the input audio signal into a frequency domain signal, and extracting the time and frequency domain features of the knocking sound, such as the duration of the knocking sound, the effective range of the frequency domain, the frequency domain energy distribution characteristics, etc. The data of the extracted knocking sound is the frame data corresponding to the energy greater than the threshold and the obvious mutation. The features of the knocking sound are input into the deep learning network, and the deep learning network allows the input features of various types of knocking sounds to show clearer distinction at a higher feature dimension; the recognition result of the deep learning network is a classification label (label), specifically different types of notes. After the training is completed, the deep learning network model has the function of distinguishing the correspondence between the knocking sound signal and the music signal.

[0249] In step 602A, in response to the triggering operation, the microphone is turned on and a recording signal of the knocking object is collected.

[0250] For example, after the model training is completed, refer to Figure 5C , Figure 5C This is a schematic diagram of the third interface of the audio processing method provided in an embodiment of the present application; audio processing interface 501C includes a recording control 502C and a sound effect generation control 503C. In response to recording control 502C being triggered, audio acquisition processing begins; in response to sound effect generation control 503C being triggered, an audio signal generated based on the acquired audio signal is output.

[0251] For example, audio collection and processing begins, that is, the recording function of the terminal device is turned on. The recorded audio collected by the terminal device through the microphone includes the knocking sound of the object that the user previously bound to the corresponding note. For example: the user knocks on a specific object to make a corresponding sound. After multiple rounds of recording signals, the relationship between the knocking sound of the object and the bound pitch height is recorded, and the user is informed that the binding of this pitch height is completed. The user then continues to bind other pitch heights according to the above steps. After all custom pitch heights are bound, the setting interface (interface 501A or interface 501B) is exited.

[0252] In step 603A, the timing of the knocking sound signal in the recording signal is detected, and the knocking sound characteristics are obtained.

[0253] For example, after the user finishes recording and clicks the "Generate Sound Effect" control, the system will use the trained deep learning network to identify the various percussion sounds in the recorded audio, and record the parameters corresponding to each percussion sound (sounding time, sound intensity, duration). Based on the binding relationship between percussion sounds and pitch height, the system extracts the pitch height corresponding to each percussion sound. The percussion sound attribute parameters are configured for each pitch height to generate a sound signal, and these percussion sounds are replaced by corresponding sound effect files through the corresponding note replacement method. Users can then play or save the sound effect files.

[0254] For example, during the application process, the timbre and pitch are determined by the relationship between the knocking sound and the pitch height generated during the training process, the duration of the sound signal corresponding to the pitch height is determined by the duration of the knocking sound (positive correlation), and the loudness of the sound signal corresponding to the pitch height is determined by the volume of the knocking sound (positive correlation).

[0255] In some embodiments, the recorded signal is multiplexed, and in step 604A, the knocking sound signal in the recorded signal is eliminated.

[0256] For example, the method of eliminating the knocking sound signal can be based on the inverse Fourier transform of the deep learning network. The input of the deep learning network model is the power spectrum obtained by Fourier transform (FFT) of the audio signal. The output of the deep learning network model is two parts: the frequency domain suppression gain of the knocking sound signal and the probability of knocking sound appearing in the current input sound signal. The frequency domain suppression gain will be multiplied by the complex spectrum corresponding to the Fourier transform output of the input signal. The multiplication result data is subjected to the inverse Fourier transform (IFFT) to obtain the sound signal after the knocking sound is eliminated.

[0257] In an embodiment of the present application, the deep learning network model can be composed of different types of units, such as multiple convolutional units (conv) and gated units (GRU), fully connected units (FC), normalization layers (sigmoid), etc.

[0258] refer to Figure 7 , Figure 7 : This is a model structure diagram of the audio processing method provided in an embodiment of the present application; the deep learning network model 700 includes a first fully connected layer 701, a first convolutional layer 702, a second convolutional layer 703, a third convolutional layer 704, a fourth convolutional layer 705, a first gated loop unit 706, a second gated loop unit 707, a second fully connected layer 708, a third fully connected layer 709, and a normalization layer 710. The audio signal is converted into a Fourier transform power spectrum through Fourier transform, and the Fourier transform power spectrum is input into the deep learning network model 700, passes through the fully connected layer and multiple convolutional layers, and is output to different branches in the first gated loop unit 706. The third fully connected layer 709 and the normalization layer 710 classify the features of the Fourier transform power spectrum to obtain the probability that the percussion sound in the audio signal is mapped to the corresponding note. The second gate recurrent unit 707 and the second fully connected layer 708 output the audio signal with suppressed gain, which is merged with the original Fourier transform power spectrum. The merged result is processed by inverse fast Fourier transform to obtain a sound signal with the knocking sound eliminated.

[0259] In some embodiments, a pre-trained neural network model is saved to the local terminal device. When a user inputs a new tapping sound, a new sound data sample is uploaded to the server. The server analyzes the new sample for similarities with existing samples of different categories on the server, and then selects the optimal model and sends it to the local terminal device for use. If the new sample differs significantly from the existing samples on the server, online learning or incremental training is initiated to train the new sample, and the newly trained model is ultimately sent to the local terminal device.

[0260] In step 605A, the sound signal corresponding to the pitch is superimposed on the recorded signal after the knocking sound signal is eliminated.

[0261] For example, the sound signal of the knocking sound is identified and eliminated, and the corresponding pitch is replaced at the position where the knocking sound signal is eliminated. The replacement can be achieved by linearly superimposing the recording signal of the knocking sound eliminated and the sound signal of the pitch corresponding to the knocking sound.

[0262] For example, the sound signal of the pitch height corresponding to the knocking sound can be determined in the following way: according to the characteristics of the knocking sound (intensity, i.e. energy size, duration, spectral energy distribution), the sound signal of the pitch height is enhanced to obtain the configured sound signal.

[0263] For example: the intensity of the knocking sound corresponds to the volume of the pitch, different intensities correspond to different levels of volume output, the duration of the knocking sound is proportional to the output duration of the pitch, and the spectral energy distribution of the knocking sound corresponds to the equalization processing of the pitch, such as the different forms of distribution of high and low frequency energy correspond to different equalizer (EQ) adjustment curves to control the gain output of different frequency bands of the pitch.

[0264] In some embodiments, the recorded signal is not multiplexed, and in step 606A, a preset pitch is generated at a position corresponding to the timing of detecting the knocking sound.

[0265] For example, the solution of not multiplexing the recorded signal is suitable for real-time generation application scenarios. For example, if a user knocks on a table, a music sound signal corresponding to the knocking sound is immediately output.

[0266] For example, a pitch is generated at the location of a detected knock. Similarly, the output pitch is adjusted based on the characteristics of the knock. This method can be converted in real time and does not require generating a recording file, which can reduce memory usage.

[0267] In some embodiments, there is no need for the user to define the relationship between the pitch and the sound signal emitted by the object sound. Figure 6B , Figure 6BThis is a flow chart of the audio processing method provided in an embodiment of the present application; the following description will be given using the terminal device as an example in which the execution subject is used.

[0268] In step 601B, in response to the triggering operation, the microphone is turned on and a recording signal of the striking object is collected.

[0269] The principle of step 601B refers to step 602A above and will not be repeated here.

[0270] In step 602B, the timing of the knocking sound signal in the recording signal is detected, and the knocking sound characteristics are obtained.

[0271] The principle of step 602B refers to step 603A above and will not be repeated here.

[0272] In some embodiments, the recorded signal is multiplexed, and in step 603B, the knocking sound signal in the recorded signal is eliminated.

[0273] The principle of step 603B refers to step 604A above and will not be repeated here.

[0274] In step 604B, a pitch height and a corresponding sound signal are generated based on a preset rule, and the sound signal corresponding to the pitch height is superimposed on the recorded signal after the knocking sound signal is eliminated.

[0275] For example, the preset rules include the following types:

[0276] 1. Random selection. For example, there is no binding relationship between knocking sound and pitch height. Each time a knocking sound signal is collected, a pitch height is randomly selected to generate a sound signal.

[0277] 2. Determine the pitch based on the similarity between the knocking sound feature and the sound feature of the sound signal corresponding to the pitch.

[0278] 3. Each volume interval corresponds to a sound signal of a different pitch, and the sound signal of the corresponding pitch is determined according to the volume interval in which the volume of the knocking sound is located.

[0279] In some embodiments, the recorded signal is not multiplexed, and in step 605B, at the position corresponding to the timing of the detection of the knocking sound, the pitch is generated according to a preset rule.

[0280] For example, the preset rules have been introduced in step 604B above and will not be repeated here.

[0281] In some embodiments, the audio processing method provided by the embodiments of the present application can be applied to music games, refer to Figure 5D , Figure 5DSchematic diagram of the fourth interface of the audio processing method provided in an embodiment of the present application; when the audio processing method provided in an embodiment of the present application is applied to a music game, in interface 501D of the music game, the note icon descends at a uniform rate. Assuming the sound signal collected is the user's clapping, when the note descends to the dotted line corresponding to the tapping time point, the user claps, and the terminal device collects the clapping sound and outputs a drumming sound signal. The user taps an object according to the tapping timing indicated on the game screen to produce a sound. The terminal device identifies from the recording whether the user's tapping time point matches the tapping time point indicated by the game, and finally calculates the score.

[0282] In some embodiments, the audio processing method provided by the embodiments of the present application can be applied in third-person or first-person games, and used to control virtual objects to perform operations. For example: the user customizes the knocking of a certain object to correspond to the operation of a certain prop in the game, specifically: the user knocks on the desktop to make a corresponding sound similar to knocking on a wooden board. The user can bind the knocking sound to the virtual props in the multiplayer competitive game to shoot before starting the game. The knocking sound of the wooden board corresponds to shooting a bullet. For example, the sound of clapping will correspond to the use of virtual props, and clapping once corresponds to throwing an explosive virtual prop, etc. Therefore, recognition based on knocking sound can be used to control corresponding physical or virtual objects in different scenes, allowing users to interact through knocking actions.

[0283] For example, the following is an example of replacing virtual props (for example, appearance parts worn by virtual objects) with virtual objects. Figure 8 , Figure 8 This is a schematic diagram of the sixth interface of the audio processing method provided in an embodiment of the present application. Virtual scene 801 includes a virtual object 802 wearing a component 803. In response to a user's clapping sound signal, virtual object 802 wearing a component 805 is displayed in virtual scene 801, along with a prompt 804 stating "Appearance Changed." By capturing the user's clapping sound, the virtual object in the virtual scene is controlled to change its appearance, improving the efficiency of human-computer interaction in virtual scenes.

[0284] In an embodiment of the present application, after the neural network model is trained, the user can randomly tap the identified object at his or her own rhythm to make a sound, and the server or terminal device will generate corresponding sound effects and music in real time or non-real time. The corresponding tapping sound can also be removed from the recording and replaced with the corresponding sound effect or instrument to generate a synthesized musical work. In the generated music, the sound effect volume and pitch height are further controlled based on the user's tapping strength and the spectral distribution analysis of the tapping sound.

[0285] In addition, when users knock on objects at will, the system will recognize the knocking sounds and randomly generate sound effects of different styles to replace the knocking sounds to generate corresponding music works. Users can control the volume of the knocking sounds by different knocking forces, and then control the volume of the generated sound signal. By controlling the tone of the knocking sounds (for example, knocking on different positions of the same object and knocking at different angles to generate different sound effects), the tone of the generated music can be controlled. In this way, the user's entertainment experience is enhanced and the user is given more space for music creation.

[0286] The audio processing method of the embodiment of the present application can be completely based on recording signal recognition. The recording function is available in ordinary mobile phones or portable devices, and there is no need to connect the terminal device to other devices. Users can make specific sounds by knocking on different objects and customize corresponding different sound effects, scales, etc. The terminal device or server recognizes and records the knocking sounds and sets the sound effects and music that match them to create corresponding musical works. After obtaining a pure music work, the user can also merge it with the human voice to generate song audio. Alternatively, if the original recording itself carries the human voice, the knocking sound signal in the original recording signal is replaced with the sound signal of the music to obtain song audio, which increases the user's freedom in audio production.

[0287] The following continues to describe the exemplary structure of the audio processing device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2A As shown, the software modules stored in the audio processing device 455 of the memory 450 may include: a display module 4551, used to display an audio processing interface, wherein the audio processing interface includes a first recording control; an acquisition module 4552, used to acquire a first audio signal in response to a trigger operation for the first recording control; an output module 4553, used to output a second audio signal in response to the first audio signal including multiple volume jump points, wherein the second audio signal is generated based on the pitch heights corresponding to the multiple first sub-audio signals, and the multiple first sub-audio signals are segmented from the first audio signal based on the multiple volume jump points.

[0288] In some embodiments, the output module 4553 is used to obtain the pitch heights corresponding to multiple first sub-audio signals before outputting the second audio signal; generate multiple second sub-audio signals based on the pitch heights corresponding to the multiple first sub-audio signals, wherein the pitch height corresponding to a first sub-audio signal is used to generate a second sub-audio signal, the first sub-audio signal and the corresponding second sub-audio signal have the same or different durations, and the timing of the multiple second sub-audio signals is the same as the timing of the first sub-audio signals corresponding to the multiple second sub-audio signals; and synthesize the second audio signal based on the multiple second sub-audio signals according to the timing of the multiple second sub-audio signals.

[0289] In some embodiments, the output module 4553 is configured to synthesize a transition audio signal corresponding to each two second sub-audio signals based on each two second sub-audio signals that are adjacent in time sequence; arrange each second sub-audio signal according to the time sequence of the plurality of second sub-audio signals, and insert a corresponding transition audio signal between each two second sub-audio signals that are adjacent in time sequence to form a second audio signal.

[0290] In some embodiments, the output module 4553 is used to perform the following processing on every two second sub-audio signals that are adjacent in time sequence: take the second sub-audio signal that is earlier in time sequence of the two second sub-audio signals that are adjacent in time sequence as the first signal to be processed, and take the second sub-audio signal that is later in time sequence as the second signal to be processed; obtain the tail segment of the first signal to be processed and the head segment of the second signal to be processed; superimpose the tail segment and the head segment to obtain a superimposed signal; perform volume adjustment processing on the superimposed signal to obtain a transition audio signal corresponding to the two second sub-audio signals that are adjacent in time sequence, wherein the volume curve of the transition audio signal is characterized by gradually decreasing from an initial volume and then gradually rising to the initial volume.

[0291] In some embodiments, the output module 4553 is configured to, before outputting the second audio signal,

[0292] Obtaining pitch heights corresponding to each of the plurality of first sub-audio signals; generating a plurality of second sub-audio signals based on the pitch heights corresponding to each of the plurality of first sub-audio signals, wherein the pitch height corresponding to one first sub-audio signal is used to generate one second sub-audio signal; and replacing each of the plurality of first sub-audio signals in the first audio signal with the second sub-audio signals corresponding to the first sub-audio signal to form a second audio signal.

[0293] In some embodiments, the output module 4553 is used to perform the following processing for each first sub-audio signal: obtain a first duration of the first sub-audio signal; determine a sound source type pre-associated with the pitch; determine a second duration of the second sub-audio signal based on the first duration of the first sub-audio signal; and generate an audio signal of a second duration based on the second duration, the sound source type, and the pitch as the second sub-audio signal.

[0294] In some embodiments, the sound source types corresponding to the pitch heights include: different types of musical instruments and preconfigured sound effects.

[0295] In some embodiments, the pitch heights corresponding to the multiple first sub-audio signals are obtained through a pre-trained neural network model, and the audio processing interface includes a second recording control, a third recording control and a model training control; the acquisition module 4552 is used to obtain the sample pitch height associated with the sample reference signal to be recorded before collecting the first audio signal in response to the trigger operation for the first recording control; in response to the first trigger operation for the second recording control, start audio acquisition processing and display the third recording control; in response to the second trigger operation for the third recording control, use the audio signal collected between the interval time of the first trigger operation and the second trigger operation as the third audio signal; use the third audio signal as the sample reference signal corresponding to the sample pitch height; in response to the trigger operation for the model training control, combine the sample reference signal and the sample pitch height into a sample pair, and train the initialized neural network model based on the sample pair to obtain a pre-trained neural network model, wherein the pre-trained neural network model is used to determine the pitch heights corresponding to the multiple first sub-audio signals.

[0296] In some embodiments, the acquisition module 4552 is used to call the neural network model based on each first sub-audio signal and perform the following processing: perform feature extraction processing on the first sub-audio signal to obtain time domain features and frequency domain features; perform prediction processing based on the time domain features and frequency domain features corresponding to the first sub-audio signal to obtain the predicted probability of mapping the first sub-audio signal to different pitch heights; and use the pitch height with the maximum predicted probability as the pitch height corresponding to the first sub-audio signal.

[0297] In some embodiments, the audio processing interface also includes: a first option list, the first option list including first options corresponding to multiple different sound source types; an acquisition module 4552, for responding to a first selection operation for the first option, using the sound source type corresponding to the first option selected in the first selection operation as a sample sound source type; displaying at least one pitch height associated with the sample sound source type; and responding to a second selection operation for the pitch height, using the pitch height selected in the second selection operation as the sample pitch height of the sample reference signal.

[0298] In some embodiments, the neural network model includes: a feature extraction layer and a feature classification layer; the feature extraction layer is used to perform feature extraction processing on the first audio signal, and the feature classification layer is used to perform prediction processing based on the result of the feature extraction processing; the acquisition module 4552 is used to call the neural network model based on the sample reference signal to perform feature extraction processing, and obtain the sample time domain features and sample frequency domain features of the sample reference signal; perform prediction processing based on the sample time domain features and the sample frequency domain features to obtain the predicted probability of mapping the sample reference signal to different pitch heights; use the pitch height with the maximum predicted probability as the predicted pitch height corresponding to the sample reference signal; determine the loss function of the neural network model based on the difference between the sample pitch height and the predicted pitch height; update the parameters of the neural network model based on the loss function to obtain a pre-trained neural network model.

[0299] In some embodiments, the acquisition module 4552 is used to, after acquiring the first audio signal, in response to the first audio signal including at least one volume jump point and the first sub-audio signal included in the first audio signal not belonging to the learned audio type, use the first audio signal as a new sample signal, wherein the learned audio type is a sample sound source type that has been learned by the neural network model; and train the neural network model based on the new sample signal to obtain an updated neural network model.

[0300] In some embodiments, the output module 4553 is used to obtain a mapping relationship table between pitch heights and label values, wherein each pitch height in the mapping relationship table corresponds one-to-one to each label value; the following processing is performed for each first sub-audio signal: random number generation processing is performed based on the first sub-audio signal to obtain a random value; in response to the first value being equal to the label value, the pitch height corresponding to the same label value is used as the pitch height corresponding to the first sub-audio signal.

[0301] In some embodiments, the acquisition module 4552 is used to start audio acquisition processing in response to a third trigger operation on the first recording control; and in response to a fourth trigger operation on the first recording control, use the audio signal collected between the interval time of the third trigger operation and the fourth trigger operation as the first audio signal.

[0302] In some embodiments, the display module 4551 is used to display a virtual scene before outputting the second audio signal, wherein the virtual scene includes a virtual object; the output module 4553 is used to display a screen of the virtual object performing a preset operation in response to the volume of the first sub-audio signal being greater than a volume threshold, and the first sub-audio signal and the reference sound signal being sound signals emitted by the same type of objects when outputting the second audio signal; wherein the reference sound signal is pre-configured, and the types of preset operations include: moving operations of virtual objects; interactive operations between virtual objects and virtual props; and interactive operations between virtual objects and other virtual objects.

[0303] In some embodiments, as Figure 2A As shown, the software modules stored in the audio processing device 455 of the memory 450 may include: an acquisition module 4552, used to obtain a first audio signal; an acquisition module 4552, used to perform volume recognition processing on the first audio signal to obtain multiple volume jump points; an acquisition module 4552, used to separate multiple first sub-audio signals from the first audio signal based on the multiple volume jump points; an output module 4553, used to obtain the pitch heights corresponding to the multiple first sub-audio signals respectively; and an output module 4553, used to generate a second audio signal based on each pitch height.

[0304] In some embodiments, as Figure 2A As shown, the software modules stored in the audio processing device 455 of the memory 450 may include: a display module 4551, used to display an audio processing interface, wherein the audio processing interface includes an audio signal list, including at least one pre-recorded audio signal; an acquisition module 4552, used to respond to a selection operation on the audio signal list and use the selected audio signal as the first audio signal.

[0305] The output module 4553 is used to output a second audio signal in response to the first audio signal including multiple volume jump points, wherein the second audio signal is generated based on the pitch heights corresponding to the multiple first sub-audio signals, and the multiple first sub-audio signals are separated from the first audio signal based on the multiple volume jump points.

[0306] The present invention provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform the audio processing method described in the present invention.

[0307] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the audio processing method provided in the embodiment of the present application, for example, Figure 3A The audio processing method shown.

[0308] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0309] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0310] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0311] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0312] To sum up, by performing audio processing based on the sub-audio signal with jumps in the collected first audio signal in the embodiment of the present application, the accuracy of the generated audio signal is improved and the computing resources required for audio processing are saved compared to the solution of performing full processing based on the collected signal; the second audio signal is generated based on the pitch heights corresponding to the first sub-audio signals, so that the first audio signal and the second audio signal can cross the sound source type, thereby improving the richness of the generated audio signal.

[0313] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An audio processing method, characterized in that: The method comprises: Displaying an audio processing interface, wherein the audio processing interface includes a first recording control; In response to a trigger operation on the first recording control, collecting a first audio signal; In response to the first audio signal including multiple volume jump points, a second audio signal is output, wherein the second audio signal is generated based on the pitch heights corresponding to multiple first sub-audio signals, and the multiple first sub-audio signals are separated from the first audio signal based on the multiple volume jump points.

2. The method according to claim 1, characterized in that Before outputting the second audio signal, the method further includes: Obtaining pitch heights corresponding to the plurality of first sub-audio signals respectively; generating a plurality of second sub audio signals based on pitches corresponding to the plurality of first sub audio signals, wherein the pitch corresponding to one of the first sub audio signals is used to generate one of the second sub audio signals, the first sub audio signal and the corresponding second sub audio signal have the same or different durations, and the timing of the plurality of second sub audio signals is the same as the timing of the first sub audio signals corresponding to the plurality of second sub audio signals; A second audio signal is synthesized based on the plurality of second sub audio signals according to a time sequence of the plurality of second sub audio signals.

3. The method according to claim 2, characterized in that The synthesizing the second audio signal based on the plurality of second sub-audio signals according to the time sequence of the plurality of second sub-audio signals includes: synthesizing, based on every two second sub-audio signals that are adjacent in time sequence, transition audio signals corresponding to every two second sub-audio signals; Arrange each of the second sub audio signals according to the time sequence of the plurality of second sub audio signals, and insert a corresponding transition audio signal between every two second sub audio signals adjacent in time sequence to form the second audio signal.

4. The method according to claim 3, characterized in that The synthesizing, based on every two second sub-audio signals that are adjacent in time sequence, the transition audio signal corresponding to every two second sub-audio signals includes: The following processing is performed on every two second sub-audio signals that are adjacent in time sequence: of the two second sub audio signals adjacent in time sequence, the second sub audio signal with the earlier one in time sequence is used as the first signal to be processed, and the second sub audio signal with the later one in time sequence is used as the second signal to be processed; Obtaining a tail segment of the first signal to be processed and a head segment of the second signal to be processed; performing superposition processing on the tail segment and the head segment to obtain a superposition signal; Volume adjustment processing is performed on the superimposed signal to obtain the transition audio signal corresponding to two time-adjacent second sub-audio signals, wherein a volume curve of the transition audio signal is characterized by gradually decreasing from an initial volume and then gradually increasing to the initial volume.

5. The method according to any one of claims 1 to 4, characterized in that Before outputting the second audio signal, the method further includes: Obtaining pitch heights corresponding to the plurality of first sub-audio signals respectively; generating a plurality of second sub audio signals based on the pitch heights respectively corresponding to the plurality of first sub audio signals, wherein the pitch height corresponding to one first sub audio signal is used to generate one second sub audio signal; The multiple first sub audio signals in the first audio signal are respectively replaced with the second sub audio signals corresponding to the first sub audio signals to form a second audio signal.

6. The method according to claim 2 or 5, characterized in that: Generating a plurality of second sub audio signals based on the pitch heights respectively corresponding to the plurality of first sub audio signals includes: The following processing is performed on each of the first sub audio signals: Obtaining a first duration of the first sub-audio signal; determining a sound source type pre-associated with the pitch of the sound; determining a second duration of the second sub audio signal based on a first duration of the first sub audio signal; An audio signal with a second duration is generated based on the second duration, the sound source type, and the pitch as the second sub audio signal.

7. The method according to claim 6, characterized in that The sound source types corresponding to the pitch height include: different types of musical instruments and pre-configured sound effects.

8. The method according to claim 2 or 5, characterized in that The pitch heights respectively corresponding to the multiple first sub-audio signals are obtained through a pre-trained neural network model, and the audio processing interface includes a second recording control, a third recording control, and a model training control; Before collecting the first audio signal in response to the triggering operation on the first recording control, the method further includes: Obtaining a sample pitch associated with a sample reference signal to be recorded; In response to a first trigger operation on the second recording control, starting audio acquisition processing and displaying a third recording control; In response to a second trigger operation on the third recording control, using an audio signal collected between an interval between the first trigger operation and the second trigger operation as a third audio signal; using the third audio signal as a sample reference signal corresponding to the sample pitch; In response to a trigger operation on the model training control, the sample reference signal and the sample pitch height are combined into a sample pair, and an initialized neural network model is trained based on the sample pair to obtain the pre-trained neural network model, wherein the pre-trained neural network model is used to determine the pitch heights corresponding to the multiple first sub-audio signals respectively.

9. The method according to claim 8, characterized in that The obtaining of the pitch heights respectively corresponding to the plurality of first sub-audio signals includes: The neural network model is called based on each of the first sub-audio signals to perform the following processing: Performing feature extraction processing on the first sub-audio signal to obtain time domain features and frequency domain features; performing prediction processing based on the time domain features and the frequency domain features corresponding to the first sub audio signal to obtain predicted probabilities of the first sub audio signal being mapped to different pitches; The pitch of the tone with the maximum prediction probability is used as the pitch corresponding to the first sub-audio signal.

10. The method according to claim 8, characterized in that The audio processing interface further includes: a first option list, the first option list including first options corresponding to a plurality of different sound source types; The obtaining of the sample pitch associated with the sample reference signal to be recorded comprises: In response to a first selection operation for the first option, taking the sound source type corresponding to the first option selected by the first selection operation as a sample sound source type; Displaying at least one pitch associated with the sample sound source type; In response to a second selection operation on the pitch height, the pitch height selected by the second selection operation is used as the sample pitch height of the sample reference signal.

11. The method according to claim 8, characterized in that The neural network model includes: a feature extraction layer and a feature classification layer; the feature extraction layer is used to perform feature extraction processing on the first audio signal, and the feature classification layer is used to perform prediction processing based on the result of the feature extraction processing; The training of the initialized neural network model based on the sample pair to obtain the pre-trained neural network model includes: Calling the neural network model to perform feature extraction processing based on the sample reference signal to obtain sample time domain features and sample frequency domain features of the sample reference signal; Performing prediction processing based on the sample time domain features and the sample frequency domain features to obtain prediction probabilities of the sample reference signal being mapped to different pitch heights; Using the pitch height with the maximum predicted probability as the predicted pitch height corresponding to the sample reference signal; determining a loss function for the neural network model based on a difference between the sample pitch and the predicted pitch; The parameters of the neural network model are updated based on the loss function to obtain the pre-trained neural network model.

12. The method according to claim 8, characterized in that After collecting the first audio signal, the method further includes: In response to the first audio signal including at least one volume jump point and the first sub-audio signal included in the first audio signal not belonging to a learned audio type, using the first audio signal as a new sample signal, wherein the learned audio type is a sample sound source type that has been learned by the neural network model; The neural network model is trained based on the new sample signal to obtain an updated neural network model.

13. The method according to claim 2 or 5, characterized in that: The obtaining of the pitch heights respectively corresponding to the plurality of first sub-audio signals includes: Obtaining a mapping relationship table between pitch heights and label values, wherein each pitch height in the mapping relationship table corresponds to each label value in a one-to-one manner; The following processing is performed on each of the first sub audio signals: performing random number generation processing based on the first sub-audio signal to obtain a random value; In response to the first value being equal to the label value, the pitch corresponding to the same label value is used as the pitch corresponding to the first sub audio signal.

14. The method according to any one of claims 1 to 13, characterized in that The collecting of the first audio signal in response to the triggering operation on the first recording control includes: In response to a third trigger operation on the first recording control, starting audio acquisition processing; In response to a fourth trigger operation on the first recording control, an audio signal collected between an interval between the third trigger operation and the fourth trigger operation is used as a first audio signal.

15. The method according to any one of claims 1 to 14, characterized in that Before outputting the second audio signal, the method further includes: displaying a virtual scene, wherein the virtual scene includes a virtual object; When outputting the second audio signal, the method further includes: In response to a volume of the first sub audio signal being greater than a volume threshold, and the first sub audio signal and the reference sound signal being sound signals emitted by the same type of object, displaying a picture of the virtual object performing a preset operation; The reference sound signal is pre-configured, and the types of preset operations include: movement operations of the virtual object; interaction operations between the virtual object and virtual props; and interaction operations between the virtual object and other virtual objects.

16. The method according to any one of claims 1 to 15, characterized in that After collecting the first audio signal in response to the triggering operation on the first recording control, the method further includes: determining a time-domain volume curve of the first audio signal, wherein different points in the time-domain volume curve represent volume values ​​of the first audio signal at different moments; Extracting a point that satisfies a volume jump condition from the time-domain volume curve as a volume jump point, wherein the volume jump condition includes: a volume value corresponding to the point is greater than a volume threshold, and an absolute value of a differential value of a volume corresponding to the point in the time-domain volume curve is greater than that of at least one adjacent point; In response to the volume between two temporally adjacent volume jump points being greater than the volume threshold, the first sub-audio signal is obtained by segmenting the first audio signal according to the moments corresponding to the two temporally adjacent volume jump points.

17. An audio processing method, characterized in that: The method comprises: displaying an audio processing interface, wherein the audio processing interface includes a list of audio signals, including at least one pre-recorded audio signal; In response to a selection operation on the audio signal list, taking the selected audio signal as the first audio signal; In response to the first audio signal including multiple volume jump points, a second audio signal is output, wherein the second audio signal is generated based on the pitch heights corresponding to multiple first sub-audio signals, and the multiple first sub-audio signals are separated from the first audio signal based on the multiple volume jump points.

18. An audio processing method, characterized in that: The method comprises: Acquire a first audio signal; performing volume recognition processing on the first audio signal to obtain a plurality of volume jump points; dividing the first audio signal into a plurality of first sub-audio signals based on the plurality of volume jump points; Obtaining pitch heights corresponding to the plurality of first sub-audio signals respectively; A second audio signal is generated based on each of the pitches.

19. An audio processing device, characterized in that: The device comprises: A display module, configured to display an audio processing interface, wherein the audio processing interface includes a first recording control; an acquisition module, configured to acquire a first audio signal in response to a trigger operation on the first recording control; an output module, configured to output a second audio signal in response to the first audio signal including a plurality of volume jump points, wherein the second audio signal is generated based on pitches respectively corresponding to a plurality of first sub-audio signals, and the plurality of first sub-audio signals are segmented from the first audio signal based on the plurality of volume jump points.

20. An audio processing device, characterized in that: The device comprises: An acquisition module, configured to acquire a first audio signal; an acquisition module, configured to perform volume recognition processing on the first audio signal to obtain a plurality of volume jump points; an acquisition module, configured to separate the first audio signal into a plurality of first sub-audio signals based on the plurality of volume jump points; an output module, configured to obtain pitch heights corresponding to the plurality of first sub-audio signals; An output module is configured to generate a second audio signal based on each of the pitches.

21. An audio processing device, characterized in that: The device comprises: a display module, configured to display an audio processing interface, wherein the audio processing interface includes a list of audio signals, including at least one pre-recorded audio signal; a collection module, configured to, in response to a selection operation on the audio signal list, take the selected audio signal as the first audio signal; an output module, configured to output a second audio signal in response to the first audio signal including a plurality of volume jump points, wherein the second audio signal is generated based on pitches respectively corresponding to a plurality of first sub-audio signals, and the plurality of first sub-audio signals are segmented from the first audio signal based on the plurality of volume jump points.

22. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the audio processing method according to any one of claims 1 to 16, or claim 17 or 18 when executing the computer-executable instructions or computer program stored in the memory.

23. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the audio processing method according to any one of claims 1 to 16, or claim 17 or 18 is implemented.

24. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the audio processing method according to any one of claims 1 to 16, or claim 17 or 18 is implemented.