Audio information processing method, audio information presentation method and device
By analyzing and dynamically modifying the template information of the audio service processing environment, the problem that audio information processing cannot be modified in real time and controlled in the prior art is solved, and the simple processing of audio information is achieved.
Patent Information
- Application Number
- CN202110513496.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-11
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-05-11
AI Technical Summary
In the prior art, the audio information processing process cannot achieve real-time modification and flexible control, resulting in cumbersome user operations.
By analyzing the template information of the audio service processing environment, the first audio track data is obtained and saved in the audio information storage hash table, dynamically modify it in response to the dynamic modification instruction, an audio frame data reader is configured, and an audio frame corresponding to the target timestamp is extracted for combination processing in response to the audio service output instruction.
Real-time modification and flexible control of audio information are realized, simplifying the audio information processing process and improving convenience.
Smart Images

Figure CN115329122B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to audio information processing technology, and particularly to an audio information processing method, an audio information presentation method, a device, an electronic device, and a storage medium. Background Art
[0002] In the related art, the forms of audio information are diverse, and the demand for audio information processing has shown an explosive growth. However, due to the complex storage process of audio information, when processing audio information, it is impossible to modify and flexibly control the audio information in real time, making the audio information processing process of users more cumbersome. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide an audio information processing method, an audio information presentation method, a device, an electronic device, and a storage medium, which can achieve real-time modification and flexible control of audio information, making the audio information processing process in response to an audio service output instruction more convenient and improving the convenience of audio information processing.
[0004] The technical solution of the embodiments of the present invention is implemented as follows:
[0005] Embodiments of the present invention provide an audio information processing method, including:
[0006] Analyzing and processing template information of an audio service processing environment to obtain first audio track data;
[0007] Saving the first audio track data into an audio information storage hash table, where the audio information storage hash table is used to save audio information, and the first audio track data is stored in the audio information;
[0008] In response to a dynamic modification instruction, dynamically modifying the first audio track data to obtain second audio track data;
[0009] In response to an audio service output instruction, obtaining audio information from the audio information storage hash table, and configuring a corresponding audio frame data reader for the audio information;
[0010] Extracting a first audio frame in the second audio track data stored in the audio information corresponding to a target timestamp through the audio frame data reader;
[0011] Combining and processing first audio frames corresponding to different target timestamps to obtain and output a second audio frame, so as to respond to the audio service output instruction through the second audio frame.
[0012] Embodiments of the present invention further provide an audio information presentation method, including:
[0013] Display a user interface and present a task control function item in the user interface, where the task control function item is used to dynamically modify the first audio track data through a condition-triggered dynamic modification instruction;
[0014] In response to a trigger operation on the task control function item, obtain animation effect information including an audio service output instruction;
[0015] Obtain a second audio frame corresponding to the audio service output instruction;
[0016] Present the animation effect information and the second audio frame in the user interface.
[0017] An embodiment of the present invention further provides an audio information processing device, and the device includes:
[0018] A first information transmission module, configured to parse and process template information of an audio service processing environment to obtain first audio track data;
[0019] A first information processing module, configured to save the first audio track data to an audio information storage hash table, where the audio information storage hash table is used to save audio information, and the audio information is used to store the first audio track data;
[0020] The first information processing module is configured to, in response to a dynamic modification instruction, dynamically modify the first audio track data to obtain second audio track data;
[0021] The first information processing module is configured to, in response to an audio service output instruction, obtain audio information from the audio information storage hash table and configure a corresponding audio frame data reader for the audio information;
[0022] The first information processing module is configured to, through the audio frame data reader, extract a first audio frame in the second audio track data stored in the audio information corresponding to a target timestamp;
[0023] The first information processing module is configured to perform a combination process on first audio frames corresponding to different target timestamps to obtain and output a second audio frame, so as to respond to the audio service output instruction through the second audio frame.
[0024] In the above solution,
[0025] The first information processing module is configured to parse the template information of the audio service processing environment to obtain the timing information of the template information;
[0026] The first information processing module is configured to parse the audio parameters corresponding to the template information according to the timing information of the template information, and obtain the audio type and track information parameters corresponding to the template information;
[0027] The first information processing module is configured to extract the template information based on the audio type and track information parameters corresponding to the template information, so as to obtain the first audio track data corresponding to the template information.
[0028] In the above solution,
[0029] The first information processing module is configured to extract audio data from the audio resource component of the template information and construct the first audio information when the audio type is a single audio;
[0030] The first information processing module is configured to extract audio data from the multimedia information component of the template information and construct the second audio information when the audio type is an audio matching the video information;
[0031] The first information processing module is configured to extract audio data from the animation resource component of the template information and construct the third audio information when the audio type is an audio matching the animation resource;
[0032] The first information processing module is configured to combine the first audio information, the second audio information, and the third audio information to obtain the first audio track data corresponding to the template information.
[0033] In the above solution,
[0034] The first information processing module is configured to receive the dynamic modification instruction, where the dynamic modification instruction includes at least one of the following:
[0035] Adjusting the start position of playing the audio track data, pausing the playing, resuming the playing, addressing the playing, conditional triggering, and scripted playing;
[0036] The first information processing module is configured to dynamically modify the first audio track data in the audio information storage hash table according to the type of the dynamic modification instruction, so as to obtain the second audio track data;
[0037] The first information processing module is configured to modify the audio information storage hash table in response to the dynamic modification instruction, so as to obtain the audio information storage hash table corresponding to the second audio track data.
[0038] In the above solution,
[0039] The first information processing module is configured to, when first responding to an audio service output instruction and obtaining audio information from the audio information storage hash table, configure a corresponding first audio frame data reader for the audio information;
[0040] The first information processing module is configured to detect the persistence status of the audio information. When the audio information persists, maintain the persistence status of the first audio frame data reader and update the data information in the audio information;
[0041] The first information processing module is configured to, when the audio information is removed, delete the first audio frame data reader and configure a second audio frame data reader according to the change of the audio information.
[0042] In the above solution,
[0043] The first information processing module is configured to, when the dynamic modification instruction is addressing playback or the playback start position is adjusted to the start position, determine a target time parameter matching the dynamic modification instruction and save the target time parameter in the audio information storage hash table;
[0044] The first information processing module is configured to, when outputting the second audio frame, perform a comparison process on the target time parameter and the target timestamp to determine a timestamp comparison result;
[0045] The first information processing module is configured to, based on the timestamp comparison result, trigger an addressing playback process to ensure that the target time parameter and the target timestamp are synchronized when the second audio frame is output.
[0046] In the above solution,
[0047] The first information processing module is configured to, when parsing the template information of the audio service processing environment and not obtaining the first audio track data, output an empty data frame matching the target timestamp as the second audio frame.
[0048] In the above solution,
[0049] The first information processing module is configured to, when the dynamic modification instruction is to adjust the playback rate of an audio frame, adjust the first audio track data in the audio information through the audio frame data reader to obtain an audio frame playback rate matching the dynamic modification instruction;
[0050] The first information processing module is configured to, when the dynamic modification instruction is to adjust the volume of an audio frame, adjust the first audio track data in the audio information through the audio frame data reader to obtain the volume of the audio frame that matches the dynamic modification instruction.
[0051] An embodiment of the present invention further provides an audio information presentation device, which includes:
[0052] The second information processing module is configured to display a user interface and present a task control function item in the user interface, where the task control function item is used to dynamically modify the first audio track data through a condition-triggered dynamic modification instruction;
[0053] The second information processing module is configured to, in response to a trigger operation on the task control function item, obtain animation effect information including an audio service output instruction;
[0054] The second information processing module is configured to obtain a second audio frame corresponding to the audio service output instruction;
[0055] The second information processing module is configured to present the animation effect information and the second audio frame in the user interface.
[0056] In the above solution,
[0057] The second information processing module is configured to display a user interface and present a task control function item in the user interface, where the task control function item is used to dynamically modify the first audio track data through an addressing playback dynamic modification instruction;
[0058] The second information processing module is configured to, in response to a trigger operation on the task control function item, obtain lyric effect information including an audio service output instruction;
[0059] The second information processing module is configured to obtain a second audio frame corresponding to the audio service output instruction;
[0060] The second information processing module is configured to present the lyric effect information and the second audio frame in the user interface.
[0061] In the above solution,
[0062] The second information processing module is configured to, in response to a viewing operation on the task control function item, present a content page including template information of the audio service processing environment and present at least one interaction function item in the content page, where the interaction function item is used to implement interaction with the audio service processing environment;
[0063] The second information processing module is configured to receive an interaction operation for the audio service processing environment triggered based on the interaction function item, so as to execute a corresponding interaction instruction.
[0064] In the above solution,
[0065] The second information processing module is configured to present first interaction prompt information in the content page, and the first interaction prompt information is used to prompt that the interaction content corresponding to the interaction operation can be presented in the user interface;
[0066] The second information processing module is configured to switch the content page to the user interface in response to an operation of switching to the user interface.
[0067] In the above solution,
[0068] The second information processing module is configured to present second interaction prompt information in the content page, and the second interaction prompt information is used to prompt that the interaction content corresponding to the interaction operation can be presented in the special effect information template library interface;
[0069] The second information processing module is configured to switch the content page to the special effect information template library interface in response to an instruction to switch to the special effect information template library interface.
[0070] In the above solution,
[0071] The second information processing module is configured to present a sharing function item for sharing the special effect information in the user interface;
[0072] The second information processing module is configured to share the special effect information with users in different audio service processing environments in response to a trigger operation on the sharing function item for the special effect information.
[0073] An embodiment of the present invention further provides an electronic device, and the electronic device includes:
[0074] A memory for storing executable instructions;
[0075] A processor, when running the executable instructions stored in the memory, implements the foregoing audio information processing method or implements the foregoing audio information presentation method.
[0076] An embodiment of the present invention further provides a computer-readable storage medium, storing executable instructions, and when the executable instructions are executed by a processor, the foregoing audio information processing method is implemented or the foregoing audio information presentation method is implemented.
[0077] The embodiment of the present invention has the following beneficial effects:
[0078] In an embodiment of the present invention, template information of an audio service processing environment is parsed to obtain first audio track data; the first audio track data is saved to an audio information storage hash table, where the audio information storage hash table is used to save audio information, and the first audio track data is stored in the audio information; in response to a dynamic modification instruction, the first audio track data is dynamically modified to obtain second audio track data; in response to an audio service output instruction, audio information is obtained from the audio information storage hash table, and a corresponding audio frame data reader is configured for the audio information; through the audio frame data reader, a first audio frame in the second audio track data stored in the audio information corresponding to a target timestamp is extracted; the first audio frames corresponding to different target timestamps are combined to obtain and output a second audio frame, so as to respond to the audio service output instruction through the second audio frame; thus, real-time modification and flexible control of audio information can be achieved, making the audio information processing process in response to an audio service output instruction more convenient and improving the convenience of audio information processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 FIG. is a schematic diagram of a usage environment of an audio information processing method provided by an embodiment of the present invention;
[0080] Figure 2 FIG. is a schematic diagram of the composition structure of an electronic device provided by an embodiment of the present invention;
[0081] Figure 3 FIG. is a schematic diagram of the structure of an architecture for organizing logic and data in an embodiment of the present invention;
[0082] Figure 4 FIG. is an optional flowchart of an audio information processing method provided by an embodiment of the present invention;
[0083] Figure 5 FIG. is a schematic diagram of the processing of audio track data by a dynamic modification instruction in an embodiment of the present invention;
[0084] Figure 6 FIG. is an optional flowchart of an audio information processing method provided by an embodiment of the present invention;
[0085] Figure 7 FIG. is an optional flowchart of an audio information processing method provided by an embodiment of the present invention;
[0086] Figure 8 FIG. is a schematic diagram of animation effect information in an embodiment of the present invention;
[0087] Figure 9 FIG. is a schematic diagram of animation effect information in an embodiment of the present invention. DETAILED DESCRIPTION
[0088] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0089] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0090] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.
[0091] 1) In response to: used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or can have a set delay; without special instructions, there is no restriction on the execution order of the multiple executed operations.
[0092] 2) Target video: various forms of video information available on the Internet, such as video files, audio information, etc. presented in a client or a smart device.
[0093] 3) Client: a carrier for implementing specific functions in a terminal. For example, a mobile client (APP) is a carrier for specific functions in a mobile terminal, such as the function of performing an online live broadcast (video streaming) or the function of playing an online video.
[0094] 4) Audio information: including but not limited to: background music in long videos (video works), music in short videos (videos uploaded by users with a length less than 1 minute), audio (such as a mv with fixed pictures or a record), and sound effects in animations.
[0095] 5) ECS: Entity Component System, an architecture pattern for organizing logic and data, where a component (Component) carries the data required for a system to run, a System is a logical module for processing the data carried on the Component, and an Entity is a representative of an entity. An entity uses several Components to carry data.
[0096] 6) Audio Process System: an audio processing system in a rendering engine implemented based on ECS.
[0097] 7) AudioInfoMap: A hash table for storing audio information, used to store the structures of multiple AudioInfos.
[0098] 8) AudioReader: A reader used to read the audio frame data in an AudioInfo.
[0099] 9) AudioInfo: That is, audio information, representing the structure of an audio data source.
[0100] Figure 1 It is a schematic diagram of the usage scenario of the audio information processing method provided by the embodiments of the present invention. Refer to Figure 1 , different corresponding clients capable of performing different functions are set on the terminals (including terminal 10-1 and terminal 10-2). Among them, the corresponding clients for the terminals (including terminal 10-1 and terminal 10-2) obtain different video information from the corresponding servers 200 through different service processes via the network 300 for browsing. The terminals are connected to the servers 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and uses a wireless link to implement data transmission. Among them, the types of audio information obtained by the terminals (including terminal 10-1 and terminal 10-2) from the corresponding servers 200 are not the same. The audio information includes, but is not limited to: long videos (such as video works), short videos (videos uploaded by users with a length less than 1 minute), audio (such as an mv with a fixed picture or a record), and audio information corresponding to the sound effects in animations. For example, the terminals (including terminal 10-1 and terminal 10-2) can obtain long videos (that is, videos carrying video information or corresponding video links) from the corresponding servers 200 through the network 300, or can also obtain short videos from the corresponding servers 400 through the network 300 for browsing. Different types of videos can be stored in the servers 200 and the servers 400. Among them, in this application, the playback environments of different types of audio information are not distinguished. In the foregoing audio service processing environment, it is necessary to process and output the audio information according to different service requirements and user usage requirements.
[0101] Taking short videos as an example, the audio information processing method provided by the present invention can be applied to short video playback. In the production and playback of short videos, audio information from different data sources can be spliced or addressed for playback processing to generate audio information in an animation effect that meets user requirements. For example, template information in an audio service processing environment can be parsed to obtain first audio track data; the first audio track data is saved to an audio information storage hash table, where the audio information storage hash table is used to save audio information and the first audio track data is stored in the audio information; in response to a dynamic modification instruction, the first audio track data is dynamically modified to obtain second audio track data; in response to an audio service output instruction, audio information is retrieved from the audio information storage hash table, and a corresponding audio frame data reader is configured for the audio information; through the audio frame data reader, the first audio frame in the second audio track data stored in the audio information corresponding to a target timestamp is extracted; the first audio frames corresponding to different target timestamps are combined to obtain and output a second audio frame, so as to respond to the audio service output instruction through the second audio frame.
[0102] Finally, an audio frame matching the corresponding audio service output instruction is presented on the user interface UI (User Interface). The audio frame obtained in this process that matches the audio service output instruction can also be called by other application programs (for example, recommended to contacts in a short video client to generate corresponding audio frames using the same animation effect). Of course, the audio information processing method provided in this application can also be applied to other audio service processing environments, such as web audio information processing processes and animation effect production applets.
[0103] Due to the increasing demand for audio information processing, the audio information processing method provided in the embodiments of this application can be implemented through cloud technology. Among them, the embodiments of the present invention can be implemented in combination with cloud technology or blockchain network technology. Cloud technology (Cloud technology) refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data calculation, storage, processing, and sharing. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. The background services of technical network systems require a large amount of computing and storage resources, such as multimedia information websites, picture websites, and more portal websites. Therefore, cloud technology needs to be supported by cloud computing.
[0104] It should be noted that cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely expandable and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a provider of the basic capabilities of cloud computing, a cloud computing resource pool platform, referred to as a cloud platform, is generally called Infrastructure as a Service (IaaS). Various types of virtual resources are deployed in the resource pool for external customers to choose from. The cloud computing resource pool mainly includes: computing devices (which can be virtual machines, including operating systems), storage devices, and network devices.
[0105] The structure of the electronic device according to the embodiments of the present invention will be described in detail below. The electronic device can be implemented in various forms, such as a dedicated terminal with audio information processing functions, such as a gateway, or a server with audio information processing functions, such as the Figure 1 server 200 described above. Figure 2 FIG. is a schematic diagram of the composition structure of the electronic device provided by the embodiments of the present invention. It can be understood that Figure 2 only shows the exemplary structure of the server rather than all structures, and can be implemented according to needs Figure 2 the partial structure or all structures shown.
[0106] The electronic device provided by the embodiments of the present invention includes: at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. Each component in the electronic device 20 is coupled together through a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 205.
[0107] Among them, the user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, a button, a button, a touchpad, or a touch screen, etc.
[0108] It can be understood that the memory 202 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The memory 202 in the embodiments of the present invention is capable of storing data to support the operation of a terminal (such as 10-1). Examples of such data include: any computer programs for operating on the terminal (such as 10-1), such as an operating system and application programs. Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs.
[0109] In some embodiments, the audio information processing device provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the audio information processing device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the audio information processing method provided by the embodiments of the present invention. For example, a processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuits), DSPs, programmable logic devices (PLDs, Programmable Logic Devices), complex programmable logic devices (CPLDs, Complex Programmable Logic Devices), field-programmable gate arrays (FPGAs, Field-Programmable Gate Arrays) or other electronic components.
[0110] As an example of the audio information processing device provided by the embodiments of the present invention implemented in a combination of software and hardware, the audio information processing device provided by the embodiments of the present invention can be directly embodied as a combination of software modules executed by the processor 201. The software modules can be located in a storage medium, and the storage medium is located in the memory 202. The processor 201 reads the executable instructions included in the software modules in the memory 202 and combines the necessary hardware (for example, including the processor 201 and other components connected to the bus 205) to complete the audio information processing method provided by the embodiments of the present invention.
[0111] As an example, the processor 201 can be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0112] As an example of the audio information processing device provided by the embodiments of the present invention implemented in hardware, the device provided by the embodiments of the present invention can be directly implemented by a processor 201 in the form of a hardware decoding processor. For example, it can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components to execute the audio information processing method provided by the embodiments of the present invention.
[0113] The memory 202 in the embodiments of the present invention is used to store various types of data to support the operation of the electronic device 20. Examples of these data include: any executable instructions for operating on the electronic device 20, such as executable instructions, and the program for implementing the audio information processing method of the embodiments of the present invention can be included in the executable instructions.
[0114] In some other embodiments, the audio information processing device provided by the embodiments of the present invention can be implemented in a software manner. Figure 2 The audio information processing device 2020 stored in the memory 202 is shown. It can be software in the form of a program and a plug-in, etc., and includes a series of modules. As an example of the program stored in the memory 202, it can include the audio information processing device 2020. The audio information processing device 2020 includes the following software modules: a first information transmission module 2081 and a first information processing module 2082. When the software modules in the audio information processing device 2020 are read into the RAM by the processor 201 and executed, the audio information processing method provided by the embodiments of the present invention will be implemented. The functions of each software module in the audio information processing device 2020 are introduced below:
[0115] The first information transmission module 2081 is used to parse and process the template information of the audio service processing environment to obtain the first audio track data.
[0116] The first information processing module 2082 is used to save the first audio track data to an audio information storage hash table, where the audio information storage hash table is used to save audio information, and the audio information is used to store the first audio track data.
[0117] The first information processing module 2082 is used to dynamically modify the first audio track data in response to a dynamic modification instruction to obtain second audio track data.
[0118] The first information processing module 2082 is configured to obtain audio information from the audio information storage hash table in response to an audio service output instruction, and configure a corresponding audio frame data reader for the audio information;
[0119] The first information processing module 2082 is configured to extract a first audio frame from the second audio track data stored in the audio information and corresponding to a target timestamp through the audio frame data reader;
[0120] The first information processing module 2082 is configured to perform combined processing on first audio frames corresponding to different target timestamps to obtain and output a second audio frame, so as to respond to the audio service output instruction through the second audio frame.
[0121] In some other embodiments, the audio information presentation device provided by the embodiments of the present invention may also be implemented in a software manner. The audio information presentation device 2021 in the memory 202 may be software in the form of a program and a plug-in, etc., and includes a series of modules. As an example of the program stored in the memory 202, it may include the audio information presentation device 2021. The audio information presentation device 2021 includes the following software modules: a second information transmission module 2083 and a second information processing module 2084. When the software modules in the audio information presentation device 2021 are read into the RAM by the processor 201 and executed, the audio information presentation method provided by the embodiments of the present invention will be implemented. The functions of each software module in the audio information presentation device 2021 are introduced below:
[0122] The second information transmission module 2083 is configured to display a user interface and present a task control function item in the user interface. The task control function item is used to dynamically modify the first audio track data through a condition-triggered dynamic modification instruction.
[0123] The second information processing module 2084 is configured to obtain animation effect information including an audio service output instruction in response to a trigger operation on the task control function item.
[0124] The second information processing module 2084 is configured to obtain a second audio frame corresponding to the audio service output instruction.
[0125] The second information processing module 2084 is configured to present the animation effect information and the second audio frame in the user interface.
[0126] According to Figure 2The electronic device shown. In one aspect of the present application, the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the different embodiments and combinations of embodiments provided in the various optional implementations of the above audio information processing method or audio information presentation method.
[0127] In combination with Figure 2 The audio information processing method provided by the embodiments of the present invention will be described in combination with the electronic device 20 shown. Before introducing the audio information processing method provided by the present invention, first, the organizational logic of the usage environment of the audio information processing method provided by the present application and the architecture of data (Entity Component System) will be introduced. Refer to Figure 3 , Figure 3 It is a schematic structural diagram of the organizational logic and data architecture in the embodiments of the present invention. Among them, the architecture mode of the organizational logic and data (EntityComponent System) includes: an audio processing system Audio Process System and a video processing system VideoProcess System. The audio processing system Audio Process System interacts with an entity (Entity) that includes two types of components, Time and Audio. The video processing system Video Process System interacts with an entity (Entity) that includes two types of components, Time and Video. Due to the complexity of the services processed by the organizational logic and data architecture, the remaining entities that do not contain compliant components will not be processed (for example Figure 3 Entity 6 and Entity 7 shown are not processed by the audio processing system Audio Process System and the video processing system Video Process System because they do not contain both Time and Audio, or both Time and Video at the same time). Figure 3 The structure shown also includes two systems: a script system (Script) and an event trigger system (Event Trigger). These two systems can modify the content of the Audio component in real time, and the audio processing system can process the latest Audio data in real time. Taking the usage environment of game animations as an example, when using Figure 3 the organizational logic and data architecture mode shown to process audio information, the entity can fill different game special effect transactions in the game program. The component can represent the data (the data) in the game. When using Figure 3When the architecture of the shown organizational logic and data is involved, the AudioProcessSystem can continuously read the components (Components are components, pure data modules, used to store the data required by the entity) contained in the entity of interest, and then process them. During use, the conversion of the motion state or the state of the audio data is achieved by updating the component data in real time. The AudioProcessSystem continuously reads these Component data to achieve real-time audio updates..
[0128] See Figure 4 , Figure 4 FIG. is an optional flowchart of the audio information processing method provided by an embodiment of the present invention. It can be understood that Figure 4 the steps shown can be executed by various terminals running the audio information processing device. For example, it can be a dedicated terminal with audio information processing functions, such as a terminal running a short video client (or a WeChat applet triggered for making animation special effect information). The following is for Figure 4 the steps shown are described.
[0129] Step 401: The audio information processing device analyzes and processes the template information of the audio service processing environment to obtain the first audio track data.
[0130] In some embodiments of the present invention, analyzing and processing the template information of the audio service processing environment to obtain the first audio track data can be achieved by the following methods:
[0131] Analyze the template information of the audio service processing environment to obtain the timing information of the template information; according to the timing information of the template information, analyze the audio parameters corresponding to the template information to obtain the audio type and track information parameters corresponding to the template information; based on the audio type and track information parameters corresponding to the template information, extract the template information to obtain the first audio track data corresponding to the template information. Since the audio type is related to the audio service processing environment and can include single audio, audio matching video information, and audio matching animation resources, it is necessary to classify and process according to the audio type and obtain the first audio track data corresponding to the template information based on the corresponding track information parameters.
[0132] Among them, taking the audio information as the background music in the long video as an example, the template information in the video data can be obtained first; then the corresponding playing duration parameter and audio track information parameter can be obtained by parsing the audio header decoding data AACDecoderSpecificInfo and audio data configuration information AudioSpecificConfig in the template information. The audio data configuration information AudioSpecificConfig is used to generate ADST (including the sampling rate, number of channels, and frame length data in the audio data). Based on the audio track information, other audio packets in the video data are obtained, and the original audio data is parsed. Finally, the AAC ES stream is packed into the ADTS format through the AAC decoder of the audio data header. Among them, a 7-byte header file ADTSheader is added before the AAC ES stream to extract the audio track data.
[0133] In some embodiments of the present invention, due to the complex composition of the audio track data, when making the audio corresponding to the animation special effect information (such as the special effect plug-in in the short video), in order to achieve independent control of different audio information in the audio track data, different types of audio information can be extracted respectively. For example: when the audio type is a single audio, the audio data is extracted from the audio resource component of the template information to construct the first audio information; when the audio type is the audio matching the video information, the audio data is extracted from the multimedia information component of the template information to construct the second audio information; when the audio type is the audio matching the animation resource, the audio data is extracted from the animation resource component of the template information to construct the third audio information; the first audio information, the second audio information, and the third audio information are combined to obtain the first audio track data corresponding to the template information. By combining different types of audio information, the audio performance of the produced animation special effect information is made more abundant, and at the same time, the control of different types of audio information is more convenient, reducing the complexity of audio information processing, being applicable to more short video users, and effectively expanding the usage range of audio information processing.
[0134] Step 402: The audio information processing device saves the first audio track data into the audio information storage hash table.
[0135] Among them, the audio information storage hash table is used to save audio information, and the first audio track data is stored in the audio information.
[0136] Step 403: The audio information processing device responds to the dynamic modification instruction and dynamically modifies the first audio track data to obtain the second audio track data.
[0137] In some embodiments of the present invention, in response to a dynamic modification instruction, the first audio track data is dynamically modified to obtain second audio track data, including:
[0138] Receiving the dynamic modification instruction, where the dynamic modification instruction includes at least one of the following:
[0139] Adjusting the start position of audio track data playback, pausing playback, resuming playback, seeking playback, conditional triggering, and scripted playback; dynamically modifying the first audio track data in the audio information storage hash table according to the type of the dynamic modification instruction to obtain second audio track data; in response to the dynamic modification instruction, modifying the audio information storage hash table to obtain an audio information storage hash table corresponding to the second audio track data. Among them, refer to Figure 5 , Figure 5 is a schematic diagram of the processing of audio track data by the dynamic modification instruction in the embodiments of the present invention. Specifically, the basic control of audio by the dynamic modification instruction includes two major categories. The first category is the operation of the audio track, including: starting playback, stopping, pausing, resuming, seeking playback (i.e., seek, for example, changing the start playback position, for example, a 60-second audio starts playing from the 10th second), intercepting (for example, only 6 seconds of a 60-second audio is played), delaying playback for a specific time (for example, delaying playback for 5 seconds), adding a new audio to the existing playback or modifying the source file path of an audio; the second category is adjusting the playback rate and volume of the audio. At the same time, the dynamic modification instruction also includes: supporting conditional triggering or script programs to control the audio track data. For example: when multiple audios can be played simultaneously through a script program, the respective playback states of each audio can be modified separately without affecting each other. When the dynamic modification instruction is conditional triggering control, a new audio can be dynamically added for playback after a specific condition occurs. For example: after successfully recognizing a face in a short video, a certain sound effect information can be played, or the background music needs to be paused and switched to another background music.
[0140] When the dynamic modification instruction is to adjust the playback rate of the audio frame, the first audio track data in the audio information is adjusted by the audio frame data reader to obtain an audio frame playback rate matching the dynamic modification instruction; when the dynamic modification instruction is to adjust the volume of the audio frame, the first audio track data in the audio information is adjusted by the audio frame data reader to obtain the volume of the audio frame matching the dynamic modification instruction.
[0141] In some embodiments of the present invention, the script program can dynamically modify the state of each audio during the process of applying the template, including operations such as removing an audio, pausing / resuming / advancing / retreating the starting playback position, etc., and can dynamically control the playback of audio information, making the playback of audio information more flexible.
[0142] Step 404: The audio information processing device responds to the audio service output instruction, obtains audio information from the audio information storage hash table, and configures a corresponding audio frame data reader for the audio information.
[0143] In some embodiments of the present invention, during the process of configuring a corresponding audio frame data reader for the audio information, when first responding to the audio service output instruction and obtaining audio information from the audio information storage hash table, a corresponding first audio frame data reader is configured for the audio information; the persistent state of the audio information is detected, and when the audio information persists, the persistent state of the first audio frame data reader is maintained, and the data information in the audio information is updated; when the audio information is removed, the first audio frame data reader is deleted, and a second audio frame data reader is configured according to the change of the audio information. In this process, because each frame of audio information needs to be synchronized with the latest audio information, if the audio frame data reader is frequently created, it will waste the processing resources of the CPU. By detecting the audio data, when first responding to the audio service output instruction and obtaining audio information from the audio information storage hash table, a corresponding audio frame data reader can be created. If the audio data persists in the subsequent processing process, the corresponding audio frame data reader can maintain its persistent state and does not release the corresponding resources, and only the member attribute information saved in the hash table needs to be updated. If an audio information is removed, the list of audio frame data readers held by AudioOutput will also synchronously delete the discarded audio frame data reader, and a new audio frame data reader will be newly created and associated in response to the newly added audio data.
[0144] Step 405: The audio information processing device extracts the first audio frame in the second audio track data corresponding to the target timestamp stored in the audio information through the audio frame data reader.
[0145] Step 406: The audio information processing device performs combined processing on the first audio frames corresponding to different target timestamps, obtains and outputs a second audio frame, so as to respond to the audio service output instruction through the second audio frame.
[0146] Continue to refer to Figure 6 , Figure 6 which is an optional flowchart of the audio information processing method provided by the embodiment of the present invention. It can be understood thatFigure 6 The steps shown can be executed by various terminals running an audio information processing device. For example, it can be a dedicated terminal with audio information processing functions, such as a terminal running a short video client (or a WeChat mini program triggered for making animation special effect information). The following is for Figure 6 the steps shown for illustration.
[0147] Step 601: When the dynamic modification instruction is addressing playback, or the playback start position is adjusted to the start position, determine the target time parameter that matches the dynamic modification instruction, and save the target time parameter in the audio information storage hash table.
[0148] When outputting the second audio frame, perform a comparison process on the target time parameter and the target timestamp to determine the timestamp comparison result.
[0149] Based on the timestamp comparison result, trigger an addressing playback process to ensure that the target time parameter and the target timestamp are synchronized when the second audio frame is output.
[0150] In some embodiments of the present invention, after the audio information starts playing, except for addressing playback or restarting playback when the audio loop plays to the end, there is no need to synchronize the current audio time. The audio frame data reader can automatically set the audio track of the corresponding audio information to the playback time of the next audio frame. When performing addressing playback or playing from the start position, the target targetTime can be set to the audio information and written into the corresponding hash table. According to the targetTime, compare it with the playback timestamp of the current audio frame data reader to calculate whether addressing processing is required.
[0151] Step 604: When parsing the template information of the audio service processing environment and no first audio track data is obtained, output an empty data frame that matches the target timestamp as the second audio frame.
[0152] Among them, when a template does not contain audio, AudioOutput will execute a response to the audio service output instruction by returning an empty data frame (all 0 data).
[0153] The following uses the playback of special effect information in a short video to illustrate the audio information processing method provided in this application. Continuing to refer to Figure 7 , Figure 7 which is an optional flowchart of the audio information processing method provided in an embodiment of the present invention. It can be understood that Figure 7 the steps shown can be executed by various terminals running an audio information processing device. For example, it can be a terminal running a short video client, and specifically includes the following steps:
[0154] Step 701: Analyze the template to construct the ECS system.
[0155] Among them, Figure 8 is the schematic diagram of animation special effect information in the embodiments of the present invention. Through the method shown in Figure 7 to process the animation special effect information in Figure 8 , specifically, it can respond to the viewing operation of the task control function item, present a content page including the template information of the audio service processing environment, and present at least one interactive function item in the content page. The interactive function item is used to implement the interaction with the audio service processing environment; receive the interactive operation triggered by the interactive function item for the audio service processing environment to execute the corresponding interactive instruction. Taking the animation special effect information shown in Figure 8 as an example, by triggering the task control function item 801 in the user interface 800, the saved first audio track data can be dynamically modified. For example, through the dynamically modified instruction triggered by the condition, the playback start position of the first audio track data can be adjusted. Further, when presenting the animation special effect information in the user interface, the animation special effect information and the second audio frame can be presented in the user interface in response to the audio service output instruction.
[0156] Step 702: Conditionally trigger the update process management of the script program.
[0157] Step 703: Obtain the Component (business-related) containing audio.
[0158] Step 704: Save the first audio track data to the audio information storage hash table.
[0159] Step 705: Dynamically modify the first audio track data in response to the dynamically modified instruction.
[0160] In some embodiments of the present invention, the first audio track data can be the complete audio track data of a song. By triggering the task control function item 801 in the user interface, the saved first audio track data is dynamically modified through the dynamically modified instruction for addressing playback, intercepting different audio frames in the song for combined processing to form the second audio frame, and the lyric special effect information can be presented in the user interface 800 and the second audio frame can be output in response to the audio service output instruction, making the audio information processing process in response to the audio service output instruction more convenient, improving the convenience of audio information processing, and enabling the user to obtain a more convenient usage experience.
[0161] Step 706: Configure the corresponding audio frame data reader for the audio information.
[0162] Step 707: Detect the continuous state of the audio information.
[0163] Step 708: When the dynamic modification instruction is address-based playback or the playback start position is adjusted to the start position, determine the target time parameter that matches the dynamic modification instruction, and save the target time parameter in the audio information storage hash table.
[0164] Step 709: Combine the first audio frames corresponding to different target timestamps to obtain a second audio frame.
[0165] Step 710: Output the second audio frame.
[0166] Reference Figure 9 , Figure 9 is a schematic diagram of animation special effect information in an embodiment of the present invention. In the content page 900, at least one interactive function item 901 can be presented in response to a viewing operation on the task control function item. Among them, the interactive function item 901 is used to implement interaction with the audio service processing environment; receive an interactive operation on the audio service processing environment triggered based on the interactive function item to execute a corresponding interactive instruction. Further, a first interactive prompt message can also be presented in the content page. The first interactive prompt message is used to prompt that the interactive content corresponding to the interactive operation can be presented in the user interface; in response to an operation of switching to the user interface, switch the content page to the user interface. Among them, when the user triggers a small program for making animation special effect information, when the first audio track data presented in the view interface does not meet the usage requirements, the user can reconfigure the first audio track data that meets the requirements through the content page 900. The user confirms through the first interactive prompt message that the information of the configured first audio track data can be presented in the content page for use by the special effect information production small program, enriching the user's choice types.
[0167] In Figure 9 the content page 900 shown, a second interactive prompt message can also be presented in the content page. The second interactive prompt message is used to prompt that the interactive content corresponding to the interactive operation can be presented in the special effect information template library interface; in response to an instruction to switch to the special effect information template library interface, switch the content page to the special effect information template library interface. Specifically, as Figure 9 shown, since the user's needs are diverse, the user can switch the content page to the special effect information template library interface through the second interactive prompt message, and can select special effect information that meets the user's usage requirements through the special effect information template library, enabling the user to obtain a richer usage experience.
[0168] In some embodiments of the present invention, a sharing function item for sharing the special effect information may also be presented in the user interface; in response to a trigger operation on the sharing function item for the special effect information, the special effect information is shared with users in different audio service processing environments. Thus, when using the audio information processing method provided in this application in a short video client, users can share the special effect information with different users through the sharing function item, and output a second audio frame that meets the service requirements in response to an audio service output instruction.
[0169] Advantageous technical effects:
[0170] In the embodiments of the present invention, by parsing and processing the template information of the audio service processing environment, first audio track data is obtained; the first audio track data is saved to an audio information storage hash table, where the audio information storage hash table is used to save audio information, and the first audio track data is stored in the audio information; in response to a dynamic modification instruction, the first audio track data is dynamically modified to obtain second audio track data; in response to an audio service output instruction, audio information is obtained from the audio information storage hash table, and a corresponding audio frame data reader is configured for the audio information; through the audio frame data reader, the first audio frame in the second audio track data corresponding to the target timestamp stored in the audio information is extracted; the first audio frames corresponding to different target timestamps are combined to obtain and output a second audio frame to respond to the audio service output instruction through the second audio frame; thus, real-time modification and flexible control of audio information can be achieved, making the audio information processing process in response to an audio service output instruction more convenient and improving the convenience of audio information processing.
[0171] The above is only the embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. An audio information processing method, characterized in that: The method comprises: Parsing the template information of the audio service processing environment to obtain first audio track data; Saving the first audio track data into an audio information storage hash table, wherein the audio information storage hash table is used to store audio information, and the first audio track data is stored in the audio information; In response to the dynamic modification instruction, dynamically modify the first audio track data to obtain second audio track data; In response to the audio service output instruction, obtain audio information from the audio information storage hash table, and configure a corresponding audio frame data reader for the audio information; Extracting, by the audio frame data reader, a first audio frame from the second audio track data corresponding to the target timestamp stored in the audio information; The first audio frames corresponding to different target timestamps are combined and processed to obtain and output a second audio frame, so as to respond to the audio service output instruction through the second audio frame.
2. The method according to claim 1, characterized in that The parsing of the template information of the audio service processing environment to obtain the first audio track data includes: Parsing the template information of the audio service processing environment to obtain timing information of the template information; Parsing the audio parameters corresponding to the template information according to the timing information of the template information to obtain the audio type and track information parameters corresponding to the template information; Based on the audio type and audio track information parameters corresponding to the template information, the template information is extracted to obtain the first audio track data corresponding to the template information.
3. The method according to claim 2, characterized in that The extracting the template information based on the audio type and audio track information parameters corresponding to the template information to obtain first audio track data corresponding to the template information includes: When the audio type is a single audio, extracting audio data from the audio resource component of the template information to construct the first audio information; When the audio type is audio that matches the video information, extracting audio data from the multimedia information component of the template information to construct second audio information; When the audio type is audio that matches the animation resource, extracting audio data from the animation resource component of the template information to construct third audio information; The first audio information, the second audio information, and the third audio information are combined to obtain first audio track data corresponding to the template information.
4. The method according to claim 1, wherein The dynamically modifying the first audio track data in response to the dynamic modification instruction to obtain the second audio track data includes: Receive the dynamic modification instruction, wherein the dynamic modification instruction includes at least one of the following: Adjust the start position of audio track data, pause it, resume it, seek to it, trigger it under certain conditions, and play it as a script. Dynamically modifying the first audio track data in the audio information storage hash table according to the type of the dynamic modification instruction to obtain second audio track data; In response to the dynamic modification instruction, the audio information storage hash table is modified to obtain an audio information storage hash table corresponding to the second audio track data.
5. The method according to claim 1, wherein The step of obtaining audio information from the audio information storage hash table in response to the audio service output instruction and configuring a corresponding audio frame data reader for the audio information comprises: When obtaining audio information from the audio information storage hash table in response to the audio service output instruction for the first time, configuring a corresponding first audio frame data reader for the audio information; detecting a persistent state of the audio information, and when the audio information persists, maintaining the persistent state of the first audio frame data reader and updating data information in the audio information; When the audio information is removed and new audio information is added, the first audio frame data reader is deleted, and a second audio frame data reader is configured according to the change of the audio information.
6. The method according to claim 1, characterized in that The method further comprises: When the dynamic modification instruction is addressing playback, or adjusting the playback start position to the start position, determining a target time parameter that matches the dynamic modification instruction, and storing the target time parameter in an audio information storage hash table; When outputting the second audio frame, comparing the target time parameter and the target timestamp to determine a timestamp comparison result; Based on the timestamp comparison result, an addressing playback process is triggered to ensure that the target time parameter and the target timestamp remain synchronized when the second audio frame is output.
7. The method according to claim 1, characterized in that The method further comprises: When the template information of the audio service processing environment is parsed and the first audio track data is not obtained, an empty data frame matching the target timestamp is output as the second audio frame.
8. The method according to claim 1, characterized in that The method further comprises: When the dynamic modification instruction is to adjust the playback rate of the audio frame, adjusting the first audio track data in the audio information by the audio frame data reader to obtain an audio frame playback rate that matches the dynamic modification instruction; When the dynamic modification instruction is to adjust the volume of the audio frame, the first audio track data in the audio information is adjusted by the audio frame data reader to obtain the volume of the audio frame that matches the dynamic modification instruction.
9. A method for presenting audio information, characterized in that: The method comprises: Displaying a user interface, and presenting a task control function item in the user interface, wherein the task control function item is used to dynamically modify the first audio track data through a dynamic modification instruction triggered by a condition to obtain second audio track data; In response to a triggering operation on the task control function item, obtaining animation special effect information including an audio service output instruction; Obtaining a second audio frame corresponding to the audio service output instruction, wherein the second audio frame is obtained by: obtaining audio information from an audio information storage hash table and configuring a corresponding audio frame data reader for the audio information; extracting, by the audio frame data reader, a first audio frame from the second audio track data corresponding to a target timestamp stored in the audio information; and combining first audio frames corresponding to different target timestamps to obtain the second audio frame, wherein the first audio track data is stored in the audio information, and the audio information is stored in the audio information storage hash table; The animation special effect information and the second audio frame are presented in the user interface.
10. The method according to claim 9, characterized in that The method further comprises: The task control function item is further used to dynamically modify the first audio track data through a dynamic modification instruction of addressing playback; In response to a triggering operation on the task control function item, obtaining lyrics special effect information including an audio service output instruction; Acquire a second audio frame corresponding to the audio service output instruction; The lyrics special effect information and the second audio frame are presented in the user interface.
11. The method according to claim 9, characterized in that The method further comprises: In response to a viewing operation on the task control function item, presenting a content page including template information of the audio service processing environment, and presenting at least one interactive function item in the content page, the interactive function item being used to implement interaction with the audio service processing environment; An interactive operation for the audio service processing environment triggered by the interactive function item is received to execute a corresponding interactive instruction.
12. An audio information processing device, characterized in that: The device comprises: A first information transmission module is used to parse the template information of the audio service processing environment to obtain first audio track data; A first information processing module is configured to save the first audio track data into an audio information storage hash table, wherein the audio information storage hash table is configured to store audio information, and the audio information is configured to store the first audio track data; The first information processing module is configured to dynamically modify the first audio track data in response to a dynamic modification instruction to obtain second audio track data; The first information processing module is configured to obtain audio information from the audio information storage hash table in response to an audio service output instruction, and configure a corresponding audio frame data reader for the audio information; The first information processing module is configured to extract, through the audio frame data reader, a first audio frame from the second audio track data corresponding to a target timestamp stored in the audio information; The first information processing module is configured to combine and process the first audio frames corresponding to different target timestamps to obtain and output a second audio frame, so as to respond to the audio service output instruction through the second audio frame.
13. An audio information presentation device, characterized in that: The device comprises: a second information transmission module, configured to display a user interface and present a task control function item in the user interface, wherein the task control function item is configured to dynamically modify the first audio track data through a conditionally triggered dynamic modification instruction to obtain second audio track data; A second information processing module is configured to obtain animation special effect information including an audio service output instruction in response to a triggering operation on the task control function item; The second information processing module is configured to obtain a second audio frame corresponding to the audio service output instruction, wherein the second audio frame is obtained by: obtaining audio information from an audio information storage hash table and configuring a corresponding audio frame data reader for the audio information; extracting, by the audio frame data reader, a first audio frame from the second audio track data corresponding to a target timestamp stored in the audio information; and combining first audio frames corresponding to different target timestamps to obtain the second audio frame, wherein the first audio track data is stored in the audio information, and the audio information is stored in the audio information storage hash table. The second information processing module is used to present the animation special effect information and the second audio frame in the user interface.
14. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable instructions; The processor is configured to implement the audio information processing method described in any one of claims 1 to 8, or the audio information presentation method described in any one of claims 9 to 11, when running the executable instructions stored in the memory.
15. A computer-readable storage medium storing executable instructions, characterized in that: When the executable instructions are executed by the processor, the audio information processing method described in any one of claims 1 to 8 is implemented, or the audio information presentation method described in any one of claims 9 to 11 is implemented.
16. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the audio information processing method according to any one of claims 1 to 8 is implemented, or the audio information presentation method according to any one of claims 9 to 11 is implemented.
Citation Information
Patent Citations
Method and system for editing audio-video
CN102638658A