Audio processing method and device, storage medium and electronic device

By acquiring the speech and semantic features of audio data, generating segmentation location data, and performing audio segmentation and visualization processing, the problem of low efficiency in the use of audio information is solved, and intelligent analysis and visualization of audio data are realized, improving the efficiency and adaptability of data application.

CN120932650APending Publication Date: 2025-11-11QINGDAO HAIER TECH +2

Patent Information

Application Number
CN202510969081.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing audio processing technologies lack intelligent analysis and processing, resulting in low efficiency in the use of audio information, making it difficult to meet application needs. This necessitates secondary processing and organization by the application side, leading to a waste of resources and costs.

Method used

By acquiring speech and semantic feature data from audio data, segmentation location data is generated. Based on this data, audio segmentation and visualization processing are performed to generate audio segment data and segment visualization data, which are then provided to the user interface.

Benefits of technology

It enables standardized preprocessing, intelligent analysis, and visualization of audio data, improving the utilization rate of audio information, reducing data processing costs in application scenarios, and enhancing data application efficiency and scenario adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932650A_ABST
    Figure CN120932650A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and device, a storage medium and an electronic device, and relates to the technical field of data processing, and the audio processing method comprises the steps: obtaining voice feature data and semantic feature data of original audio data; the original audio data is audio data extracted from the target audio and video data; generating segmentation position data according to the voice feature data and the semantic feature data; and generating audio paragraph data and paragraph visualization data according to the segmentation position data, and providing the paragraph visualization data to a user interaction interface. According to the technical scheme of the embodiment of the invention, the utilization rate of the audio information can be improved, the data processing cost in an application scene is remarkably reduced, and the data application efficiency and scene adaptability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more specifically, to an audio processing method, apparatus, storage medium, and electronic device. Background Technology

[0002] Existing audio and video processing technologies often focus solely on video content processing, neglecting the analysis and presentation of audio content. Audio processing technologies are also limited to speech-to-text transcription and basic physical characteristics of audio, such as noise reduction, enhancement, and compression, lacking intelligent analysis and processing at the audio content level. This results in low efficiency in the use of audio information. In increasingly common application scenarios, such as meeting minutes, lectures, and voice-guided reading, existing audio processing technologies are insufficient to meet application requirements and achieve the desired results. This necessitates secondary processing and organization of data on the application side, leading to a waste of resources and costs. Summary of the Invention

[0003] This application provides an audio processing method, apparatus, storage medium, and electronic device, aiming to improve the utilization rate of audio information, realize standardized preprocessing, intelligent analysis, and visualization of audio data, significantly reduce data processing costs in application scenarios, and improve data application efficiency and scenario adaptability.

[0004] According to one aspect of the embodiments of this application, an audio processing method is provided, including:

[0005] Acquire speech feature data and semantic feature data from the raw audio data; the raw audio data is the audio data extracted from the target audio and video data;

[0006] Based on the speech feature data and the semantic feature data, segmentation location data is generated;

[0007] Audio segment data and segment visualization data are generated based on the segmentation location data, and the segment visualization data is provided to the user interface.

[0008] In an optional implementation, the semantic feature data includes: speech semantic data; the process of acquiring the speech feature data and semantic feature data of the original audio data includes:

[0009] Obtain the text-converted data of the original audio data;

[0010] Natural language processing is performed on the text-to-text data to obtain the speech-semantic data.

[0011] In an optional implementation, generating audio segment data and segment visualization data based on the segmentation location data includes:

[0012] The original audio data is segmented based on the segmentation location data to generate the audio segment data and the segment marker data of the original audio data;

[0013] Based on the segmentation location data, the text conversion data is segmented to generate text paragraph data and paragraph timestamp data of the text paragraph data;

[0014] The original audio data, the paragraph mark data, the text paragraph data, and the paragraph timestamp data are processed for output visualization to generate the paragraph visualization data.

[0015] In an optional implementation, the semantic feature data includes: image semantic data; the acquisition of the speech feature data and semantic feature data of the original audio data includes:

[0016] Extract video frame data from the target audio and video data;

[0017] Obtain the semantic data of the video frame data.

[0018] In an optional implementation, generating segmentation location data based on the speech feature data and the semantic feature data includes:

[0019] The original audio data and the video frame data are subjected to time-series alignment processing;

[0020] Based on the result of the temporal alignment process, the speech feature data and the semantic feature data are subjected to decision fusion processing to generate the segmentation location data.

[0021] In an optional implementation, generating audio segment data and segment visualization data based on the segmentation location data includes:

[0022] The original audio data is segmented based on the segmentation location data to generate the audio segment data and the segment marker data of the original audio data;

[0023] Based on the segmentation location data and the image semantic data, the video image data is processed to extract images and generate scene navigation data and key frame identification data.

[0024] The original audio data, the paragraph marker data, the scene navigation data, and the keyframe identifier data are processed for output visualization to generate the paragraph visualization data.

[0025] In an optional implementation, after generating audio segment data and segment visualization data based on the segmentation location data, the method further includes:

[0026] Interactive link data is generated from the paragraph visualization data.

[0027] According to another aspect of the embodiments of this application, an audio processing apparatus is provided, comprising:

[0028] The feature acquisition module is used to acquire speech feature data and semantic feature data of the original audio data; the original audio data is the audio data extracted from the target audio and video data;

[0029] The data segmentation module is used to generate segmentation location data based on the speech feature data and the semantic feature data;

[0030] The data generation module is used to generate audio segment data and segment visualization data based on the segmentation position data, and to provide the segment visualization data to the user interface.

[0031] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein the program, when executed, performs the audio processing method provided in any embodiment of the present invention.

[0032] According to another aspect of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute an audio processing method provided in any embodiment of the present invention through the computer program.

[0033] This application provides an audio processing method, apparatus, storage medium, and electronic device. By acquiring speech and semantic feature data of audio data, it determines an intelligent segmentation scheme for the audio data, generates segmented audio segment data and corresponding segment visualization data, and provides the segment visualization data to the user interface. This solves the problems of low efficiency in the use of audio information, low level of intelligence in audio processing, and difficulty in meeting application needs. It realizes standardized preprocessing, intelligent analysis, and visualization of audio data, improves the utilization rate of audio information, significantly reduces data processing costs in application scenarios, and enhances data application efficiency and scenario adaptability. Attached Figure Description

[0034] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0035] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a schematic diagram of the hardware environment for an audio processing method according to an embodiment of this application;

[0037] Figure 2 This is a flowchart of an audio processing method provided in an embodiment of the present invention;

[0038] Figure 3 This is a flowchart of an audio processing method provided in an embodiment of the present invention;

[0039] Figure 4 This is a flowchart of an audio processing method provided in an embodiment of the present invention;

[0040] Figure 5 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present invention;

[0041] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0042] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0043] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0044] According to one aspect of the embodiments of this application, an audio processing method is provided. This audio processing method is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the above-mentioned audio processing method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0045] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0046] This embodiment provides an audio processing method applied to a server. Figure 2 A flowchart of an audio processing method provided in an embodiment of the present invention includes the following steps:

[0047] Step 210: Obtain the speech feature data and semantic feature data of the original audio data.

[0048] The original audio data refers to the audio data extracted from the target audio / video data. The original audio data can be any speech data carrying semantic information that requires intelligent analysis and processing. The target audio / video data can be any audio / video data containing the original speech data, determined based on the application scenario. Speech feature data can be data describing the acoustic features of speech. Semantic feature data can be data describing the content meaning features of speech.

[0049] In an optional implementation, before acquiring the speech feature data and semantic feature data of the original audio data, the method may further include: extracting the original audio data from the target audio and video data.

[0050] The target audio and video data can be any audio and video data containing original voice data, determined according to the application scenario. The target audio and video data can be data in any existing audio and video format, such as MP4 (Moving Pictures Experts Group-4), AVI (Audio Video Interleaved), MP3 (MPEG-1 Audio Layer III or MPEG-2 Audio Layer III), or WAV (Waveform Audio File Format), and is not limited here. The original audio data can be extracted from the target audio and video data using any method mastered by those skilled in the art, such as using open-source libraries like FFmpeg (Fast Forward MPEG), ensuring that the quality of the extracted audio data is not compromised.

[0051] In a typical example, this is applied to a scenario where audio and video data containing a character's monologue or dialogue is provided to a user for language learning. The target audio and video data can be any pre-stored instructional video data, such as an instructional video selected by the user using the terminal, while the original audio data can be audio extracted from the instructional audio and video data.

[0052] After acquiring the raw audio data, it can be analyzed and processed in any feasible way to obtain speech feature data and semantic feature data.

[0053] In an optional implementation, before acquiring the speech feature data and semantic feature data of the original audio data, the process may further include: performing audio preprocessing on the original audio data.

[0054] Audio preprocessing can be an operation to improve the quality of audio data. Optionally, audio preprocessing can include audio noise reduction, background noise removal, and volume normalization. Specifically, Gaussian filtering can be used to remove random noise from the original audio data; Fourier transform-based spectral analysis can be used to identify and filter environmental noise in the original audio data; and dynamic range compression can be used to balance volume and improve speech clarity. Audio preprocessing can improve the quality of the original audio data, thereby improving the accuracy of speech feature data and semantic feature data.

[0055] Speech feature data can reflect the content characteristics of speech through its physical attributes. For example, it can identify the speaker's emotions based on intonation and speech rate, and it can identify the intervals between sentences based on pauses. Semantic feature data can reflect the content characteristics of speech through its linguistic attributes. For example, it can identify the topic of the speech content based on semantic content and contextual relationships.

[0056] In an optional implementation, the speech feature data may include phoneme feature data, prosodic feature data, energy feature data, and time-frequency feature data.

[0057] Among these, phoneme feature data can be data describing the basic phoneme units of speech. Prosodic feature data can be data describing the intonation, pauses, and speech rate of speech. Energy feature data can be data describing the time-domain distribution of the energy of speech. Time-frequency feature data can be data describing the time-domain distribution of the frequency components of speech. Key parameters of the acoustic properties of speech, such as pitch, timbre, speech rate, and pauses, can be obtained through techniques such as spectral analysis, energy calculation, and prosodic modeling, thus enabling the acquisition of speech feature data from the original audio data.

[0058] Based on different phoneme sets in phoneme feature data, such as initials and finals in Chinese, and vowels and consonants in English, the language or dialect in audio data can be distinguished; phoneme transformations can also help determine the division between sentences. Prosodic feature data can help determine the division of sentences and semantic paragraphs; for example, long pauses usually correspond to the end of a sentence or paragraph, changes in speech rate may be accompanied by semantic transitions or speaker changes, and intonation fluctuations can distinguish between interrogative, declarative, and negative tones. Energy feature data can quickly filter out noisy audio data without speakers, and energy mutation points can locate audio data with sudden interference, and can also provide a basis for speaker identification. Time-frequency feature data can also identify different speakers and environmental noise based on the frequency of speech.

[0059] In an optional implementation, semantic feature data may include word segmentation part-of-speech data, syntactic feature data, semantic feature data, and discourse feature data.

[0060] The data includes word segmentation and part-of-speech tagging, which describes the words and their parts of speech within the audio content. Syntactic feature data describes the sentence structure of the audio content, including components such as the subject, predicate, and object. Semantic feature data describes the semantic content and contextual relationships of the audio. Discourse feature data describes the discourse structure and logical relationships within the cross-sentence context of the audio content. Semantic feature data can be obtained by converting audio content into text and further parsing it using technologies such as speech recognition and natural language processing.

[0061] Based on word segmentation and part-of-speech data, the integrity of semantic units and topic shifts can be identified through words and their parts of speech, assisting in locating sentence segmentation positions and paragraph semantics. Based on syntactic feature data, semantic integrity can be judged by analyzing sentence structure, identifying the boundaries of complete semantic paragraphs. Based on semantic feature data, paragraph boundaries can be determined by the coherence of semantic content and contextual relevance, identifying topic shifts and changes in logical hierarchy. Based on discourse feature data, the macro-structure of paragraphs can be identified through cross-sentence discourse structure and logical relationships, clarifying the levels of argumentation.

[0062] In an optional implementation, the semantic feature data may include: speech semantic data; obtaining speech feature data and semantic feature data from the original audio data may include: obtaining text-to-text conversion data from the original audio data; performing natural language processing on the text-to-text conversion data to obtain speech semantic data.

[0063] Among them, speech semantic data can be data describing the semantics of speech content. Text conversion data can be data that records the text converted from speech content in audio data.

[0064] By converting the speech content in the original audio data into text, we obtain text-converted data. Natural language processing is then performed on this text-converted data to obtain its semantic features, which can then be used to generate semantic feature data of the original audio data. Optionally, a deep learning model, combining convolutional neural networks and recurrent neural networks, can be used to achieve high-accuracy speech recognition of the original audio data and output text-converted data.

[0065] In one alternative implementation, the text conversion data may include text timestamp data.

[0066] The text timestamp data can be data that marks the temporal position of the original audio data corresponding to each character and / or word in the text-to-text conversion data. The text corresponding to the speech at each temporal position in the original audio data can be generated by generating text timestamp data based on the speech's temporal position. Based on the text timestamp data, the text-to-text conversion data and the original audio data can be aligned in the temporal domain. This allows for the determination of the speech semantic data at the same temporal position in the original audio data based on the semantic meaning of the text content corresponding to each timestamp in the text-to-text conversion data, ensuring the accuracy of the features of the original audio data determined based on the speech semantic data.

[0067] In an optional implementation, the semantic feature data may include: image semantic data; obtaining speech feature data and semantic feature data from the original audio data may include: extracting video image data from the target audio and video data; obtaining image semantic data from the video image data.

[0068] The semantic data of the images can be data describing the meaning of the images corresponding to the speech content. The video image data can be the images in the audio and video data, and video image data can be extracted from the target audio and video data by any method known to those skilled in the art, without any limitation.

[0069] In target audio and video data, the same time-domain location includes both the original audio data and its corresponding video frame data. Therefore, by extracting the video frame data, the meaning of the video frame data at each time-domain location can be determined, which can then be used to determine the speech semantics of the original audio data at each time-domain location, generating the video semantic data of the original audio data.

[0070] In one optional implementation, the image semantic data may include scene segmentation data, character recognition data, motion recognition data, and subtitle extraction data.

[0071] The data includes: scene segmentation data, which describes the relationship between adjacent frames of a video image and reflects scene changes; person recognition data, which describes people present in the video and reflects character transitions; motion recognition data, which describes the actions and activities of people in the video and reflects the meaning of the images; and subtitle extraction data, which describes the subtitle text in the video and reflects the meaning of the images.

[0072] Scene segmentation data can be obtained by analyzing the scene correlation between adjacent video frames to identify scene transitions at any temporal location. This can include changes in visual features such as color, composition, and object type. Scene transitions typically correspond to content themes and / or spatiotemporal shifts, allowing the segment boundaries of scene transitions to be determined based on scene segmentation data. Character recognition data can be obtained by tracking the appearance and disappearance of characters in an image using facial recognition technology. When the main character in an image changes, segments dominated by different characters can be identified, thus determining segment boundaries based on the change in the main character. Action recognition data can be obtained by analyzing character actions, object movements, or action flows. Based on action recognition data, segment boundaries can be determined based on complete action or activity logic. Subtitle extraction data can be obtained by recognizing and extracting text content from an image. Subtitle extraction data can include chapter titles and / or spoken text. Segment boundaries can be determined directly based on chapter titles, or by semantically identifying themes and / or scene transitions based on spoken text, thus determining segment boundaries.

[0073] Step 220: Generate segmentation location data based on speech feature data and semantic feature data.

[0074] The segmentation location data can be data describing the temporal location of speech segment boundaries.

[0075] Based on the semantic content and paragraph structure of the speech determined by the speech feature data and semantic feature data, the paragraph boundaries of the audio can be accurately located in the time domain, thereby generating segmentation location data.

[0076] In one optional implementation, the temporal domain of paragraph boundaries can be initially located based on speech feature data. Basic phoneme units can be determined based on phoneme feature data, and the regional boundaries between speech and silence can be quickly identified by combining energy feature data. By combining time-frequency feature data, the temporal locations of speaker switching or significant changes in voice quality can be detected, thereby determining the temporal locations of candidate paragraph boundaries. Based on prosodic feature data describing features such as falling intonation and pauses at the end of sentences, natural sentence breaks can be further marked within continuous speech segments, determining the initially determined paragraph boundary locations that conform to the rhythm of natural language.

[0077] In one optional implementation, paragraph boundary positions can be confirmed based on semantic feature data. Word segmentation and part-of-speech data, along with syntactic feature data, can be used to verify whether the initially determined paragraph boundary positions are located at semantic unit boundaries. For example, if a preliminarily determined paragraph boundary position corresponds to a period, question mark, or other punctuation mark in the text, or is located at the end of a clause in a complex sentence, it can be retained as a paragraph boundary position; if a preliminarily determined paragraph boundary position is located in the middle of a word, it can be discarded. Semantic feature data and discourse feature data can be used to detect semantic breaks or logical level changes across sentences, marking macroscopic paragraph boundaries in the time domain. Therefore, segmentation position data can be generated based on the final obtained paragraph boundary positions.

[0078] In an optional implementation, the probability of a segment boundary existing at each time domain location can be determined by assigning different weight values ​​to the speech feature data and semantic feature data and performing fusion calculations. The segment boundary location can then be determined based on the magnitude of this probability, thereby generating segmentation location data.

[0079] In an optional implementation, generating segmentation location data based on speech feature data and semantic feature data may include: determining topic variation features, speaker variation features, semantic segmentation features, and time interval features of the original audio data based on the speech feature data and semantic feature data, and generating segmentation location data based on the topic variation features, speaker variation features, semantic segmentation features, and time interval features.

[0080] Specifically, topic variation features can be the characteristics of paragraph boundaries formed by significant changes in the topic of speech content. Speaker variation features can be the characteristics of paragraph boundaries formed by speech content produced by different speakers. Semantic segmentation features can be the characteristics of paragraph boundaries formed by the end of any complete semantic segment. Time interval features can be the characteristics of paragraph boundaries formed by longer pauses in speech.

[0081] When the topic of discussion in the audio content changes significantly, segmentation can be performed. Audio content from different speakers can also be segmented into different segments. Furthermore, semantic segmentation features can be used to ensure that each segment has complete semantics. Longer pauses can also be identified as natural segmentation points. Therefore, combining these features can generate segmentation location data.

[0082] In an optional implementation, generating segmentation location data based on speech feature data and semantic feature data may include: determining scene change features, on-screen character features, visual theme features, and multimodal consistency features of the original audio data based on the speech feature data and semantic feature data, and generating segmentation location data based on the scene change features, on-screen character features, visual theme features, and multimodal consistency features.

[0083] Among these, scene change features can be the segment boundaries formed by significant changes in the scene within the video frame corresponding to the audio content. Character features can be the segment boundaries formed by changes in the characters within the video frame corresponding to the audio content. Visual theme features can be the segment boundaries formed by changes in the visual theme of the video frame corresponding to the audio content. Multimodal consistency features can be the segment boundaries formed when both the audio content and the video frame content change.

[0084] When the scene in the video changes significantly, it indicates a significant semantic change in the audio content, allowing for segmentation. The appearance of different characters in the video also facilitates segmentation. Changes in the video's theme can also reflect segment changes. Furthermore, multimodal consistency ensures that changes in both audio and video content exist simultaneously at the boundaries of segmented segments. Therefore, based on these features, more multidimensional information can be integrated to generate segmentation location data.

[0085] The above implementation method, by combining speech recognition and natural language processing technologies, achieves intelligent analysis and processing of audio content, enabling segmentation without manual intervention and significantly improving work efficiency. It analyzes audio content from multiple dimensions, including semantic integrity, topic changes, speaker changes, and time intervals, ensuring the scientific and rational nature of segmentation.

[0086] In an optional implementation, generating segmentation location data based on speech feature data and semantic feature data may include: performing temporal alignment processing on the original audio data and video frame data; and performing decision fusion processing on the speech feature data and semantic feature data based on the result of the temporal alignment processing to generate segmentation location data.

[0087] Temporal alignment processing can be an operation that determines the correspondence between raw audio data and video image data at the same time domain location. Decision fusion processing can be an operation that fuses and calculates speech feature data and semantic feature data according to preset rules.

[0088] Temporal alignment ensures that segment boundaries determined from the semantic data of video frame data can be accurately located in the corresponding original audio data at the same temporal position, thus identifying semantic feature data at the corresponding temporal position in the original audio data. Therefore, in decision fusion processing, temporal feature fusion can be performed on speech feature data and semantic feature data to determine the speech and semantic features at any temporal position. The two types of features can then be combined to determine whether the temporal position can be identified as a segment boundary. Furthermore, weights can be assigned based on changes in speech features and changes in the semantic data of video frame data, prioritizing the capture of consistent features across multimodal data. For example, if the speech feature data describes a change in speaker and a shift in topic at the same temporal position, while the semantic data describes a scene transition, then that temporal position can be given a higher weight and identified as a strong segment boundary. Conversely, if the speech or semantic feature data at any temporal position is insufficient to identify it as a segment boundary, the segment boundary weight value at that temporal position can be reduced. The final segmented position data not only includes the speech feature segments of the original audio data but also synchronously aligns with the semantic changes in the image content of the video frame data.

[0089] In an optional implementation, generating segmentation location data based on speech feature data and semantic feature data may include: performing temporal alignment processing on the original audio data and text-to-text conversion data; and performing decision fusion processing on the speech feature data and semantic feature data based on the temporal alignment processing result to generate segmentation location data.

[0090] Temporal alignment ensures that paragraph boundaries determined from the speech and semantic data of the text-to-speech data can be accurately located in the original audio data at the corresponding temporal position, thereby identifying the semantic feature data at the corresponding temporal position in the original audio data. In decision fusion processing, the speech feature data of the original speech data at any temporal position, the speech and semantic data of the text-to-speech data, and / or the image semantic data of the video image data can be combined to determine whether the position is a paragraph boundary.

[0091] In one optional implementation, time alignment can be performed based on the text timestamp data in the text conversion data. Alternatively, after generating the segmentation position data, the segmentation position timestamp can be determined from the text timestamp data, thereby enabling the text conversion data to be segmented into segments that correspond to the meaning of the original audio data.

[0092] The multimodal information fusion method provided by the above embodiments combines information from both audio and video modalities, overcoming the limitations of single-modal analysis and improving the accuracy of segmentation.

[0093] Step 230: Generate audio segment data and segment visualization data based on the segmentation location data, and provide the segment visualization data to the user interface.

[0094] The audio segment data can be the speech segments extracted from the audio data. The segment visualization data can be the data that shows the segment structure of the audio data. The user interface can be the interface used to present the visualization content to the user.

[0095] The original audio data can be segmented into at least one audio segment by dividing it at the temporal location of the speech segment boundaries described by the segmentation location data. The segmentation strategy of the original audio data and the resulting audio segment data can be visualized through further generated segment visualization data, which can then be provided to a user interface for application by users or application scenarios.

[0096] This application provides an audio processing method, apparatus, storage medium, and electronic device. By acquiring speech and semantic feature data of audio data, it determines an intelligent segmentation scheme for the audio data, generates segmented audio segment data and corresponding segment visualization data, and provides the segment visualization data to the user interface. This solves the problems of low efficiency in the use of audio information, low level of intelligence in audio processing, and difficulty in meeting application needs. It realizes standardized preprocessing, intelligent analysis, and visualization of audio data, improves the utilization rate of audio information, significantly reduces data processing costs in application scenarios, and enhances data application efficiency and scenario adaptability.

[0097] In an optional implementation, generating audio segment data and segment visualization data based on the segmentation location data may include: performing audio segmentation processing on the original audio data based on the segmentation location data to generate audio segment data and segment marker data of the original audio data; performing text segmentation processing on the text conversion data based on the segmentation location data to generate text segment data and segment timestamp data of the text segment data; and performing output visualization processing on the original audio data, segment marker data, text segment data, and segment timestamp data to generate segment visualization data.

[0098] The audio segmentation process involves splitting the speech content of the original audio data into temporal segments. Segment tagging data is used to label each audio segment in the original audio data. The text segmentation process involves splitting the text content of the text-to-text conversion data into temporal segments. Text segment data can be the text segments split from the text-to-text conversion data, or the segments in the text-to-text conversion data corresponding to each audio segment. Output visualization processing involves visualizing the data. Segment timestamp data can be used to mark the start and end times of the corresponding audio segments in the text segment data.

[0099] The original audio data is segmented based on the segmentation location data. This process generates segment marker data within the original audio data, indicating the location of each segment. Therefore, the final segment visualization data can include each audio segment as well as the complete original audio file with segment marker data.

[0100] Furthermore, since the semantic segments of the text-to-text conversion data corresponding to the original audio data can correspond to the original audio data, text segmentation can be performed at the temporal segmentation positions described by the segmentation position data in the text-to-text conversion data. Thus, each resulting text segment can correspond one-to-one with each audio segment. Specifically, the temporal segmentation positions described by the segmentation position data in the text-to-text conversion data can be determined based on the text timestamp data. Correspondingly, the segment timestamp data of the text segment can be generated based on the temporal start and end positions of the audio segment data corresponding to the text segment data; alternatively, the segment timestamp data can be determined from the text timestamp data in the text-to-text conversion data. Therefore, the final generated segment visualization data can also include text segment data with segment timestamp data.

[0101] In an optional implementation, the paragraph visualization data may further include a visualized paragraph structure diagram. The visualized paragraph structure diagram may include the duration of each audio paragraph and text describing the paragraph content. The visualized paragraph structure diagram can be provided to users or application scenarios to provide an intuitive understanding of the paragraph structure of the original audio data and the main content of each segment, improving the convenience of data application and its applicability in different scenarios.

[0102] In the example provided in this embodiment, where audio and video data of monologues or dialogues are provided to users for language learning, the original audio data in the instructional video data can be extracted, broken down into audio segment data corresponding to each sentence, and corresponding segment marker data can be generated from the original audio data. The resulting segment visualization data can be applied, for example, to scenarios providing users with voice-guided reading functionality. Audio segment data can be acquired and played sequentially, with time intervals inserted between each segment for the user to follow along. The original audio data with segment marker data can be used to generate an adjustable voice playback progress bar for the user. Furthermore, text-converted data of the original audio data can be generated, broken down into text segment data and segment timestamp data, which can be provided to the user as a text reference for follow-up reading and to determine the text content corresponding to different audio segments. A correspondingly generated visualized segment structure diagram can provide the user with an overall understanding of the segment structure and content of the original audio data.

[0103] In one optional implementation, Figure 3 This is a flowchart illustrating an audio processing method provided in an embodiment of the present invention. Figure 3 As shown, the process begins with selecting audio and video files, supporting formats such as MP4, AVI, MP3, and WAV. This leads to audio extraction, followed by preprocessing of the extracted audio to achieve effects such as noise reduction, background noise removal, and volume normalization. Further, speech feature extraction can be performed on the preprocessed audio, extracting features such as phonemes, prosody, energy, and time-frequency. Speech recognition is then performed on the audio, converting it into timestamped text based on a deep learning model. Natural language processing is then applied to the text, including word segmentation, syntax, semantics, and discourse analysis. Finally, paragraph segmentation is performed, and the results are output.

[0104] In an optional implementation, generating audio segment data and segment visualization data based on segmentation location data may include: performing audio segmentation processing on the original audio data based on the segmentation location data to generate audio segment data and segment tagging data of the original audio data; performing image extraction processing on the video image data based on the segmentation location data and image semantic data to generate scene navigation data and keyframe identification data; and performing output visualization processing on the original audio data, segment tagging data, scene navigation data, and keyframe identification data to generate segment visualization data.

[0105] The image extraction process can be the operation of acquiring image frames from any frame of video frame data. Scene navigation data can be the data of image frames used to show the meaning of each audio segment. Keyframe identification data can be used to mark the image frames in the video frame data that can reflect the meaning of each audio segment.

[0106] If the segments of video frame data corresponding to the original audio data can be matched with the original audio data, then through frame extraction processing, one image frame can be extracted from each segment corresponding to each audio segment in the video frame data. This extracted image frame serves as the scene navigation data for each audio segment, and keyframe identification data is generated to mark the extracted image frames. Therefore, the final generated segment visualization data can include audio segment data with scene navigation data. Furthermore, the image frames marked by the keyframe identification data can be obtained and combined with the text segment data provided in the above embodiments of the present invention to generate a multimodal segment summary containing text and keyframe frames, corresponding to the audio segment data.

[0107] In an optional implementation, after generating audio segment data and segment visualization data based on the segmentation location data, the method may further include generating interactive link data in the segment visualization data.

[0108] Interactive link data can be data provided to users as an operation channel and used to perform corresponding operations based on user actions. For example, interactive link data can be a paragraph navigation interface containing scene navigation data, provided to users and allowing them to click on any scene navigation data, and returning the audio paragraph data corresponding to the clicked scene navigation data to the user. Interactive link data can be provided to users through a user interface.

[0109] In the example provided in this embodiment, which uses audio and video data of a character's monologue or dialogue to provide to a user for language learning, a segment navigation interface can be generated based on the scene navigation data corresponding to each audio segment in the teaching video data, so as to allow the user to select the audio segment data to be learned based on the segment content displayed by the scene navigation data.

[0110] In one optional implementation, Figure 4 This is a flowchart illustrating an audio processing method provided in an embodiment of the present invention. Figure 4 As shown, the process begins with selecting audio and video files, performing video content analysis and audio extraction, and then preprocessing, extracting speech features, and performing speech recognition on the extracted audio. Finally, multimodal information fusion is performed, including temporal alignment between multimodal data, direct feature fusion, and decision-level fusion. Based on natural language processing, enhanced segmentation based on audio, text, and video is then performed, and the final output is given.

[0111] The above implementation provides multiple output formats, including segmented text, audio clips, and visual structure diagrams, to meet the application needs of different scenarios. It is applicable to audio and video files in various fields, such as meeting minutes, course lectures, interview programs, and news reports, and has broad application prospects. It solves the problems of insufficient audio content analysis and unintelligent segmentation in the existing technology, and provides new technical means for the content understanding and utilization of audio and video files, which has important practical value.

[0112] According to another aspect of the embodiments of the present invention, Figure 5 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the device includes a feature acquisition module 510, a data segmentation module 520, and a data generation module 530, wherein:

[0113] The feature acquisition module 510 is used to acquire speech feature data and semantic feature data of the original audio data; the original audio data is the audio data extracted from the target audio and video data;

[0114] The data segmentation module 520 is used to generate segmentation location data based on speech feature data and semantic feature data;

[0115] The data generation module 530 is used to generate audio segment data and segment visualization data based on the segmentation position data, and to provide the segment visualization data to the user interface.

[0116] In an optional implementation, the semantic feature data may include: speech semantic data; the feature acquisition module 510 may include: a text conversion submodule for acquiring text conversion data of the original audio data; and a natural language submodule for performing natural language processing on the text conversion data to obtain speech semantic data.

[0117] In an optional implementation, the data generation module 530 may include: an audio segmentation submodule, used to perform audio segmentation processing on the original audio data according to the segmentation position data, generating audio segment data and segmentation mark data of the original audio data; a text segmentation submodule, used to perform text segmentation processing on the text conversion data according to the segmentation position data, generating text segment data and segmentation timestamp data of the text segment data; and a first visualization submodule, used to perform output visualization processing on the original audio data, segmentation mark data, text segment data and segmentation timestamp data, generating segment visualization data.

[0118] In an optional implementation, the semantic feature data may include: image semantic data; the feature acquisition module 510 may include: a video extraction submodule for extracting video image data from the target audio and video data; and an image semantic submodule for acquiring image semantic data from the video image data.

[0119] In an optional implementation, the data segmentation module 520 may include: a temporal alignment submodule for performing temporal alignment processing on the original audio data and video frame data; and a decision fusion submodule for performing decision fusion processing on the speech feature data and semantic feature data based on the result of the temporal alignment processing to generate segmentation position data.

[0120] In an optional implementation, the data generation module 530 may include: an audio segmentation submodule, used to perform audio segmentation processing on the original audio data according to the segmentation position data, generating audio segment data and segment marker data of the original audio data; a scene extraction submodule, used to perform scene extraction processing on the video scene data according to the segmentation position data and scene semantic data, generating scene navigation data and keyframe identification data; and a second visualization submodule, used to perform output visualization processing on the original audio data, segment marker data, scene navigation data and keyframe identification data, generating segment visualization data.

[0121] In an optional implementation, the apparatus may further include an interactive link module for generating interactive link data in the segment visualization data after generating audio segment data and segment visualization data based on the segmentation location data.

[0122] According to another aspect of the embodiments of the present invention, Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, the electronic device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the electronic device can be one or more. Figure 6Taking a processor 610 as an example; the processor 610, memory 620, input device 630, and output device 640 in the electronic device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0123] The memory 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the audio processing method in this embodiment of the invention (e.g., the feature acquisition module 510, data segmentation module 520, and data generation module 530 in the audio processing device). The processor 610 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 620, thereby implementing the aforementioned audio processing method.

[0124] The memory 620 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 620 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 620 may further include memory remotely located relative to the processor 610, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0125] Input device 630 can be used to receive input digital or character information, and generate key signal inputs related to user settings and function control of the electronic device. It can also be a camera for acquiring images and a sound pickup device for acquiring audio data. Output device 640 may include display devices such as a screen, and audio devices such as a speaker. It should be noted that the specific composition of input device 630 and output device 640 can be set according to actual conditions. Processor 610 executes various functional applications and data processing of the electronic device by running software programs, instructions, and modules stored in memory 620.

[0126] According to another aspect of the present invention, the present invention also provides a computer-readable storage medium comprising a stored program, wherein the program, when executed, performs the following audio processing method:

[0127] Acquire speech feature data and semantic feature data from the raw audio data; the raw audio data is the audio data extracted from the target audio and video data;

[0128] Based on the speech feature data and the semantic feature data, segmentation location data is generated;

[0129] Audio segment data and segment visualization data are generated based on the segmentation location data, and the segment visualization data is provided to the user interface.

[0130] Of course, the computer-executable instructions provided in the embodiments of the present invention are not limited to the method operations described above, but can also perform related operations in the audio processing method provided in any embodiment of the present invention.

[0131] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0132] It is worth noting that in the embodiments of the above-mentioned audio processing device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0133] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An audio processing method, characterized in that, include: Acquire speech feature data and semantic feature data from the raw audio data; the raw audio data is the audio data extracted from the target audio and video data; Based on the speech feature data and the semantic feature data, segmentation location data is generated; Audio segment data and segment visualization data are generated based on the segmentation location data, and the segment visualization data is provided to the user interface.

2. The method according to claim 1, characterized in that, The semantic feature data includes: speech semantic data: the speech feature data and semantic feature data obtained from the original audio data include: Obtain the text-converted data of the original audio data; Natural language processing is performed on the text-to-text conversion data to obtain the speech-semantic data.

3. The method according to claim 2, characterized in that, The step of generating audio segment data and segment visualization data based on the segmented position data includes: The original audio data is segmented based on the segmentation location data to generate the audio segment data and the segment marker data of the original audio data; Based on the segmentation location data, the text conversion data is segmented to generate text paragraph data and paragraph timestamp data of the text paragraph data; The original audio data, the paragraph mark data, the text paragraph data, and the paragraph timestamp data are processed for output visualization to generate the paragraph visualization data.

4. The method according to claim 1, characterized in that, The semantic feature data includes: image semantic data; the acquisition of the speech feature data and semantic feature data of the original audio data includes: Extract video frame data from the target audio and video data; Obtain the semantic data of the video frame data.

5. The method according to claim 4, characterized in that, The step of generating segmentation location data based on the speech feature data and the semantic feature data includes: The original audio data and the video frame data are subjected to time-series alignment processing; Based on the result of the temporal alignment process, the speech feature data and the semantic feature data are subjected to decision fusion processing to generate the segmentation location data.

6. The method according to claim 4, characterized in that, The step of generating audio segment data and segment visualization data based on the segmented position data includes: The original audio data is segmented based on the segmentation location data to generate the audio segment data and the segment marker data of the original audio data; Based on the segmentation location data and the image semantic data, the video image data is processed to extract images and generate scene navigation data and key frame identification data. The original audio data, the paragraph marker data, the scene navigation data, and the keyframe identifier data are processed for output visualization to generate the paragraph visualization data.

7. The method according to claim 6, characterized in that, After generating audio segment data and segment visualization data based on the segmented position data, the method further includes: Interactive link data is generated from the paragraph visualization data.

8. An audio processing apparatus, characterized in that, include: The feature acquisition module is used to acquire speech feature data and semantic feature data of the original audio data; the original audio data is the audio data extracted from the target audio and video data; The data segmentation module is used to generate segmentation location data based on the speech feature data and the semantic feature data; The data generation module is used to generate audio segment data and segment visualization data based on the segmentation position data, and to provide the segment visualization data to the user interface.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the audio processing method according to any one of claims 1 to 7.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the audio processing method of any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • End-to-end news program structuring method and structuring framework system thereof

    CN110012349A

  • Background audio construction method and device

    CN112584062A

  • Video processing method and device, equipment and medium

    CN113672765A

  • Broadcast video splitting method and device, storage medium and electronic equipment

    CN118264848A

  • Video segmentation method and device based on multi-modal features, equipment and storage medium

    CN119478763A

Cited By

  • Voice segmentation intelligent editing system based on deep learning

    CN121260170A

  • Intelligent voice analysis system based on multiple Agents

    CN121459795A