Audio and video data processing method and device, electronic equipment and medium

By processing audio and video data and extracting and matching the time information of the voice element set, the problems of low audio-video text matching accuracy and high manual cost are solved, and efficient and accurate subtitle output is achieved.

CN113849689BActive Publication Date: 2025-11-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111125712.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-24
Publication Date
2025-11-28
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Existing technologies have low matching accuracy between text and audio/video when adding text to audio and video, resulting in high manual costs and cumbersome operations.

Method used

By processing audio and video data, extracting the set of speech elements and time information, and matching them with the set of speech elements associated with text data, the time information of the text data is determined, and finally the text and audio/video data are linked and output.

Benefits of technology

It improves the efficiency and accuracy of subtitle matching, and reduces labor costs and operational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113849689B_ABST
    Figure CN113849689B_ABST
Patent Text Reader

Abstract

The present disclosure discloses an audio and video data processing method and device, equipment, medium and product, and relates to the technical field of voice. The audio and video data processing method comprises: processing audio and video data to obtain a first speech element set and first time information for the first speech element set, matching the first speech element set with a second speech element set, wherein the second speech element set is associated with text data; determining second time information for the text data based on the matching result between the first speech element set and the second speech element set and the first time information; and associatively outputting the text data and the audio and video data based on the second time information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, particularly to the field of voice technology, and more specifically, to an audio and video data processing method, apparatus, electronic device, medium, and program product. Background Technology

[0002] In audio and video processing scenarios, it is often necessary to add corresponding text to the audio and video, such as adding subtitles. However, current technologies for adding text to audio and video suffer from low matching accuracy, high manual labor costs, and cumbersome operations. Summary of the Invention

[0003] This disclosure provides an audio and video data processing method, apparatus, electronic device, storage medium, and program product.

[0004] According to one aspect of this disclosure, an audio and video data processing method is provided, comprising: processing audio and video data to obtain a first set of speech elements and first time information for the first set of speech elements; matching the first set of speech elements with a second set of speech elements, wherein the second set of speech elements is associated with text data; determining second time information for the text data based on the matching result between the first set of speech elements and the second set of speech elements and the first time information; and outputting the text data and the audio and video data in association based on the second time information.

[0005] According to another aspect of this disclosure, an audio / video data processing apparatus is provided, comprising: a processing module, a matching module, a determining module, and an output module. The processing module is used to process audio / video data to obtain a first set of speech elements and first time information for the first set of speech elements; the matching module is used to match the first set of speech elements with a second set of speech elements, wherein the second set of speech elements is associated with text data; the determining module is used to determine second time information for the text data based on the matching result between the first set of speech elements and the second set of speech elements and the first time information; and the output module is used to output the text data and the audio / video data in association based on the second time information.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the audio / video data processing method described above.

[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing the computer to perform the above-described audio and video data processing method.

[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described audio and video data processing method.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 The system architecture of an audio and video data processing method and apparatus according to an embodiment of the present disclosure is illustrated schematically;

[0012] Figure 2 A flowchart illustrating an audio / video data processing method according to an embodiment of the present disclosure is shown schematically.

[0013] Figure 3 The schematic diagram illustrates the principle of an audio / video data processing method according to an embodiment of the present disclosure;

[0014] Figures 4A-4B The illustration shows a schematic diagram of an audio / video data processing method according to an embodiment of the present disclosure;

[0015] Figure 5 A block diagram of an audio / video data processing apparatus according to an embodiment of the present disclosure is schematically shown; and

[0016] Figure 6 This is a block diagram of an electronic device for performing audio and video data processing, used to implement embodiments of the present disclosure. Detailed Implementation

[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0020] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).

[0021] This disclosure provides an audio / video data processing method. The method includes: processing audio / video data to obtain a first set of speech elements and first temporal information for the first set of speech elements; then matching the first set of speech elements with a second set of speech elements, the second set of speech elements being associated with text data; and determining second temporal information for the text data based on the matching result between the first and second sets of speech elements and the first temporal information; finally, outputting the text data and audio / video data in association based on the second temporal information.

[0022] Figure 1 The illustration schematically depicts the system architecture of an audio / video data processing method and apparatus according to an embodiment of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0023] like Figure 1 As shown, the system architecture 100 according to this embodiment may include clients 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between clients 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0024] Users can use clients 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on clients 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0025] Clients 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers. Clients 101, 102, and 103 in this embodiment of the disclosure can, for example, run applications.

[0026] Server 105 can be a server providing various services, such as a backend management server supporting websites browsed by users using clients 101, 102, and 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the clients. Alternatively, server 105 can also be a cloud server, meaning server 105 has cloud computing capabilities.

[0027] It should be noted that the audio and video data processing method provided in this embodiment can be executed by server 105. Correspondingly, the audio and video data processing device provided in this embodiment can be located in server 105. The audio and video data processing method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with clients 101, 102, 103 and / or server 105. Correspondingly, the audio and video data processing device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with clients 101, 102, 103 and / or server 105.

[0028] For example, audio / video data and text data can be sent through clients 101, 102, and 103. After server 105 receives the audio / video data and text data from clients 101, 102, and 103 through network 104, server 105 can obtain time information for the text data based on the audio / video data and text data, and output the text data and audio / video data in association based on the time information.

[0029] It should be understood that Figure 1 The number of clients, networks, and servers shown is merely illustrative. Depending on implementation needs, there can be any number of clients, networks, and servers.

[0030] This disclosure provides an audio and video data processing method, which will be described below in conjunction with... Figure 1 The system architecture, referencing Figures 2 to 4B This describes an audio / video data processing method according to exemplary embodiments of the present disclosure. The audio / video data processing method of the embodiments of the present disclosure may, for example, be derived from... Figure 1 The server 105 shown is used to execute this.

[0031] Figure 2 A flowchart illustrating an audio / video data processing method according to an embodiment of the present disclosure is shown.

[0032] like Figure 2 As shown, the audio and video data processing method 200 of this disclosure embodiment may include, for example, operations S210 to S240.

[0033] In operation S210, audio and video data are processed to obtain a first set of voice elements and first time information for the first set of voice elements.

[0034] In operation S220, the first set of speech elements is matched with the second set of speech elements, and the second set of speech elements is associated with the text data.

[0035] In operation S230, based on the matching result between the first set of speech elements and the second set of speech elements and the first time information, the second time information for the text data is determined.

[0036] During operation S240, text data and audio / video data are output in association based on the second time information.

[0037] For example, when editing audio and video data, users can upload audio and video data along with corresponding text data to associate and output the text data and audio and video data. For instance, speech recognition can be performed on the audio in the audio and video data to obtain a first set of speech elements. This first set of speech elements includes multiple first speech elements, each of which includes, for example, phonemes, such as vowels, consonants, etc. After obtaining the first set of speech elements, the time information of each first speech element appearing in the audio and video data is determined, and the time information corresponding to multiple first speech elements is determined as the first time information for the first set of speech elements.

[0038] Then, a second set of speech elements for the text data is obtained. This second set of speech elements includes, for example, multiple second speech elements, each of which includes, for example, phonemes, such as vowels, consonants, etc. For example, for each word in the text data, the phonemes of each word are determined to obtain the second set of speech elements.

[0039] After obtaining the first set of speech elements and the second set of speech elements, the first speech elements in the first set of speech elements and the second speech elements in the second set of speech elements can be matched to identify the second speech element that matches the first speech element. The first time information of the first speech element is then used as the time information of the second speech element to obtain the second time information of the text data.

[0040] The second time information, for example, indicates the time when the text data appears in the audio and video data. Therefore, based on the second time information, the text data and the audio and video data can be output in a related manner so that the text data can be output at the corresponding time when the audio and video are played, thus realizing the output of the text data as subtitle data of the audio and video data.

[0041] According to embodiments of this disclosure, a first set of voice elements is obtained by processing audio and video data. This first set of voice elements is then matched with a second set of voice elements for text data. Based on the matching result and first time information for the first set of voice elements, second time information for the text data is determined. Finally, the text data and audio / video data are output in association according to the second time information. Therefore, the technical solution of this disclosure enables the matching of corresponding subtitle data to audio and video data, improving the efficiency and accuracy of subtitle matching, reducing the manual costs required for subtitle matching, and simplifying the operation of subtitle matching.

[0042] Figure 3 The schematic diagram illustrates the principle of an audio / video data processing method according to an embodiment of the present disclosure.

[0043] like Figure 3 As shown, for audio / video data 310, multiple audio frames are extracted from the audio / video data 310, for example, n audio frames are extracted, where n is an integer greater than or equal to 1. Then, the multiple audio frames are processed to obtain multiple audio features corresponding one-to-one with the multiple audio frames. For example, feature extraction is performed on the n audio frames to obtain n audio features 320. Exemplarily, when processing the audio frames to obtain audio features, feature extraction can be performed using a pre-trained acoustic model. The acoustic model includes, for example, a time-delay neural network model. Then, based on the time information of the audio / video data, the time information of each audio frame is determined as the first time information for the first set of speech elements. For example, the time information corresponding one-to-one with the n audio features 320 is t1 to t2. n , t1~t n As the first piece of information. t1~t n Any one of them can be a moment or a time period.

[0044] For the text data, a second set of speech elements corresponding to the text data is determined. This second set of speech elements is represented, for example, by state diagram 330. State diagram 330 includes, for example, multiple second speech elements, such as "f", "u:", "b", "a:", etc. For example, at least one speech element corresponding to each character in the text data is determined sequentially, and state diagram 330 is obtained by arranging the speech elements corresponding to all characters in the text data sequentially.

[0045] For example, for multiple audio features 320, each audio feature can be identified by an acoustic model to obtain multiple first speech elements that correspond one-to-one with the multiple audio features 320, which are used as a set of first speech elements.

[0046] For example, for each audio feature 320, multiple candidate speech elements corresponding to the audio feature 320 and multiple target probabilities corresponding to the multiple candidate speech elements are determined. Each target probability represents the probability that the recognition result of the audio feature is the corresponding candidate speech element. For example, for the first audio feature 320, the acoustic model outputs four candidate speech elements "f", "u:", "b", and "a:" corresponding to the audio feature 320, and four probabilities of 0.7, 0.1, 0.1, and 0.1 corresponding to the four candidate speech elements "f", "u:", "b", and "a:".

[0047] In one example, the candidate speech element "f" corresponding to the highest probability can be used as the first speech element corresponding to the first audio feature 320, thus obtaining the first speech element corresponding to each audio feature. The first speech elements corresponding to multiple audio features are then used as a set of first speech elements.

[0048] For example, the multiple candidate speech elements for each audio feature 320 may include multiple second speech elements “f”, “u:”, “b”, “a:” in state diagram 330.

[0049] In another example, for each audio feature 320, based on multiple target probabilities and audio semantic information of that audio feature 320, a candidate speech element is determined from multiple candidate speech elements as the first speech element corresponding to that audio feature. The audio semantic information may be, for example, the contextual relationships in the audio / video data 310.

[0050] For example, taking the third audio feature 320 as an example, this audio feature 320 corresponds to four candidate speech elements "f", "u:", "b", and "a:", and four probabilities of 0.5, 0.4, 0.05, and 0.05 corresponding to these four candidate speech elements "f", "u:", "b", and "a:". Based on the probability and context, the candidate speech element "u:" corresponding to this audio feature 320 is determined as the first speech element corresponding to the third audio feature 320. It can be understood that this method comprehensively considers probability and context to determine the first speech element for each audio feature, making the recognition result of the first speech element more accurate. For example, if the context indicates that the pronunciation of the nearby audio is "fu", then when the first speech element corresponding to the second audio feature 320 is "f", the first speech element corresponding to the third audio feature 320 is likely to be "u:" based on the context.

[0051] In another example, regarding state diagram 330, each of the plurality of second speech elements includes at least one speech state. This embodiment of the disclosure uses an example where each second speech element includes three speech states. The three speech states corresponding to each second speech element can be different; different speech states may manifest in differences in sound velocity, timbre, pitch, etc. State diagram 330 is composed of the plurality of states corresponding to each second speech element.

[0052] The target probability corresponding to each audio feature 320 can include the probabilities of multiple speech states corresponding to each second speech element, i.e., each speech state corresponds to a target probability. Then, based on the target probability and audio semantic information, each first speech element in the first speech element set is matched with each speech state to obtain the matching result of each first speech element and speech state, thus matching each audio feature to the state graph 330 and obtaining the matching path 331.

[0053] Next, for each first speech element that matches a speech state, the first time information corresponding to the first speech element is determined as the time information for each speech state. For example, for the first speech element that matches the first speech state of the second speech element "f" (corresponding to the first audio feature), the first time information corresponding to the first speech element is t1, and the first time information t1 is determined as the time information for the first speech state of the second speech element "f".

[0054] For each speech state, the time information is determined as the time information of the second speech element corresponding to that speech state. For example, for the first audio feature 320 and the second audio feature 320, which correspond to the first and second speech states of the second speech element "f" respectively, the first time information of the first speech state of the second speech element "f" is t1, and the first time information of the second speech state is t2, which are used as the time information of the second speech element "f". It can be understood that matching speech elements is achieved by matching speech states, which improves the granularity of speech element matching and thus improves matching accuracy.

[0055] Then, based on the time information for each second speech element, second time information for the text data is determined. For example, after obtaining the second time information for the text data, the time period corresponding to each sentence in the text data is determined. For example, the text between every two punctuation marks is a sentence, and the punctuation marks can be commas, periods, etc. After obtaining the time period corresponding to each sentence, the corresponding sentence is output in each time period when outputting audio and video data, thus realizing the output of subtitle data in the audio and video data.

[0056] According to embodiments of this disclosure, multiple phonemes are obtained by processing audio and video data, and these phonemes are matched with multiple phonemes in text data. Based on the matching results, the time information corresponding to the phonemes in the audio and video data is used as the time information of the phonemes in the text data. This facilitates the association and output of text data and audio and video data based on the time information of the phonemes in the text data, thereby enabling the output of subtitle data within the audio and video data. It is evident that embodiments of this disclosure achieve the matching of corresponding subtitle data to audio and video data, improving the efficiency and accuracy of subtitle matching, reducing the manual costs required for subtitle matching, and simplifying the operation of subtitle matching.

[0057] Figures 4A-4B The illustration shows a schematic diagram of an audio / video data processing method according to an embodiment of the present disclosure.

[0058] like Figures 4A-4BAs shown, when a user edits audio and video through the client, they can upload audio and video data 410 to the application and import text data 420 corresponding to the audio and video data 410 into the application. The application sends the audio and video data 410 and text data 420 to the server for processing. The server can execute the method described above to obtain second time information for the text data 420. Then, based on the second time information, the audio and video data 410 and text data 420 are output in association. The output result 430 includes using the text data 420 as subtitle data for the audio and video data 410, and the output result 430 can be displayed on the client. For example, when the audio and video plays to the scene of "Hello everyone," the text "Hello everyone" is displayed as a subtitle.

[0059] Figure 5 A block diagram of an audio / video data processing apparatus according to an embodiment of the present disclosure is shown schematically.

[0060] like Figure 5 As shown, the audio and video data processing apparatus 500 of this embodiment includes, for example, a processing module 510, a matching module 520, a determining module 530, and an output module 540.

[0061] Processing module 510 can be used to process audio and video data to obtain a first set of voice elements and first time information for the first set of voice elements. According to embodiments of this disclosure, processing module 510 can, for example, execute the functions described above. Figure 2 The operation S210 described herein will not be repeated here.

[0062] The matching module 520 can be used to match a first set of speech elements with a second set of speech elements, wherein the second set of speech elements is associated with text data. According to embodiments of this disclosure, the matching module 520 can, for example, perform the operations described above. Figure 2 The operation S220 described herein will not be repeated here.

[0063] The determining module 530 can be used to determine second time information for text data based on the matching result between the first set of speech elements and the second set of speech elements and the first time information. According to embodiments of this disclosure, the determining module 530 can, for example, perform the above-mentioned reference... Figure 2 The operation S230 described herein will not be repeated here.

[0064] The output module 540 can be used to output text data and audio / video data in a correlated manner based on second time information. According to embodiments of this disclosure, the output module 540 can, for example, perform the functions described above. Figure 2 The operation S240 described will not be repeated here.

[0065] According to embodiments of this disclosure, the processing module 510 includes: an extraction submodule, a processing submodule, a first determining submodule, and a second determining submodule. The extraction submodule is used to extract multiple audio frames from the audio and video data; the processing submodule is used to process the multiple audio frames to obtain multiple audio features corresponding one-to-one with the multiple audio frames; the first determining submodule is used to determine multiple first speech elements corresponding one-to-one with the multiple audio features, as a first speech element set; the second determining submodule is used to determine the time information of each audio frame in the multiple audio frames as first time information based on the time information of the audio and video data.

[0066] According to embodiments of this disclosure, for each of a plurality of audio features, the first determining submodule includes: a first determining unit and a second determining unit. The first determining unit is configured to determine a plurality of candidate speech elements corresponding to the audio feature and a plurality of target probabilities corresponding to the plurality of candidate speech elements, wherein each target probability represents the probability that the recognition result of the audio feature is the corresponding candidate speech element; the second determining unit is configured to determine one candidate speech element from the plurality of candidate speech elements based on the plurality of target probabilities and audio semantic information, as the first speech element corresponding to the audio feature.

[0067] According to an embodiment of this disclosure, the second voice element set includes a plurality of second voice elements, and each of the plurality of second voice elements includes at least one voice state; the matching module 520 is further configured to: match each first voice element in the first voice element set with each voice state.

[0068] According to embodiments of this disclosure, the determining module 530 includes: a third determining submodule, a fourth determining submodule, and a fifth determining submodule. The third determining submodule is used to determine, for each first voice element matching a voice state, the first time information corresponding to the first voice element as time information for each voice state; the fourth determining submodule is used to determine, for each voice state, the time information as time information for a second voice element corresponding to the voice state; and the fifth determining submodule is used to determine, based on the time information for the second voice element, second time information for the text data.

[0069] According to an embodiment of this disclosure, the output module 540 is further configured to: output text data as subtitle data of audio and video data based on second time information.

[0070] According to embodiments of this disclosure, the first speech element in the first speech element set includes phonemes, and the second speech element in the second speech element set includes phonemes.

[0071] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0072] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0073] Figure 6 This is a block diagram of an electronic device for performing audio and video data processing, used to implement embodiments of the present disclosure.

[0074] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0075] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0076] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0077] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as audio and video data processing methods. For example, in some embodiments, the audio and video data processing methods may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the audio and video data processing methods described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform audio and video data processing methods by any other suitable means (e.g., by means of firmware).

[0078] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0079] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable audio / video data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0080] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0081] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0082] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0083] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0084] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0085] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for processing audio-video data, comprising: processing audio-video data to obtain a first set of speech elements and first time information corresponding to the first set of speech elements; matching the first set of speech elements with a second set of speech elements, wherein the second set of speech elements is associated with text data; the second set of speech elements comprises a plurality of second speech elements, each of the plurality of second speech elements comprises a different speech state; different speech states differ in speed, tone or pitch; the processing of the audio-video data to obtain the first set of speech elements and the first time information corresponding to the first set of speech elements comprises: extracting a plurality of audio frames from the audio-video data; processing the plurality of audio frames to obtain a plurality of audio features corresponding to the plurality of audio frames one-to-one; any audio frame corresponds to a single audio feature; determining a plurality of first speech elements corresponding to the plurality of audio features one-to-one as the first set of speech elements; any audio feature corresponds to a single first speech element; determining time information of each audio frame in the plurality of audio frames as the first time information according to time information of the audio-video data; the matching of the first set of speech elements with the second set of speech elements comprises: based on probabilities that audio features of first speech elements respectively correspond to speech states of the plurality of second speech elements and audio semantic information, matching each first speech element in the first set of speech elements with each speech state of the plurality of second speech elements respectively; determining second time information corresponding to the text data based on a matching result between the first set of speech elements and the second set of speech elements and the first time information; and based on the second time information, associatively outputting the text data and the audio-video data.

2. The method of claim 1, wherein, the determining of the plurality of first speech elements corresponding to the plurality of audio features one-to-one as the first set of speech elements comprises, for each audio feature in the plurality of audio features: determining a plurality of candidate speech elements corresponding to the audio feature and a plurality of target probabilities corresponding to the plurality of candidate speech elements, wherein each target probability in the plurality of target probabilities represents a probability that an identification result of the audio feature is a corresponding candidate speech element; and based on the plurality of target probabilities and audio semantic information, determining a candidate speech element from the plurality of candidate speech elements as a first speech element corresponding to the audio feature.

3. The method of claim 1, wherein, the determining of the second time information corresponding to the text data based on the matching result between the first set of speech elements and the second set of speech elements and the first time information comprises: for each first speech element matched with a speech state, determining first time information corresponding to the first speech element as time information corresponding to the speech state; determining the time information corresponding to each speech state as time information of a second speech element corresponding to the speech state; and determine second time information for the text data based on the time information for the second speech elements.

4. The method of any of claims 1-3, wherein, the text data and the audio-video data are associatedly output based on the second time information. the text data is output as subtitle data of the audio-video data based on the second time information.

5. The method of any of claims 1-3, wherein, the first speech elements in the first set of speech elements comprise phonemes, and the second speech elements in the second set of speech elements comprise phonemes. 6.An audio-video data processing apparatus, comprising: a processing module configured to process audio-video data to obtain a first set of speech elements and first time information for the first set of speech elements; a matching module configured to match the first set of speech elements with a second set of speech elements, wherein the second set of speech elements is associated with text data, the second set of speech elements comprises a plurality of second speech elements, and each second speech element in the plurality of second speech elements comprises a different speech state, wherein the different speech states are different in terms of speed, tone or pitch; the matching module is further configured to match each first speech element in the first set of speech elements with each speech state of the plurality of second speech elements based on probabilities that audio features of the first speech element correspond to the speech states of the plurality of second speech elements and audio semantic information; a determining module configured to determine second time information for the text data based on a matching result between the first set of speech elements and the second set of speech elements and the first time information; and an output module configured to associatedly output the text data and the audio-video data based on the second time information. the processing module comprises an extracting sub-module configured to extract a plurality of audio frames from the audio-video data, a processing sub-module configured to process the plurality of audio frames to obtain a plurality of audio features corresponding to the plurality of audio frames, wherein any audio frame corresponds to a single audio feature, a first determining sub-module configured to determine a plurality of first speech elements corresponding to the plurality of audio features as the first set of speech elements, wherein any audio feature corresponds to a single first speech element, and a second determining sub-module configured to determine time information of each audio frame in the plurality of audio frames as the first time information according to time information of the audio-video data.

7. The apparatus of claim 6, wherein, for each audio feature in the plurality of audio features, the first determining sub-module comprises: a first determining unit configured to determine a plurality of candidate speech elements corresponding to the audio feature and a plurality of target probabilities corresponding to the plurality of candidate speech elements, wherein each target probability in the plurality of target probabilities represents a probability that an identification result of the audio feature is a corresponding candidate speech element; and a second determining unit configured to determine a candidate speech element from the plurality of candidate speech elements as the first speech element corresponding to the audio feature based on the plurality of target probabilities and audio semantic information.

8. The apparatus of claim 6, wherein, the determining module comprises: a third determining sub-module, configured to determine, for each speech state, time information corresponding to a first speech element matched with the speech state, as time information of the speech state; a fourth determining sub-module, configured to determine, for each speech state, the time information of the speech state, as time information of a second speech element corresponding to the speech state; and a fifth determining sub-module, configured to determine, based on the time information of the second speech element, second time information of the text data.

9. The apparatus of any of claims 6-8, wherein, The output module is further configured to: output, based on the second time information, the text data as subtitle data of the audio and video data.

10. The apparatus of any of claims 6-8, wherein, The first speech element in the first speech element set comprises a phoneme, and the second speech element in the second speech element set comprises a phoneme. 11.An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

12. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-5. 13.A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Information processing method and device, computer equipment and storage medium

    CN112837401A

  • System and method for speech-based audio and text alignment

    CN113112996A

  • Methods and apparatus for automatic speech recognition

    CN1783213A