ASR training data acquisition method and system, storage medium and electronic equipment

By identifying the matching of video subtitle coordinates and audio files, the problem of cumbersome and inefficient acquisition of ASR training data is solved, and the rapid and intelligent acquisition of training data is achieved, which improves the convenience and accuracy of ASR training data.

CN120260550APending Publication Date: 2025-07-04SHANGHAI MIDU DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510566610.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing ASR training data acquisition methods are cumbersome and inefficient, and lack automation.

Method used

By identifying the video subtitle coordinates, disassembly frames to obtain the frame sequence number, and combining the visual language model to update the coordinates, perform optical character recognition, calculate the subtitle frame attributes, and match the voice in the audio file to obtain training data.

Benefits of technology

It realizes fast and intelligent subtitle content recognition and matching, improving the convenience and accuracy of ASR training data acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260550A_ABST
    Figure CN120260550A_ABST
Patent Text Reader

Abstract

The invention provides an ASR training data acquisition method and system, a storage medium and electronic equipment. The method comprises the following steps: recognizing subtitle coordinates and reading an audio file based on a target video; performing identification based on the caption coordinates to obtain caption contents; calculating the frame attribution of the subtitle content based on the audio file in combination with the subtitle content; and obtaining a target voice in the audio file, and extracting a correspondingly matched target subtitle based on the frame attribution to obtain ASR training data. According to the ASR training data acquisition method and system, the storage medium and the electronic equipment, the subtitle content in the video can be quickly recognized, the corresponding subtitle content can be matched according to the voice of the video, the overall scheme is high in intelligent degree and high in practicability, the convenience degree of ASR training data acquisition is improved, and the accuracy rate is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video image processing, and particularly relates to a method, a system, a storage medium and an electronic device for obtaining ASR training data. Background Art

[0002] Speech recognition technology, also known as Automatic Speech Recognition (ASR), aims to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0003] In actual use, model training needs to be carried out according to training data. However, the current ways to obtain training data can be through direct interaction between users and devices using voice or oral speech, or manual input in advance. The existing ways to obtain training data have the following deficiencies:

[0004] (1) The acquisition method is cumbersome and not simple enough;

[0005] (2) The acquisition efficiency is low and it lacks automation. Summary of the Invention

[0006] In view of the above-mentioned deficiencies of the prior art, the purpose of the present invention is to provide a method, a system, a storage medium and an electronic device for obtaining ASR training data, which are used to solve the problem of low efficiency and lack of intelligence in obtaining ASR training data.

[0007] In a first aspect, the present invention provides a method for obtaining ASR training data, and the method includes the following steps:

[0008] Identifying subtitle coordinates based on a target video and reading an audio file;

[0009] Identifying subtitle content based on the subtitle coordinates;

[0010] Calculating the frame attribution of the subtitle content based on the audio file in combination with the subtitle content;

[0011] Obtaining target speech in the audio file, and extracting corresponding matching target subtitles based on the frame attribution to obtain ASR training data.

[0012] In a possible implementation manner of the present application, the identifying subtitle coordinates based on a target video specifically includes:

[0013] Splitting a target video into frames to obtain frame numbers;

[0014] Combining the frame numbers and preset characters to obtain arranged input data;

[0015] The coordinate position of the arranged input data is iteratively updated using a preset visual language model until the last frame of the target video is reached to obtain the subtitle coordinates.

[0016] In a possible implementation of the present application, the target video is deframed based on a preset sampling rate, wherein the sampling rate includes 30 frames per second, the preset characters include image characters and position characters, and the arrangement input data includes image characters, a first frame of video, position characters, coordinate positions, image characters, a second frame of video, and an arrangement order of position characters.

[0017] In a possible implementation of the present application, obtaining the subtitle content based on the subtitle coordinate recognition specifically includes:

[0018] Based on the subtitle coordinates, a picture of a frame subtitle area in the target video is intercepted to obtain a target picture;

[0019] Optical character recognition is performed on the target image to obtain the subtitle content.

[0020] In a possible implementation of the present application, the calculating the frame attribution of the subtitle content based on the audio file in combination with the subtitle content specifically includes:

[0021] Calculate the error rate of subtitles in two adjacent frames;

[0022] The size of the preset ratio is compared based on the word error rate, wherein if the word error rate is smaller than the preset ratio, it is determined that the current two adjacent frames are the same segment of speech; otherwise, they are different segments of speech.

[0023] In a possible implementation of the present application, obtaining the target speech in the audio file and extracting the corresponding matching target subtitles based on the frame attribution to obtain ASR training data specifically includes:

[0024] Acquire the target speech based on the audio file combined with all speech time periods;

[0025] Extracting the subtitles of the corresponding frames based on different speech time periods to obtain the target subtitles;

[0026] The ASR training data is obtained based on the target speech combined with the corresponding target subtitles.

[0027] In a possible implementation of the present application, the method further includes acquiring the target video, and obtaining the target video based on input data and / or response data, wherein the input data includes video data input by a user terminal.

[0028] In a second aspect, the present invention provides an ASR training data acquisition system, the system comprising:

[0029] A reading module, configured to identify subtitle coordinates based on a target video and read an audio file;

[0030] An identification module, configured to identify subtitle content based on the subtitle coordinates;

[0031] A calculation module, configured to calculate the frame attribution of the subtitle content based on the audio file in combination with the subtitle content;

[0032] An extraction module, configured to obtain target speech in the audio file, and extract corresponding matching target subtitles based on the frame attribution to obtain ASR training data.

[0033] In a third aspect, the present invention provides an electronic device, which includes: a processor and a memory;

[0034] The memory is used to store a computer program;

[0035] The processor is configured to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned ASR training data acquisition method.

[0036] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by an electronic device, the above-mentioned ASR training data acquisition method is implemented.

[0037] As described above, the ASR training data acquisition method, system, storage medium and electronic device of the present invention have the following beneficial effects:

[0038] (1) It can quickly identify subtitle content in a video;

[0039] (2) It can match corresponding subtitle content according to the speech in the video;

[0040] (3) It has a high degree of intelligence, great practicality, improves the convenience of obtaining ASR training data, and has a high accuracy rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It shows a schematic diagram of a scenario of the electronic device of the present invention in an embodiment;

[0042] Figure 2 It shows a flowchart of the ASR training data acquisition method of the present invention in an embodiment;

[0043] Figure 3 It shows a schematic diagram of steps of the ASR training data acquisition method of the present invention in an embodiment;

[0044] Figure 4Shown is a schematic structural diagram of the ASR training data acquisition system according to an embodiment of the present invention;

[0045] Figure 5 Shown is a schematic structural diagram of the electronic device according to an embodiment of the present invention.

[0046] Description of component numbers

[0047] Steps S202 to S208 40 ASR training data acquisition system

[0048] 41 Reading module

[0049] 42 Recognition module

[0050] 43 Calculation module

[0051] 44 Extraction module Detailed implementation manners

[0052] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0053] It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0054] The following embodiments of the present invention provide an ASR training data acquisition method, which can be applied to an electronic device as shown in Figure 1 The electronic device described in the present invention may include a mobile phone 11 with a wireless charging function, a tablet computer 12, a notebook computer 13, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The specific type of the electronic device in the embodiments of the present invention is not limited in any way.

[0055] For example, the electronic device may be a station (STAION, ST) in a WLAN with wireless charging function, a cellular phone with wireless charging function, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless charging function, a computing device or other processing device, a computer, a laptop computer, a handheld communication device, a handheld computing device, and / or other devices for communicating on a wireless system, as well as next-generation communication systems, such as mobile terminals in a 5G network, mobile terminals in a future evolved Public Land Mobile Network (PLMN), or mobile terminals in a future evolved Non-terrestrial Network (NTN), etc.

[0056] For example, the electronic device may communicate with the network and other devices through wireless communication. The above wireless communication may use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).

[0057] The technical solutions in the embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0058] Specifically, please refer to Figure 2 , in an embodiment of the invention, the method for obtaining ASR training data of the present invention includes the following steps:

[0059] Step S202, identifying subtitle coordinates based on a target video and reading an audio file;

[0060] Step S204, identifying subtitle content based on the subtitle coordinates;

[0061] Step S206, calculating the frame attribution of the subtitle content based on the audio file in combination with the subtitle content;

[0062] Step S208, obtaining the target speech in the audio file, and extracting the corresponding matching target subtitles based on the frame attribution to obtain ASR training data.

[0063] It should be noted that in this embodiment, a target video needs to be determined first. Among them, the target video can be obtained based on input data and / or response data. The input data includes video data input by the user terminal. Specifically, the input data can be input by the user terminal, and the response data can be automatically obtained by a program in the electronic device to call the response to obtain the corresponding video.

[0064] Further, in an embodiment of the invention, after the target video is obtained, subtitle coordinates are identified based on the target video. Specifically, it includes: splitting the target video into frames to obtain frame numbers; combining the frame numbers and preset characters to obtain arranged input data; using a preset vision-language model to iteratively update the coordinate positions of the arranged input data until the last frame of the target video is reached to obtain the subtitle coordinates.

[0065] It should be noted that in this embodiment, the target video is split into frames based on a preset sampling rate to obtain corresponding frame numbers, such as the first frame, the second frame, etc. Correspondingly, the sampling rate includes 30 frames per second, and the preset characters include image characters and position characters <loc>, correspondingly, the arranged input data is obtained by combining the frame sequence number and the preset characters, where the arranged input data includes image characters , the first frame of video (with subtitles), and position characters <loc>, coordinate position, image character , the second frame of video, and position character <loc>The arrangement order is used, and finally, the preset vision-language model is used to iteratively update the coordinate positions of the arranged input data until the last frame of the target video is reached to obtain the subtitle coordinates. The vision-language model is the VLM (Vision-Language Models) model. The subtitle coordinates in different frames of the video are obtained through iterative update, and finally, the coordinates where the subtitles appear in all frames are obtained to get the subtitle coordinates.

[0066] Further, in an embodiment of the invention, the obtaining the subtitle content based on the subtitle coordinates specifically includes: intercepting a picture of the frame subtitle area in the target video based on the subtitle coordinates to obtain a target picture; performing optical character recognition on the target picture to obtain the subtitle content.

[0067] It should be noted that in this embodiment, after obtaining the subtitle coordinates, it is necessary to perform character recognition on them to obtain the subtitle content. Specifically, a picture of the frame subtitle area in the target video is intercepted based on the subtitle coordinates to obtain a target picture, and then optical character recognition is performed on the target picture to obtain the subtitle content. Optical character recognition (OCR, Optical Character Recognition) is a computer input technology that uses character recognition technology to convert image information into usable text, so that the subtitle content can be recognized from the target picture. Since the OCR technology is a prior art, the specific recognition process is not described in detail in this embodiment and is only used as an application to obtain the subtitle content.

[0068] Further, in an embodiment of the invention, the calculating the frame attribution of the subtitle content based on the audio file and the subtitle content specifically includes: calculating the character error rate of the subtitles in two adjacent frames; comparing the size of a preset ratio based on the character error rate. If the character error rate is less than the preset ratio, it is determined that the current two adjacent frames are of the same segment of speech; otherwise, they are of different segments of speech.

[0069] It should be noted that the above embodiments illustrate the acquisition of subtitle coordinates and the recognition of subtitle content. In this embodiment, it specifically illustrates the calculation of frame attribution for different frames to determine whether the subtitles of adjacent frames are of the same segment of speech. Specifically, the speech belongingness is judged by calculating the character error rate (CER, Character Error Rate). Among them, CER specifically calculates the ratio of the number of misrecognized characters to the total number of characters. The lower the CER, the higher the recognition accuracy. Specifically, the character error rate of the subtitles in two adjacent frames is calculated, and the size of a preset ratio is compared based on the character error rate. The preset ratio is taken as "0.05". If the character error rate is less than the preset ratio, it is determined that the current two adjacent frames are of the same segment of speech; otherwise, they are of different segments of speech. And each frame of subtitle of the same segment of speech is corresponding to this segment of speech.

[0070] Further, in an embodiment of the invention, obtaining the target speech in the audio file and extracting the corresponding and matching target subtitles based on the frame attribution to obtain ASR training data specifically includes the following steps: obtaining the target speech based on the audio file in combination with all speech time periods; extracting the subtitles of corresponding frames based on different speech time periods to obtain the target subtitles; and obtaining the ASR training data based on the target speech in combination with the corresponding target subtitles.

[0071] It should be noted that in this embodiment, the audio file read based on the target video needs to obtain the target speech according to all speech time periods, and different subtitles correspond to different target speeches. Therefore, it is necessary to extract the subtitles of corresponding frames based on different speech time periods to obtain the target subtitles. After obtaining the target speech and the matching target subtitles, the ASR training data can be obtained.

[0072] Further, as Figure 3 shown, it is a schematic flowchart of the ASR training data acquisition method of the present application, which is divided into four steps, corresponding to the first step: finding the positions of all frame subtitles, that is, corresponding to obtaining subtitle coordinates, which specifically includes: (1) splitting the video frames: ① splitting the frames at a sampling rate of 30 frames per second; ② determining the start time point and end time point of the frame appearance through the sampling rate and the frame number. For example, the start time point of the first frame is "0", and the end time point is "1 / 30". Another example is that the start time point of the second frame is "1 / 30", and the end time point is "2 / 30". It can be expressed that the start time point of the Nth frame is (N - 1) × 1 / 30, and the end time point is N × 1 / 30. It should be noted that: the unit of the time point is seconds, × represents multiplication, and / represents division; (2) setting 2 special characters , <loc>, corresponding to preset characters, respectively representing an image and a position; (3) According to , the first frame (including subtitles), <loc>, coordinate position, , second frame, <loc>In the order, they are arranged as a whole input, received by the VLM (Visual Language Model), and the coordinate positions where the subtitles in the second frame appear are given; (4) The specific form of the coordinate positions is [upper left corner of the x-axis, upper left corner of the y-axis, lower right corner of the x-axis, lower right corner of the y-axis]; (5) After obtaining the coordinate positions of the second frame, rearrange the input of the VLM, replace the first frame (including subtitles) with the second frame, and replace the coordinates with the coordinates output by the VLM to obtain a new round of input; (6) Repeat obtaining the output coordinates and repeating the replacement of the input until the last frame, and the coordinate positions where the subtitles appear in all frames can be obtained, corresponding to iteratively updating the subtitle coordinates.

[0073] The second step corresponds to: identifying the content of the subtitles, that is, identifying the subtitle content based on the subtitle coordinates, specifically including: (1) Through the coordinate positions, intercept the picture of the frame subtitle area, that is, the corresponding target picture; (2) For the picture of the subtitle area, through OCR, obtain the specific content of the subtitles, that is, perform optical character recognition on the target picture to obtain the subtitle content.

[0074] The third step corresponds to: finding all the time periods of the speech, that is, corresponding to calculating the frame attribution of the subtitle content, specifically including: (1) For the first-frame subtitle and the second-frame subtitle, calculate the CER error rate of misspelled words. If it is less than "0.05", it is considered that these two frames belong to the same speech segment, and the start time point of this speech segment is the start time point of the first frame. Then calculate the error rate of misspelled words between the second frame and the third frame. If it is still less than "0.05", continue to calculate the error rate of misspelled words between the third frame and the fourth frame until the error rate of misspelled words between the (M - 1)-th frame and the M-th frame is greater than or equal to "0.05". Then the end time point of this speech segment is the end time point of the (M - 1)-th frame, and the subtitle of this speech segment is the first-frame subtitle, so as to obtain the subtitle corresponding to the same speech segment; (2) After finding the end time point of a speech segment, start to find the start time point and end time point of the next speech segment. The start time point of this speech segment is the start time point of the M-th frame, and then continue to search in the manner of step (1) until the corresponding speech segment and its subtitle content are found.

[0075] The fourth step corresponds to: determining the training data, specifically including (1) Reading the audio file in the video, that is, reading the audio file based on the target video and intercepting all the speech segments according to all the speech time periods; (2) Finding the speech segments and their corresponding subtitle content from the third step, and then constructing the required ASR training data set from the speech and subtitles.

[0076] The embodiments of the present application further provide an ASR training data acquisition system. The ASR training data acquisition system can implement the ASR training data acquisition method described in the present application. However, the implementation devices of the ASR training data acquisition method described in the present application include, but are not limited to, the structures of the ASR training data acquisition systems listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principles of the present application are included in the protection scope of the present application.

[0077] Please refer to Figure 4 , in one embodiment, an ASR training data acquisition system 40 provided in this embodiment, the system includes:

[0078] A reading module 41, configured to identify subtitle coordinates based on a target video and read an audio file;

[0079] An identification module 42, configured to identify subtitle content based on the subtitle coordinates;

[0080] A calculation module 43, configured to calculate the frame attribution of the subtitle content based on the audio file in combination with the subtitle content;

[0081] An extraction module 44, configured to obtain a target voice in the audio file, and extract a corresponding matching target subtitle based on the frame attribution to obtain ASR training data.

[0082] Since the specific implementation manner of this embodiment corresponds to the foregoing method embodiment, the same details will not be repeated here. Those skilled in the art should also understand that Figure 4 The division of each module in the embodiment is only a logical function division. In actual implementation, it can be fully or partially integrated into one or more physical entities, and these modules can all be implemented in the form of software called by a processing element, or all in the form of hardware, or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware.

[0083] In several embodiments provided by the present invention, it should be understood that the disclosed system, device or method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules or units can be in an electrical, mechanical or other form.

[0084] The modules / units described as separate components may or may not be physically separated, and the components shown as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the objectives of the embodiments of the present invention. For example, in various embodiments of the present invention, the functional modules / units can be integrated in one processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.

[0085] Those of ordinary skill in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0086] The embodiments of the present invention also provide a computer-readable storage medium. Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing a processor through a program, and the program can be stored in a computer-readable storage medium. The storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0087] Embodiments of the present application may further provide a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, or data center to another website, computer, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.).

[0088] When the computer program product is executed by a computer, the computer executes the method described in the foregoing method embodiments. The computer program product may be a software installation package. In the case where the foregoing method is required, the computer program product may be downloaded and executed on the computer.

[0089] The descriptions of the processes or structures corresponding to the above respective drawings have their own emphases. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0090] Embodiments of the present invention further provide an electronic device. The electronic device includes a processor and a memory.

[0091] The memory is used to store a computer program.

[0092] The memory includes various media that can store program codes, such as ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disc.

[0093] The processor is connected to the memory and is used to execute the computer program stored in the memory, so that the electronic device executes the above ASR training data acquisition method.

[0094] Preferably, the processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0095] As shown Figure 5 in FIG. 1, the electronic device of the present invention is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors or processing units 51, a memory 52, and a bus 53 that connects different system components (including the memory 52 and the processing unit 51).

[0096] The bus 53 represents one or more of several types of bus architectures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0097] The electronic device typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.

[0098] The memory 52 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 521 and / or cache memory 522. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 523 may be used for reading and writing non-removable, non-volatile magnetic media ( Figure 5 not shown, typically referred to as a "hard disk drive"). Although Figure 5 not shown in FIG. 1, a disk drive for reading and writing removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing removable non-volatile optical disks (such as CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 53 through one or more data media interfaces. The memory 52 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.

[0099] A program / utility 524 having a set (at least one) of program modules 5241 may be stored, for example, in the memory 52. Such program modules 5241 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 5241 generally perform the functions and / or methods described in the embodiments of the present invention.

[0100] The electronic device can also communicate with one or more external devices (such as keyboards, pointing devices, displays, etc.), and can also communicate with one or more devices that enable users to interact with the electronic device, and / or communicate with any device that enables the electronic device to communicate with one or more other computing devices (such as network cards, modems, etc.). This communication can be carried out through the input / output (I / O) interface 54. Moreover, the electronic device can also communicate with one or more networks (such as local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through the network adapter 55. As Figure 5 shown, the network adapter 55 communicates with other modules of the electronic device through the bus 53. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0101] The above embodiments are only illustrative of the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.< / loc> < / loc> < / loc> < / loc> < / loc> < / loc>

Claims

1. A method for obtaining ASR training data, characterized in that, include: Identify subtitle coordinates based on the target video and read audio files; Obtaining subtitle content based on the subtitle coordinate recognition; Calculate the frame attribution of the subtitle content based on the audio file combined with the subtitle content; The target speech in the audio file is obtained, and the corresponding matching target subtitles are extracted based on the frame attribution to obtain ASR training data.

2. The ASR training data acquisition method according to claim 1, wherein The identifying subtitle coordinates based on the target video specifically includes: Deframe the target video to obtain the frame number; Combining the frame number and the preset character to obtain the arranged input data; The coordinate position of the arranged input data is iteratively updated using a preset visual language model until the last frame of the target video is reached to obtain the subtitle coordinates.

3. The ASR training data acquisition method according to claim 2, wherein The target video is deframed based on a preset sampling rate, wherein the sampling rate includes 30 frames per second, the preset characters include image characters and position characters, and the arrangement input data includes image characters, a first frame of video, position characters, coordinate positions, image characters, a second frame of video, and an arrangement order of position characters.

4. The ASR training data acquisition method according to claim 3, characterized in that The obtaining of subtitle content based on the subtitle coordinate identification specifically includes: Based on the subtitle coordinates, a picture of a frame subtitle area in the target video is intercepted to obtain a target picture; Optical character recognition is performed on the target image to obtain the subtitle content.

5. The method for obtaining ASR training data according to claim 4, wherein The calculating the frame attribution of the subtitle content based on the audio file in combination with the subtitle content specifically includes: Calculate the error rate of subtitles in two adjacent frames; The size of the preset ratio is compared based on the word error rate, wherein if the word error rate is smaller than the preset ratio, it is determined that the current two adjacent frames are the same segment of speech; otherwise, they are different segments of speech.

6. The ASR training data acquisition method according to claim 5, wherein The step of obtaining the target speech in the audio file and extracting the corresponding matching target subtitles based on the frame attribution to obtain ASR training data specifically includes: Acquire the target speech based on the audio file combined with all speech time periods; Extracting the subtitles of the corresponding frames based on different speech time periods to obtain the target subtitles; The ASR training data is obtained based on the target speech combined with the corresponding target subtitles.

7. The ASR training data acquisition method according to claim 1, characterized in that The method further includes acquiring the target video, and obtaining the target video based on input data and / or response data, wherein the input data includes video data input by a user end.

8. An ASR training data acquisition system, characterized in that, include: A reading module, used to identify subtitle coordinates based on the target video and read audio files; An identification module, used for identifying subtitle content based on the subtitle coordinates; A calculation module, used for calculating the frame attribution of the subtitle content based on the audio file in combination with the subtitle content; The extraction module is used to obtain the target speech in the audio file, and extract the corresponding matching target subtitles based on the frame attribution to obtain ASR training data.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the ASR training data acquisition method described in any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that, The electronic device comprises: a processor and a memory; wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the ASR training data acquisition method as described in any one of claims 1 to 7.