Video distribution methods, devices, readable media and electronic devices

By calculating the confidence and probability values ​​of language tags for non-silent audio segments in videos, the problem of insufficient language probability assessment in video distribution is solved, enabling accurate video distribution and filtering.

CN115035914BActive Publication Date: 2025-10-31BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210621894.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-10-31
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

In existing technologies, the video distribution process lacks probability assessment for each language in the video, making it impossible to effectively distribute videos based on language.

Method used

By acquiring the speech recognition results of the non-silent audio segments of the video, calculating the language label confidence score of each segment, and calculating the language label probability value of the video based on the confidence score, the distribution or discard of the video is determined.

Benefits of technology

It enables precise distribution or filtering of videos based on language probability values, improving the accuracy and efficiency of video distribution and meeting the needs of different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035914B_ABST
    Figure CN115035914B_ABST
Patent Text Reader

Abstract

This disclosure relates to a video distribution method, apparatus, readable medium, and electronic device. The method includes: obtaining a confidence level of a language tag for each non-silent audio segment based on the confidence level of each character identified in that segment; obtaining a probability value of a language tag for the video based on the confidence level of the language tag for each non-silent audio segment; and processing the video based on the probability value of the video's language tag. Through this technical solution, during video distribution, only the probability value requirement for each language for each service provider / user needs to be set. Videos meeting the requirements can be sent to the corresponding target user, or videos that do not meet the requirements for any user can be discarded, facilitating video distribution based on language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video application technology, and more specifically, to a video distribution method, apparatus, readable medium, and electronic device. Background Technology

[0002] In video scenarios, long videos are usually segmented, and each audio segment is fed into a speech recognition system to obtain the language label, transcription result and confidence score of each audio segment. However, there is no probability for each language label of the entire video, that is, there is no video-level probability for each language. This is not conducive to distributing videos according to language. Summary of the Invention

[0003] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, this disclosure provides a video distribution method, the method comprising:

[0005] Obtain the speech recognition results of all non-silent audio segments in the video. The speech recognition results of each non-silent audio segment include: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment.

[0006] The confidence level of the language label for each non-silent audio segment is obtained based on the confidence level of each character identified in each non-silent audio segment.

[0007] Based on the confidence level of the language tag for each non-silent audio segment, the probability value of the language tag for the video is obtained;

[0008] The video is processed based on the probability value of its language tag, wherein the processing includes distributing the video to a target user or discarding the video.

[0009] Secondly, this disclosure provides a video distribution device, comprising:

[0010] The speech recognition module is used to obtain the speech recognition results of all non-silent audio segments of the video. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment.

[0011] The segment confidence acquisition module is used to obtain the confidence of the language tag of each non-silent audio segment based on the confidence of each character identified in each non-silent audio segment.

[0012] The language probability acquisition module is used to obtain the probability value of the language tag of the video based on the confidence level of the language tag of each non-silent audio segment;

[0013] The processing module is configured to process the video based on the probability value of the video's language tag, wherein the processing includes distributing the video to a target user or discarding the video.

[0014] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described method.

[0015] Fourthly, this disclosure provides an electronic device, comprising:

[0016] A storage device having at least one computer program stored thereon;

[0017] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the above method.

[0018] Using the above technical solution, the probability value of each video being in each language can be obtained. During video distribution, it is only necessary to set the probability value requirement for each business party / user for each language, and videos that meet the requirements can be sent to the corresponding business party / user (target user), or videos that do not meet any business party / user requirements can be discarded, facilitating video distribution based on language.

[0019] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0021] Figure 1 This is a flowchart of a video distribution method provided according to one embodiment of the present disclosure.

[0022] Figure 2 This is a schematic diagram of the structure of a speech recognition model provided according to one embodiment of the present disclosure.

[0023] Figure 3 This is a block diagram of a video distribution apparatus provided according to one embodiment of the present disclosure.

[0024] Figure 4This is a schematic diagram of the structure of an electronic device provided according to one embodiment of the present disclosure. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0031] All actions involving the acquisition of signals, information, or data in this disclosure are carried out in accordance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.

[0032] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0033] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0034] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0035] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0036] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0037] Figure 1 This is a flowchart of a video distribution method provided according to one embodiment of the present disclosure, such as... Figure 1 As shown, the method provided in this disclosure may include steps 11 to 14.

[0038] In step S11, the speech recognition results of all non-silent audio segments of the video are obtained.

[0039] The speech recognition result for each non-silent audio segment includes: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment. The speech recognition result may also include: the transcription result of each character recognized in the non-silent audio segment.

[0040] In step S12, the confidence level of the language tag for each non-silent audio segment is obtained based on the confidence level of each character identified in each non-silent audio segment.

[0041] There are multiple ways to implement step S12, which are not limited here. For example, the confidence level of a language tag for a non-silent audio segment can be equal to the average confidence level of each character in the non-silent audio segment.

[0042] In step S13, the probability value of the language tag of the video is obtained based on the confidence level of the language tag of each non-silent audio segment.

[0043] In step S14, the video is processed according to the probability value of the language tag of the video, wherein the processing includes distributing the video to the target user or discarding the video.

[0044] Discarding the video means not distributing the video to any user.

[0045] Using the above technical solution, the probability value of each video being in each language can be obtained. During video distribution, it is only necessary to set the probability value requirement for each business party / user for each language, and videos that meet the requirements can be sent to the corresponding business party / user (target user), or videos that do not meet any business party / user requirements can be discarded, facilitating video distribution based on language.

[0046] Optionally, step S13 includes:

[0047] For each language tag in the video, based on the confidence level of that language tag and the confidence levels of all language tags in the video, the probability value of that language tag in the video is obtained.

[0048] Based on the speech recognition results of the video, the video may have one language tag or at least two language tags. If the video has one language tag, then each language tag in the video refers to that language tag. If the video has at least two language tags, then each language tag in the video refers to either of those two language tags.

[0049] Optionally, the probability value of the language tag of the video is obtained based on all confidence scores of the language tag and the confidence scores of all language tags of the video, including:

[0050] Based on all the confidence scores of the language tag, select the confidence scores that are greater than or equal to the first threshold, sum the selected confidence scores to obtain the sum of confidence scores of the language tag, and divide the sum of confidence scores of the language tag by the sum of confidence scores of all language tags of the video to obtain the probability value of the language tag of the video.

[0051] The first threshold can be flexibly set according to needs and is not restricted here. To illustrate the method described in step S13, the following example is given. For instance, a video includes 10 non-silent audio segments, identifying three language labels: Cantonese, Mandarin, and Sichuanese. Six of the non-silent audio segments are labeled as Cantonese, with confidence levels of 0.2, 0.3, 0.6, 0.7, 0.7, and 0.8 respectively; three are labeled as Mandarin, with confidence levels of 0.5, 0.7, and 0.9 respectively; and one is labeled as Sichuanese, with a confidence level of 0.2. The first threshold is 0.6. For the Cantonese language tag, obtain the confidence scores of all non-silent audio segments tagged as Cantonese: 0.2, 0.3, 0.6, 0.7, 0.7, 0.8; select the confidence scores greater than or equal to 0.6: 0.6, 0.7, 0.7, 0.8; sum the selected confidence scores: 0.6 + 0.7 + 0.7 + 0.8 = 2.8; obtain the sum of confidence scores for this language tag: 2.8; divide this sum of confidence scores by the sum of confidence scores for all language tags in the video: The probability of obtaining the Cantonese language tag for this video is 0.5. Similarly, the probability of obtaining the Mandarin language tag for this video is 0.29, and the probability of obtaining the Sichuan dialect language tag for this video is 0.

[0052] The above analysis shows that setting a first threshold can prevent segment-level misclassification. For example, for a Cantonese language tag, non-silent audio segments with confidence levels of 0.2 and 0.3 are misclassified, as they are likely not Cantonese but rather English. If the sum of the confidence levels of all non-silent audio segments tagged as Cantonese is directly divided by the sum of the confidence levels of all non-silent audio segments in the video, the resulting probability value for the Cantonese language tag will be inaccurate, leading to misclassification.

[0053] Optionally, the video includes at least two language tags, and step S14 includes:

[0054] The probability sum is obtained by summing the probability values ​​of all language tags in the video.

[0055] If the sum of the probabilities is less than a second threshold, the video is discarded.

[0056] The second threshold can be set flexibly as needed and is not restricted here. To illustrate the method described in step S14, the following example is given. For instance, a video identifies three language tags: Cantonese, Mandarin, and Sichuanese. The probability of the video being tagged as Cantonese is 0.03, the probability of it being tagged as Mandarin is 0.05, and the probability of it being tagged as Sichuanese is 0.02, resulting in a sum of probabilities of 0.1. Assuming the second threshold is 0.3, then 0.1 is less than 0.3, and the video is discarded.

[0057] Analysis of the probability values ​​of language tags reveals that they reflect not only the confidence level of identifying each non-silent audio segment of the video as belonging to that language tag, but also the number of non-silent audio segments in the video containing that language tag. Therefore, a low probability value for a language tag indicates low confidence in the video belonging to that language tag, and / or a small number of non-silent audio segments in the video belonging to that language tag. Consequently, when the sum of the probabilities of all language tags in a video is less than the second threshold, it indicates that the video does not belong to the target language (the probability of identifying it as belonging to the target language using a model trained on multiple target languages ​​is low), and is therefore discarded. For example, continuing the example above, suppose we highly value Cantonese, Mandarin, and Sichuanese (the target languages). If the probability of identifying the video as a Cantonese language is 0.03, as a Mandarin language is 0.05, and as a Sichuanese language is 0.02, then the video is neither Cantonese (very little Cantonese content), nor Mandarin (very little Mandarin content), nor Sichuanese (very little Sichuanese content), and is not the language we value (not the target language). Therefore, we discard the video.

[0058] Optionally, the video includes at least two language tags, and step S14 further includes:

[0059] If the sum of the probabilities is greater than or equal to the second threshold, the language tag corresponding to the highest probability value among all language tags of the video is taken as the distribution language of the video.

[0060] That is, when a video should not be discarded, the language tag corresponding to the highest probability value is selected as the distribution language of the video. For example, if a video identifies three language tags: Cantonese, Mandarin, and Sichuanese, and the probability of the video being a Cantonese language tag is 0.5, the probability of the video being a Mandarin language tag is 0.3, and the probability of the video being a Sichuanese language tag is 0.1, then the sum of the probabilities is 0.9. Assuming the second threshold is 0.3, then 0.9 is greater than 0.3, so the Cantonese language tag is selected as the distribution language of the video.

[0061] Based on the association between the distribution language and the user, the video is distributed to the target user, wherein the target user is the user associated with the distribution language.

[0062] The association between the distribution language and the user can be flexibly set according to the requirements of the business party / user. Continuing the example above, this step can distribute the video to users / business parties associated with Cantonese.

[0063] Using the above technical solution, videos can be distributed based on the language tag corresponding to the highest probability value, provided that the video should not be discarded. This approach is applicable to business scenarios where language exclusivity is relatively high, such as when the target audience for the video is primarily users who speak English, French, or Mandarin. In this case, a single video will only be distributed to a single user group.

[0064] Similarly, optionally, the video includes only one language tag, and step S14 includes:

[0065] If the probability value of the language tag in the video is less than the third threshold, the video is discarded.

[0066] Optionally, the video includes only one language tag, and step S14 further includes:

[0067] If the probability value of the language tag of the video is greater than or equal to the third threshold, the video is distributed to the target user, which is the user associated with the language tag.

[0068] Alternatively, in another embodiment, step S14 includes:

[0069] If the probability value of the target language tag of the video is greater than or equal to the target threshold, the video is recalled into the machine learning resource pool of the target language represented by the target language tag.

[0070] The target threshold is the threshold corresponding to the target language tag. That is, the target threshold may be different for each target language tag. For example, if a video identifies three language tags: Cantonese, Mandarin, and Sichuanese, and the probability of the video being a Cantonese language tag is 0.5, the probability of the video being a Mandarin language tag is 0.4, and the probability of the video being a Sichuanese language tag is 0.01, then the target threshold for the Cantonese language tag is set to 0.4, the target threshold for the Mandarin language tag is set to 0.3, and the target threshold for the Sichuanese language tag is set to 0.3; then the video will be recalled to the Cantonese machine learning resource pool and the Mandarin machine learning resource pool.

[0071] By employing the above technical solution, videos whose probability value of the target language label is greater than or equal to the target threshold are recalled to the machine learning resource pool of the target language represented by the target language label, thereby increasing the video size of the machine learning resource pool. Furthermore, when using videos recalled to the machine learning resource pool of the target language as training data for annotation, the transcription results and language labels identified in step S11 can be directly used to revise the annotations, replacing a complete re-annotation and improving annotation efficiency.

[0072] Optionally, step S11 includes:

[0073] Extract the audio from the video.

[0074] Endpoint detection is performed on the audio to extract all non-silent audio segments.

[0075] Endpoint detection, also known as Voice Activity Detection (VAD), aims to distinguish between speech and non-speech regions. Simply put, endpoint detection accurately locates the start and end points of speech within noisy speech, removing silent and noisy portions.

[0076] Input all non-silent audio segments into the speech recognition model to obtain the speech recognition results of all non-silent audio segments. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment, each character recognized in the non-silent audio segment, and the confidence score of each recognized character.

[0077] The speech recognition model is trained on at least two languages. Furthermore, as... Figure 2 As shown, the speech recognition model includes an encoder, a decoder, and an attention mechanism. The encoder processes each non-silent audio segment (X1, ..., X...) T Encoded as a hidden layer vector h enc The decoder takes as input the character y predicted in the previous time step. u-1 and context vector c u To predict the character at the current moment, the context vector c u Through the encoder's hidden layer vector h enc and the hidden state of the decoder at the previous time step Calculated.

[0078] The above technical solution employs an end-to-end modeling approach to construct a speech recognition model. This speech recognition model does not require mandatory alignment of speech and text. Furthermore, this speech recognition model is trained on at least two languages, resulting in higher recognition accuracy compared to conventional language recognition-speech recognition cascade systems (which require high accuracy from the preceding language recognition model; if the language recognition is incorrect, the downstream speech recognition model will be unable to effectively transcribe the content).

[0079] Optionally, the speech recognition model can also be trained based on a language.

[0080] Optionally, in order to ensure that the speech recognition model outputs corresponding language labels during the prediction process, we add language labels to the beginning of the text corresponding to the training data, as shown below:

[0081] <yue>The three of you just didn't want me to log off, hahaha, I think I'm so narcissistic.

[0082] <yue>If you turn on the fan while my pants are tucked in, you'll know how hot I am.

[0083] <zh>Thank you, thank you, roses.

[0084] in, <yue>Indicates Cantonese. <zh>It represents Mandarin Chinese.

[0085] Therefore, this speech recognition model can determine the language (language) corresponding to the speech by the first predicted character.

[0086] Based on the above technical concept, this disclosure also provides a video distribution device. Figure 3 A block diagram of a video distribution apparatus provided according to one embodiment of this disclosure. (See diagram below.) Figure 3 As shown, the video distribution apparatus provided in this disclosure includes:

[0087] The speech recognition module is used to obtain the speech recognition results of all non-silent audio segments of the video. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment.

[0088] The segment confidence acquisition module is used to obtain the confidence of the language tag of each non-silent audio segment based on the confidence of each character identified in each non-silent audio segment.

[0089] The language probability acquisition module is used to obtain the probability value of the language tag of the video based on the confidence level of the language tag of each non-silent audio segment;

[0090] The processing module is configured to process the video based on the probability value of the video's language tag, wherein the processing includes distributing the video to a target user or discarding the video.

[0091] Using the above technical solution, the probability value of each video being in each language can be obtained. During video distribution, it is only necessary to set the probability value requirement for each business party / user for each language, and videos that meet the requirements can be sent to the corresponding business party / user (target user), or videos that do not meet any business party / user requirements can be discarded, facilitating video distribution based on language.

[0092] Optionally, the language probability acquisition module is specifically used to obtain the probability value of the language tag of the video based on the confidence level of each language tag in the video and the confidence levels of all language tags in the video.

[0093] Optionally, the probability value of the language tag of the video is obtained based on all confidence scores of the language tag and the confidence scores of all language tags of the video, including:

[0094] Based on all the confidence scores of the language tag, select the confidence scores that are greater than or equal to the first threshold, sum the selected confidence scores to obtain the sum of confidence scores of the language tag, and divide the sum of confidence scores of the language tag by the sum of confidence scores of all language tags of the video to obtain the probability value of the language tag of the video.

[0095] Setting a first threshold can prevent segment-level misclassification. For example, for a Cantonese language tag, non-silent audio segments with confidence levels of 0.2 and 0.3 are considered misclassifications, as they are likely not Cantonese but rather English. If the sum of the confidence levels of all non-silent audio segments tagged as Cantonese is directly divided by the sum of the confidence levels of all non-silent audio segments in the video, the resulting probability value for the Cantonese language tag will be inaccurate, leading to misclassification.

[0096] Optionally, the processing module includes:

[0097] The probability and calculation submodule is used to sum the probability values ​​of all language tags in the video to obtain the probability sum;

[0098] A discard submodule is used to discard the video if the probability sum is less than a second threshold.

[0099] Analysis of the probability values ​​of language tags reveals that they reflect not only the confidence level of identifying each non-silent audio segment of the video as belonging to that language tag, but also the number of non-silent audio segments in the video containing that language tag. Therefore, a low probability value for a language tag indicates low confidence in the video belonging to that language tag, and / or a small number of non-silent audio segments in the video belonging to that language tag. Consequently, when the sum of the probabilities of all language tags in a video is less than the second threshold, it indicates that the video does not belong to the target language (the probability of identifying it as belonging to the target language using a model trained on multiple target languages ​​is low), and is therefore discarded.

[0100] Optionally, the processing module further includes:

[0101] The language distribution acquisition submodule is used to, when the sum of the probabilities is greater than or equal to a second threshold, select the language tag corresponding to the highest probability value among all language tag probability values ​​of the video as the distribution language of the video.

[0102] The distribution submodule is used to distribute the video to a target user based on the association between the distribution language and the user, wherein the target user is the user associated with the distribution language.

[0103] Using the above technical solution, videos can be distributed based on the language tag corresponding to the highest probability value, provided that the video should not be discarded. This approach is applicable to business scenarios where language exclusivity is relatively high, such as when the target audience for the video is primarily users who speak English, French, or Mandarin. In this case, a single video will only be distributed to a single user group.

[0104] Optionally, the processing module is specifically configured to, when the probability value of the target language tag of the video is greater than or equal to the target threshold, recall the video into the machine learning resource pool of the target language represented by the target language tag.

[0105] By employing the above technical solution, videos whose probability value of the target language label is greater than or equal to the target threshold are recalled to the machine learning resource pool of the target language represented by the target language label, thereby increasing the video size of the machine learning resource pool. Furthermore, when using videos recalled to the machine learning resource pool of the target language as training data for annotation, the transcription results and language labels identified in step S11 can be directly used to revise the annotations, replacing a complete re-annotation and improving annotation efficiency.

[0106] Optionally, the speech recognition module includes:

[0107] The extraction submodule is used to extract the audio from the video.

[0108] The endpoint detection submodule is used to perform endpoint detection on the audio and extract all non-silent audio segments of the audio.

[0109] The recognition submodule is used to input all non-silent audio segments into the speech recognition model to obtain the speech recognition results of all non-silent audio segments. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment, each character recognized in the non-silent audio segment, and the confidence score of each recognized character.

[0110] The speech recognition model is trained on at least two languages. Furthermore, as... Figure 2 As shown, the speech recognition model includes an encoder, a decoder, and an attention mechanism. The encoder encodes each non-silent audio segment into a hidden layer vector. The decoder takes the character predicted in the previous time step and the context vector as input to predict the character in the current time step. The context vector is calculated using the hidden layer vector of the encoder and the hidden layer state of the decoder in the previous time step.

[0111] The above technical solution employs an end-to-end modeling approach to construct a speech recognition model. This speech recognition model does not require mandatory alignment of speech and text. Furthermore, this speech recognition model is trained on at least two languages, resulting in higher recognition accuracy compared to conventional language recognition-speech recognition cascade systems (which require high accuracy from the preceding language recognition model; if the language recognition is incorrect, the downstream speech recognition model will be unable to effectively transcribe the content).

[0112] Optionally, in order to ensure that the speech recognition model outputs corresponding language labels during the prediction process, we add language labels to the beginning of the text corresponding to the training data, as shown below:

[0113] <yue>The three of you just didn't want me to log off, hahaha, I think I'm so narcissistic.

[0114] <yue>If you turn on the fan while my pants are tucked in, you'll know how hot I am.

[0115] <zh>Thank you, thank you, roses.

[0116] in, <yue>Indicates Cantonese. <zh>It represents Mandarin Chinese.

[0117] Therefore, this speech recognition model can determine the language (language) corresponding to the speech by the first predicted character.

[0118] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0119] The following is for reference. Figure 4 This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0120] like Figure 4 As shown, electronic device 600 may include processing unit 601 (e.g., central processing unit, graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0121] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0122] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0123] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0124] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0125] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0126] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire speech recognition results for all non-silent audio segments of a video, wherein the speech recognition result for each non-silent audio segment includes: a language tag for the non-silent audio segment and a confidence score for each character recognized in the non-silent audio segment; obtain a confidence score for the language tag of the non-silent audio segment based on the confidence score for each character recognized in the non-silent audio segment; obtain a probability value for the language tag of the video based on the confidence score for the language tag of the non-silent audio segment; and process the video based on the probability value for the language tag of the video, wherein the processing includes distributing the video to a target user or discarding the video.

[0127] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0129] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, a language probability acquisition module can also be described as "a module that acquires the probability values ​​of a video as each language".

[0130] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0131] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0132] According to one or more embodiments of this disclosure, Example 1 provides a video distribution method, the method comprising:

[0133] Obtain the speech recognition results of all non-silent audio segments in the video. The speech recognition results of each non-silent audio segment include: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment.

[0134] The confidence level of the language label for each non-silent audio segment is obtained based on the confidence level of each character identified in each non-silent audio segment.

[0135] Based on the confidence level of the language tag for each non-silent audio segment, the probability value of the language tag for the video is obtained;

[0136] The video is processed based on the probability value of its language tag, wherein the processing includes distributing the video to a target user or discarding the video.

[0137] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, which obtains the probability value of the language tag of the video based on the confidence level of the language tag of each non-silent audio segment, including:

[0138] For each language tag in the video, based on the confidence level of that language tag and the confidence levels of all language tags in the video, the probability value of that language tag in the video is obtained.

[0139] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, which obtains the probability value of the language tag of the video based on all confidence scores of the language tag and the confidence scores of all language tags of the video, including:

[0140] Based on all the confidence scores of the language tag, select the confidence scores that are greater than or equal to the first threshold, sum the selected confidence scores to obtain the sum of confidence scores of the language tag, and divide the sum of confidence scores of the language tag by the sum of confidence scores of all language tags of the video to obtain the probability value of the language tag of the video.

[0141] According to one or more embodiments of this disclosure, Example 4 provides a method as described in any one of Examples 1-3, wherein the video has at least two language tags, and processing the video based on the probability values ​​of the language tags of the video includes:

[0142] The probability sum is obtained by summing the probability values ​​of all language tags in the video.

[0143] If the sum of the probabilities is less than a second threshold, the video is discarded.

[0144] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, which further includes processing the video based on the probability value of the video's language tag:

[0145] If the sum of the probabilities is greater than or equal to the second threshold, the language tag corresponding to the highest probability value among all language tags of the video is taken as the distribution language of the video.

[0146] Based on the association between the distribution language and the user, the video is distributed to the target user, wherein the target user is the user associated with the distribution language.

[0147] According to one or more embodiments of this disclosure, Example 6 provides a method as described in any one of Examples 1-3, which processes the video based on the probability value of the video's language tag, including:

[0148] If the probability value of the target language tag of the video is greater than or equal to the target threshold, the video is recalled into the machine learning resource pool of the target language represented by the target language tag, wherein the target threshold is the threshold corresponding to the target language tag.

[0149] According to one or more embodiments of this disclosure, Example 7 provides a method as described in any one of Examples 1-3, which obtains the speech recognition results of all non-silent audio segments of a video, including:

[0150] Extract the audio from the video;

[0151] Endpoint detection is performed on the audio to extract all non-silent audio segments.

[0152] Input all non-silent audio segments into the speech recognition model to obtain the speech recognition results of all non-silent audio segments. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment, each character recognized by the non-silent audio segment, and the confidence score of each recognized character.

[0153] The speech recognition model is trained on at least two languages ​​and includes an encoder, a decoder, and an attention mechanism. The encoder encodes each non-silent audio segment into a hidden layer vector. The decoder takes the character predicted in the previous time step and the context vector as input to predict the character in the current time step. The context vector is calculated using the hidden layer vector of the encoder and the hidden layer state of the decoder in the previous time step.

[0154] According to one or more embodiments of this disclosure, Example 8 provides a video distribution apparatus, including:

[0155] The speech recognition module is used to obtain the speech recognition results of all non-silent audio segments of the video. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment.

[0156] The segment confidence acquisition module is used to obtain the confidence of the language tag of each non-silent audio segment based on the confidence of each character identified in each non-silent audio segment.

[0157] The language probability acquisition module is used to obtain the probability value of the language tag of the video based on the confidence level of the language tag of each non-silent audio segment;

[0158] The processing module is configured to process the video based on the probability value of the video's language tag, wherein the processing includes distributing the video to a target user or discarding the video.

[0159] According to one or more embodiments of this disclosure, Example 9 provides the apparatus of Example 8, wherein the language probability acquisition module is specifically configured to, for each language tag of the video, obtain a probability value of the language tag of the video based on all confidence levels of the language tag and the confidence levels of all language tags of the video.

[0160] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 9, which obtains the probability value of the language tag of the video based on all confidence levels of the language tag and the confidence levels of all language tags of the video, including:

[0161] Based on all the confidence scores of the language tag, select the confidence scores that are greater than or equal to the first threshold, sum the selected confidence scores to obtain the sum of confidence scores of the language tag, and divide the sum of confidence scores of the language tag by the sum of confidence scores of all language tags of the video to obtain the probability value of the language tag of the video.

[0162] According to one or more embodiments of this disclosure, Example 11 provides an apparatus of any one of Examples 8-10, wherein the processing module includes:

[0163] The probability and calculation submodule is used to sum the probability values ​​of all language tags in the video to obtain the probability sum;

[0164] A discard submodule is used to discard the video if the probability sum is less than a second threshold.

[0165] According to one or more embodiments of this disclosure, Example 12 provides an apparatus of Example 11, wherein the processing module further includes:

[0166] The language distribution acquisition submodule is used to, when the sum of the probabilities is greater than or equal to a second threshold, select the language tag corresponding to the highest probability value among all language tag probability values ​​of the video as the distribution language of the video.

[0167] The distribution submodule is used to distribute the video to a target user based on the association between the distribution language and the user, wherein the target user is the user associated with the distribution language.

[0168] According to one or more embodiments of this disclosure, Example 13 provides an apparatus of any one of Examples 8-10, wherein the processing module is specifically configured to recall the video into a machine learning resource pool of the target language represented by the target language tag if the probability value of the target language tag of the video is greater than or equal to a target threshold.

[0169] According to one or more embodiments of this disclosure, Example 14 provides an apparatus of any one of Examples 8-10, wherein the speech recognition module includes:

[0170] An extraction submodule is used to extract the audio from the video;

[0171] The endpoint detection submodule is used to perform endpoint detection on the audio and extract all non-silent audio segments of the audio.

[0172] The recognition submodule is used to input all non-silent audio segments into the speech recognition model to obtain the speech recognition results of all non-silent audio segments. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment, each character recognized by the non-silent audio segment, and the confidence score of each recognized character.

[0173] The speech recognition model is trained on at least two languages ​​and includes an encoder, a decoder, and an attention mechanism. The encoder encodes each non-silent audio segment into a hidden layer vector. The decoder takes the character predicted in the previous time step and the context vector as input to predict the character in the current time step. The context vector is calculated using the hidden layer vector of the encoder and the hidden layer state of the decoder in the previous time step.

[0174] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-7.

[0175] According to one or more embodiments of this disclosure, Example 16 provides an electronic device comprising:

[0176] A storage device having at least one computer program stored thereon;

[0177] At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method described in any one of Examples 1-7.

[0178] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0179] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0180] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.< / zh> < / yue> < / zh> < / yue> < / yue> < / zh> < / yue> < / zh> < / yue> < / yue>

Claims

1. A video distribution method, characterized in that, The method includes: Obtain the speech recognition results of all non-silent audio segments in the video. The speech recognition results of each non-silent audio segment include: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment. The confidence level of the language label for each non-silent audio segment is obtained based on the confidence level of each character identified in each non-silent audio segment. Based on the confidence level of the language tag for each non-silent audio segment, the probability value of the language tag for the video is obtained; The video is processed based on the probability value of its language tag, wherein the processing includes distributing the video to a target user or discarding the video; The step of obtaining the probability value of the language tag of the video based on the confidence level of the language tag of each non-silent audio segment includes: For each language tag in the video, based on the confidence level of that language tag and the confidence levels of all language tags in the video, the probability value of that language tag in the video is obtained.

2. The video distribution method according to claim 1, characterized in that, Based on the confidence scores of all language tags for that language and the confidence scores of all language tags for the video, the probability value of the language tag for that video is obtained as follows: Based on all the confidence scores of the language tag, select the confidence scores that are greater than or equal to the first threshold, sum the selected confidence scores to obtain the sum of confidence scores of the language tag, and divide the sum of confidence scores of the language tag by the sum of confidence scores of all language tags of the video to obtain the probability value of the language tag of the video.

3. The video distribution method according to claim 1 or 2, characterized in that, The video includes at least two language tags, and the video is processed based on the probability values ​​of the language tags, including: The probability sum is obtained by summing the probability values ​​of all language tags in the video. If the sum of the probabilities is less than a second threshold, the video is discarded.

4. The video distribution method according to claim 3, characterized in that, Processing the video based on the probability value of its language tags further includes: If the sum of the probabilities is greater than or equal to the second threshold, the language tag corresponding to the highest probability value among all language tags of the video is taken as the distribution language of the video. Based on the association between the distribution language and the user, the video is distributed to the target user, wherein the target user is the user associated with the distribution language.

5. The video distribution method according to claim 1 or 2, characterized in that, The processing of the video based on the probability value of its language tags includes: If the probability value of the target language tag of the video is greater than or equal to the target threshold, the video is recalled into the machine learning resource pool of the target language represented by the target language tag, wherein the target threshold is the threshold corresponding to the target language tag.

6. The video distribution method according to claim 1 or 2, characterized in that, The speech recognition results for all non-silent audio segments of the video include: Extract the audio from the video; Endpoint detection is performed on the audio to extract all non-silent audio segments. Input all non-silent audio segments into the speech recognition model to obtain the speech recognition results of all non-silent audio segments. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment, each character recognized by the non-silent audio segment, and the confidence score of each recognized character. The speech recognition model is trained on at least two languages ​​and includes an encoder, a decoder, and an attention mechanism. The encoder encodes each non-silent audio segment into a hidden layer vector. The decoder takes the character predicted in the previous time step and the context vector as input to predict the character in the current time step. The context vector is calculated using the hidden layer vector of the encoder and the hidden layer state of the decoder in the previous time step.

7. A video distribution device, characterized in that, include: The speech recognition module is used to obtain the speech recognition results of all non-silent audio segments of the video. The speech recognition result of each non-silent audio segment includes: the language label of the non-silent audio segment and the confidence score of each character recognized in the non-silent audio segment. The segment confidence acquisition module is used to obtain the confidence of the language tag of each non-silent audio segment based on the confidence of each character identified in each non-silent audio segment. The language probability acquisition module is used to obtain the probability value of the language tag of the video based on the confidence level of the language tag of each non-silent audio segment; A processing module is configured to process the video based on the probability value of the video's language tags, wherein the processing includes distributing the video to a target user or discarding the video; The language probability acquisition module is used to obtain the probability value of each language tag in the video based on the confidence level of each language tag and the confidence levels of all language tags in the video.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-6.

9. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method, system and apparatus for providing multilingual program

    CN101437149A

  • Language recognition method and device, electronic equipment and storage medium

    CN112017630A