How to set up the server and subtitle location

By calculating the text matching degree between the subtitles and different image areas in the target image, and setting the subtitles to display in the area with the smallest matching degree, the problem of subtitles blocking content in the prior art is solved, and the user experience is improved.

CN115866312BActive Publication Date: 2025-05-23青岛聚看云科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111119843.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-24
Publication Date
2025-05-23
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Subtitles added by existing AI subtitles technology are usually fixed in an area below the screen, which may block the content that users need to watch and affect the user experience.

Method used

By receiving a subtitle request, voice recognition is performed in response to the speech stream, the text matching degree between the subtitles and different image areas in the target image is calculated, and the coordinate area of ​​the subtitles is set in an image area with a matching degree less than the maximum value.

Benefits of technology

Reduce the occlusion of the target image by subtitles and improve the user's experience of watching subtitles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115866312B_ABST
    Figure CN115866312B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a server and a method for setting the position of subtitles, wherein the server is configured to: receive a subtitle request; in response to the subtitle request, when receiving a voice stream, perform voice recognition on the voice stream to obtain subtitles; calculate the matching degree between the subtitles and the text in each image area, wherein the image area is a local display area of ​​a target image corresponding to the voice stream, and the target image includes multiple image areas; and set the coordinate area of ​​the subtitles within the image area where the matching degree is less than the maximum value. The present application reduces the occlusion effect of subtitles on image content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of subtitles, and in particular to a server and a method for setting a subtitle position. Background Art

[0002] Subtitles are an important way for people to understand the content of videos when watching them. For pre-recorded videos such as movies and TV series, subtitles are usually added by staff after the video is recorded. However, for real-time recorded videos such as conference videos, online education videos, training videos, etc., it is difficult and costly for staff to add subtitles in real time. AI (Artificial Intelligence) subtitle technology is usually used to automatically add subtitles.

[0003] In related technologies, subtitles added by AI subtitle technology are usually fixed in an area at the bottom of the screen. However, this area may contain content that users need to watch, and the subtitles will block the content that users need to watch, affecting the user experience. Summary of the invention

[0004] In order to solve the technical problem of low subtitle accuracy, the present application provides a method for setting a server and subtitle location.

[0005] In a first aspect, the present application provides a server, which is configured to:

[0006] Receive subtitle requests;

[0007] In response to the subtitle request, when receiving a voice stream, performing voice recognition on the voice stream to obtain subtitles;

[0008] Calculating the matching degree between the subtitle and the text in each image area, wherein the image area is a local display area of ​​a target image corresponding to the voice stream, and the target image includes multiple image areas;

[0009] The coordinate area of ​​the subtitle is set within the image area where the matching degree is less than the maximum value.

[0010] In some embodiments, the target text is text recognized by performing an optical character recognition method on the target image.

[0011] In some embodiments, calculating the matching degree between the subtitle and the text in each image region includes:

[0012] Calculating the matching degree of each subtitle segmentation word with the corresponding target segmentation word in each image area, wherein the subtitle segmentation word is the segmentation word obtained by segmenting the subtitle word, and the target segmentation word is the segmentation word obtained by segmenting the text in the image area;

[0013] All matching degrees in each image region are weighted to obtain the matching degree between the subtitle and the text in each image region.

[0014] In some embodiments, the multiple image areas of the target image are divided according to the distribution of text in the target image in the target image.

[0015] In a second aspect, the present application provides a method for setting a subtitle position, the method comprising:

[0016] Receive subtitle requests;

[0017] In response to the subtitle request, when receiving a voice stream, performing voice recognition on the voice stream to obtain subtitles;

[0018] Calculating the matching degree between the subtitle and the text in each image area, wherein the image area is a local display area of ​​a target image corresponding to the voice stream, and the target image includes multiple image areas;

[0019] The coordinate area of ​​the subtitle is set within the image area where the matching degree is less than the maximum value.

[0020] The beneficial effects of the server and subtitle location setting method provided by this application include:

[0021] The present application calculates the matching degree between subtitles and texts in different image areas of the target image, and sets the subtitles in the image area with the smallest matching degree, thereby reducing the impact of subtitles on the understanding of the voice stream caused by occluding the target image, and improving the user's experience of viewing subtitles. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 Schematic diagram of an operation scenario between a display device and a control device according to some embodiments is exemplarily shown in FIG.

[0024] Figure 2 A schematic diagram of a process of generating subtitles according to some embodiments is exemplarily shown in FIG.

[0025] Figure 3 exemplarily shows a schematic diagram of an interface of a target image according to some embodiments;

[0026] Figure 4 A schematic diagram of a subtitle display interface according to some embodiments is exemplarily shown in FIG.

[0027] Figure 5 A schematic diagram of a process for setting a subtitle position according to some embodiments is exemplarily shown in FIG.

[0028] Figure 6 exemplarily shows a schematic diagram of an interface of a target image according to some embodiments;

[0029] Figure 7 A schematic diagram of a subtitle display interface according to some embodiments is exemplarily shown in FIG.

[0030] Figure 8 A schematic diagram of a subtitle display interface according to some embodiments is exemplarily shown in FIG.

[0031] Fig. 9 exemplarily shows a timing diagram of starting a shared desktop according to some embodiments;

[0032] Fig.10 exemplarily shows a timing diagram of subtitle generation according to some embodiments;

[0033] Fig.11 A timing diagram of subtitle generation according to some embodiments is exemplarily shown in FIG. DETAILED DESCRIPTION

[0034] In order to make the purpose and implementation method of the present application clearer, the exemplary implementation method of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0035] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0036] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.

[0037] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0038] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0039] Figure 1 FIG. 1 is a schematic diagram of an operation scenario between a display device and a control device according to an embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control apparatus 100 .

[0040] In some embodiments, the control device 100 may be a remote controller, and the communication between the remote controller and the display device includes infrared protocol communication or Bluetooth protocol communication, and other short-range communication methods, and the display device 200 is controlled wirelessly or wired. The user may control the display device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, etc.

[0041] In some embodiments, a smart device 300 (such as a mobile terminal, a tablet computer, a computer, a laptop computer, etc.) may also be used to control the display device 200. For example, the display device 200 is controlled using an application running on the smart device.

[0042] In some embodiments, the display device 200 can also be controlled in a manner other than the control device 100 and the smart device 300. For example, the user's voice command control can be directly received through a module for obtaining voice commands configured inside the display device 200, or the user's voice command control can be received through a voice control device set outside the display device 200.

[0043] In some embodiments, the display device 200 also communicates data with the server 400. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactions to the display device 200. The server 400 may be one cluster or multiple clusters, and may include one or more types of servers.

[0044] In some embodiments, the display device can run multiple applications, one of which can be a conference application, and the interface of the conference application can be provided with a desktop sharing control and an audio input control. The desktop sharing control can be configured to share the display interface of the current device with other display devices participating in the current conference in response to a trigger, so that the other display devices participating in the current conference display the display interface; the audio input control can be a microphone control, which can be configured to share the audio received by the current device with other display devices participating in the current conference in response to a trigger, so that the other display devices participating in the current conference play the audio.

[0045] For example, the participants of a conference include participant 1 and participant 2. The terminal device used by participant 1 to participate in the conference is display device 1, and the terminal device used by participant 2 to participate in the conference is display device 2. When participant 1 needs to share the content displayed by display device 1 with participant 2, he can click the desktop sharing control, and the server can control display device 2 to display the display interface of display device 1; when participant 1 needs to explain the content displayed by display device 1, he can click the microphone control and then explain the content in the display interface, and the server can control display device 2 to play the audio of participant 1's explanation.

[0046] For ease of understanding, in the embodiment of the present application, participant 1 is a speaker in a meeting, and participant 2 is an audience. Of course, during the actual meeting, the identities of the two can be switched.

[0047] In some embodiments, to facilitate understanding of the speaker's speech content, the conference application provides an AI subtitle function. When the audience turns on this function, the server can perform voice recognition on the speaker's audio recorded by the speaker's display device and generate subtitles based on the recognition results. However, the accuracy of voice recognition is limited, resulting in limited accuracy of the subtitles.

[0048] In some embodiments, the subtitles generated by the AI ​​subtitle function are usually in a fixed display area, which may cause the generated subtitles to block the content that the viewer needs to watch.

[0049] In actual implementation, there is more than one speaker and audience in a meeting. This application introduces the process of subtitle generation by taking a meeting including one speaker and one audience as an example. The scenario of multiple speakers or multiple audiences can refer to the embodiments of this application for adaptive adjustments.

[0050] In order to solve the problem of low accuracy of subtitles, the present application embodiment provides a subtitle generation method, see Figure 2 , the method may include the following steps:

[0051] Step S110: receiving a subtitle request.

[0052] In some embodiments, the participants of a conference include two users, namely, participant 1 and participant 2. Participant 1 is a speaker, participant 2 is an audience, the terminal device used by participant 1 to participate in the conference is display device 1, and the terminal device used by participant 2 to participate in the conference is display device 2.

[0053] In some embodiments, after participant 1 clicks on the shared desktop control on display device 1, display device 1 may, in response to the shared desktop control being triggered, package the shared desktop command, the current screen image of display device 1, and the device ID of display device 1 and send them to the server, wherein the shared desktop command is a preset command corresponding to the shared desktop control, which is used to enable the server to control the display devices of other participants to display the screen image of participant 1. After participant 1 clicks on the audio input control on display device 1, display device 1 may, in response to the audio input control being triggered, start the microphone to record the audio of participant 1's speech in real time, and package the audio, conference ID, and the device ID of display device 2 and send them to the server, and the audio may also be called a voice stream.

[0054] During the meeting, participant 1 may adjust the current display interface of the display device, for example, adjust the current display interface from the first image to the second image on the next page of the first image. The display device may be configured to send the changed display interface and the page turning message to the server when determining that the current screen change is a preset screen change, such as page turning. The server may determine that the screen of display device 1 has changed based on the page turning message received from display device 1. Alternatively, the server may also determine that the screen of display device 1 has changed based on the new screen image received from display device 1.

[0055] In some embodiments, after participant 2 triggers the subtitle control on display device 2, display device 2 may generate a subtitle request, which may include the conference ID of the current conference and the device ID of display device 2. The conference ID may be the conference number. After generating the subtitle request, display device 2 sends the subtitle request to the server.

[0056] In some embodiments, participant 2 may trigger the subtitle control at any time after participant 2 joins the meeting.

[0057] In some embodiments, the conference application may be configured to automatically enable the subtitle function after a participant joins the conference, and if the participant has enabled the shared desktop function, the subtitle function for the participant is exited.

[0058] Step S120: In response to the subtitle request, when a voice stream is received, segmentation processing is performed on the semantic text corresponding to the voice stream to obtain a plurality of segmented words to be corrected.

[0059] In some embodiments, after receiving the subtitle request, the server can obtain the shared desktop command corresponding to the conference ID based on the conference ID in the subtitle request, and determine display device 1 as the target display device because the device ID corresponding to the shared desktop command is the device ID of display device 1. The screen image of the target display device is determined as the target image for generating subtitles. Subtitles need to be generated on the screen image sent by display device 1 so that display device 2 of participant 2 can display the subtitles on the screen image.

[0060] In some embodiments, the target image may also refer to a reference image for which subtitles are to be generated.

[0061] For example, the target image may be all page images of a document uploaded to the server by display device 1 or display device 2, or a portion of the page images, such as the current page image, or the current page image and a preset number of page images before and after. The current page image is an image displayed by display device 1 and display device 2, which may be uploaded to the server by display device 1. The server may determine the image most recently uploaded by display device 1 as the current page image, may identify the page number from the current page image, and then obtain the page images of a preset number of pages before and after the document, which may be 2, i.e., the server may determine the current page image, the page images of the first two pages, and the page images of the last two pages as the target images of the received voice stream.

[0062] For example, the target image may also be a screen image that participant 1 has recently sent to the server a preset number of times, and the preset number of times may be 3. If a message indicating a preset screen change is received from display device 1, such as a page turning message, the server may update the target image. If the target image is an image, the target image is updated to the screen image of display device 1 corresponding to the page turning message.

[0063] In some embodiments, the server is configured to control only display devices with subtitle function enabled to display subtitles. Of course, the server can also be configured to display subtitles on all participating display devices by default.

[0064] In some embodiments, after acquiring the target image, the server may perform text recognition on the target image to obtain the text on the target image, and use the text on the target image as the target text.

[0065] In some embodiments, the text recognition method may be an optical character recognition method or other general text recognition methods.

[0066] In some embodiments, after the target text is obtained, the target text may be segmented to facilitate comparison with the text recognized by the speech stream.

[0067] In some embodiments, when the server receives the voice stream sent by the display device 1, it can determine that the voice stream corresponds to the current target image. The voice stream is subjected to voice recognition to obtain a semantic text. The semantic text is segmented to obtain a plurality of segmented words to be corrected. To facilitate the distinction between different segmented words, each segmented word to be corrected can be provided with a segmented word number, which is the order determined by the segmented word processing. For example, for the semantic text ABCCDD, the segmented results are: AB, CC, DD, and the segmented word numbers are: 1, 2, 3.

[0068] Step S130: For each participle to be revised, a group of candidate words is obtained respectively, wherein one of the candidate words includes the participle to be revised.

[0069] In some embodiments, for each participle to be corrected, the first candidate word may be determined to be the participle to be corrected, the weight is a preset weight, such as 10, and the Nth candidate word may be obtained from the pronunciation confusion set, where N is greater than or equal to 2. Of course, it is also possible that the confusion set does not contain the candidate word corresponding to the participle to be corrected, so the number of candidate words for each participle to be corrected is greater than or equal to 1.

[0070] In some embodiments, a pronunciation confusion set can be pre-set, which contains a large number of confusing phrases whose pronunciations are easily confused. Each confusing phrase can be set with a weight, which can represent the pronunciation similarity. The range of pronunciation similarity can be 0 to 1. The smaller the weight, the less likely it is to be confused, and the larger the weight, the easier it is to be confused.

[0071] For example, in the confusion set, the weight of AA-AB is 0.8, and the weight of AA-AC is 0.6, which means that the probability of AA being confused as AB is higher than the probability of AA being confused as AC. Of course, in the pronunciation confusion set, easily confused words can also be stored in other ways besides confusing phrases, such as tree diagrams.

[0072] Taking a participle to be corrected as an example, in the pronunciation confusion set, all confused phrases containing the participle to be corrected, or confused phrases containing the corrected participle and with a weight greater than a third threshold value can be obtained, wherein, illustratively, the third threshold value can be 0.6. In the obtained confused phrases, words other than the participle to be corrected are used as candidate words for the participle to be corrected. For example, for AA-AB, if the participle to be corrected is AB, AA is used as a candidate word. For each participle to be corrected, at least one candidate word can be obtained, and as a group of candidate words, a maximum of a preset number of candidate words can be obtained, and the preset number can be 5.

[0073] The above method for obtaining candidate words is only an example. In actual implementation, candidate words can also be obtained by other methods.

[0074] Step S140: for each to-be-corrected segmentation, respectively calculate the pronunciation similarity and glyph similarity of each candidate word with the target text; if there is a segmentation in the target text whose pronunciation similarity with one of the candidate words reaches a first threshold, and whose glyph similarity with the to-be-corrected segmentation does not reach a second threshold, the segmentation is determined as the target segmentation corresponding to the to-be-corrected segmentation; otherwise, if there is no segmentation in the target text whose pronunciation similarity with one of the candidate words reaches the first threshold, and whose glyph similarity with the to-be-corrected segmentation does not reach the second threshold, the to-be-corrected segmentation is not corrected, and the to-be-corrected segmentation is determined as the target segmentation, wherein the target text is the text obtained from the target image corresponding to the speech stream.

[0075] In some embodiments, the word segmentation to be revised may or may not need to be revised. Whether the word segmentation to be revised needs to be revised can be determined based on the two indicators of glyph similarity and pronunciation similarity. Among them, the calculation method of glyph similarity and pronunciation similarity can be obtained according to some existing calculation methods, and the embodiments of this application will not be repeated.

[0076] The scenario that needs to be corrected is as follows: for a segmented word to be corrected, if the pronunciation similarity between a segmented word in the target text and one of the candidate words reaches the first threshold, and the glyph similarity with the segmented word to be corrected does not reach the second threshold, indicating that the segmented word to be corrected is similar in pronunciation to a segmented word in the target text, but has a large glyph deviation, then the segmented word in the target text can be determined as the target segmented word. Exemplarily, the first threshold can range from 0.5 to 1, and the second threshold can range from 0.8 to 1.

[0077] Scenarios that do not require correction include scenarios other than the above scenarios. For example, the pronunciation similarity between a segmentation in the target text and one of the candidate words reaches a first threshold, and the glyph similarity with the segmentation to be corrected reaches a second threshold, indicating that the segmentation to be corrected is the same as a segmentation in the target text and does not need to be corrected. For another example, the pronunciation similarity between a segmentation in the target text and one of the candidate words reaches a first threshold, and the glyph similarity between the segmentation to be corrected and the segmentation to be corrected reaches a second threshold, indicating that the segmentation to be corrected is the same as a segmentation in the target text and does not need to be corrected. For another example, the pronunciation similarity between a segmentation in the target text and one of the candidate words does not reach the first threshold, indicating that the pronunciation of the segmentation to be corrected and the segmentation in the target text are quite different, and the accuracy of correction based on the target text is low, so correction cannot be made based on the target text.

[0078] In some embodiments, each to-be-corrected segmented word may be corrected according to one or more correction principles. Taking a to-be-corrected segmented word as an example, the correction principles may include a text recurrence principle and a pronunciation recurrence principle:

[0079] 1) Principle of text reproduction.

[0080] A text recurrence principle is: for a word segmentation to be corrected, if one of the candidate words appears in the target text, the weight of the candidate word is set to the largest among the word segmentation parameters of the group of candidate words; if multiple candidate words appear in the target text, the original weights of the multiple candidate words are compared, and the weight of the candidate word with the largest original weight is set to the largest among the group of candidate words, where the original weight is the weight of the candidate word corresponding to the word segmentation to be corrected in the pronunciation confusion set.

[0081] In a group of candidate words, a method of setting the weight of one of the candidate words to be the largest in the group of candidate words may be to set the weight of the candidate word to 100.

[0082] 2) Pronunciation reproduction principle.

[0083] One pronunciation repetition principle is to compare the pronunciation of each candidate word with the pronunciation of the target text. The similarity considerations may include pronunciation and tone, and these two considerations may be weighted. The same pronunciation means that both the pronunciation and tone are the same, in which case the similarity is the highest, and the similarity of other cases is lower than this case.

[0084] After comparing the pronunciations, if the pronunciation of one of the candidate words appears in the pronunciation of the target text, the text corresponding to this pronunciation in the target text is added as a new candidate word to the candidate words corresponding to the word segmentation parameters, and the weight of the new candidate word is set to the largest among the candidate words corresponding to the word segmentation parameters.

[0085] After comparing the pronunciations, if the pronunciations of multiple candidate words appear in the pronunciation of the target text, the original weights of the multiple candidate words are compared, and the weight of the candidate word with the largest original weight is set to be the largest in the group of candidate words.

[0086] In a group of candidate words, a method of setting the weight of one of the candidate words to be the largest in the group of candidate words may be to set the weight of the candidate word to 100.

[0087] In some embodiments, the priority of the text recurrence principle can be pre-set to be higher than the pronunciation recurrence principle, that is, after a successful correction based on the text recurrence principle, correction will no longer be performed based on the pronunciation recurrence principle, wherein a successful correction based on the text recurrence principle means that one or more candidate words appear in the target text, and if any candidate word does not appear in the target text, the correction fails and correction will continue to be performed based on the pronunciation recurrence principle.

[0088] In some embodiments, after both the text recurrence principle and the pronunciation recurrence principle have failed to correct, the original weight of each candidate word may not be changed, wherein the failure of the correction according to the pronunciation recurrence principle means that the pronunciation of each candidate word is less similar to the pronunciation of the target text than a preset threshold, indicating that the pronunciations are not similar, and the success of the correction according to the pronunciation recurrence principle means that the pronunciation of at least one candidate word is more similar to the pronunciation of the target text than or equal to the threshold.

[0089] In some embodiments, the correction principle is not limited to the text recurrence principle and the pronunciation recurrence principle, and the priority is not limited to the text recurrence principle being higher than the pronunciation recurrence principle, as long as the word segmentation is corrected according to the target text.

[0090] In some embodiments, after the correction is completed, the candidate word with the highest weight corresponding to each participle to be corrected may be determined as the target participle corresponding to the participle to be corrected.

[0091] According to the above correction rules, a method for determining a target participle may be: for each participle to be corrected, comparing the glyph similarity of each candidate word with the target text, if there is a candidate word with a glyph similarity greater than a second threshold, determining the candidate word with a glyph similarity greater than the second threshold as the target participle, otherwise, not determining the target participle for the participle to be corrected, wherein the target text is a text obtained from a target image corresponding to the speech stream; for the participle to be corrected for which the target participle has not been determined, comparing the pronunciation similarity of each candidate word with the target text, if there is a candidate word with a pronunciation similarity greater than a first threshold, determining the candidate word with a pronunciation similarity greater than the first threshold as the target participle, otherwise, determining the participle to be corrected for which the target participle has not been determined as the target participle.

[0092] Step S150: combining the target segmented words corresponding to each segmented word to be corrected into subtitles.

[0093] In some embodiments, after all the segmented words to be corrected are corrected, the target segmented words of all the segmented words to be corrected can be sequentially combined into a sentence according to the grouping number, that is, the subtitles to be displayed on the audience's display device, and the subtitles are returned to the display device of the audience corresponding to the conference ID.

[0094] According to the above subtitle generation method, an example of subtitle generation is:

[0095] For example, the speaker's speech content is: "In the current large screen optimization solution", the speech stream voice_strem of the speech content is voice recognized to obtain the semantic text candidate_text, for example, candidate_text = {in the line tight large bottle optimization solution}. The semantic text is segmented to obtain 6 segmented words to be corrected: In the line tight large bottle optimization solution, it can be set:

[0096] candidate_text[1]=[{"text":"line tight","weight":10}];

[0097] candidate_text[2]=[{"text": "of", "weight": 10}];

[0098] candidate_text[3]=[{"text":"Big Bottle","weight":10}];

[0099] candidate_text[4] = [{"text":"optimization","weight":10}];

[0100] candidate_text[5] = [{"text":"scheme","weight":10}];

[0101] candidate_text[6]=[{"text":"medium","weight":10}].

[0102] Among them, candidate_text[1]~candidate_text[6] represent 6 candidate words to be revised, text represents the text of the candidate word, weight represents the weight of the candidate word, and the weight of each word to be revised obtained according to the semantic text is 10.

[0103] For each word segment to be corrected, a set of candidate words and their weights are obtained from the pronunciation confusion set and added to candidate_text[1]~candidate_text[6], and the following result is obtained: candidate_text[1] = [

[0105] {"text":"Line Tight","weight":10},

[0106] {"text":"close first","weight":8},

[0107] {"text":"Advanced","weight":5},

[0108] {"text":"sink","weight":5} ];

[0110] …,

[0111] candidate_text[3]= [

[0113] {"text": "Large bottle", "weight": 10},

[0114] {"text": "Large screen", "weight": 9},

[0115] {"text": "Draw level", "weight": 8} ;

[0117] …

[0118] It can be seen that for candidate_text[1], if the recognition result of the speech recognition algorithm is directly adopted, the determined target word segmentation is "tight line", which does not match the speech content of the speaker. For candidate_text[3], if the recognition result of the speech recognition algorithm is directly adopted, the determined target word segmentation is "large bottle", which does not match the speech content of the speaker.

[0119] Through the screen image corresponding to the speech stream, i.e., the target image, the word segmentation to be corrected can be corrected. For a word segmentation to be corrected, first compare whether there is a word segmentation in the target text screen_text in the screen image that is the same as one of the candidate words of the word segmentation to be corrected. If so, update the weight of the same word segmentation.

[0120] For example, the target image is Figure 3 the image shown. The target text recognized by the target image is: "In the current large screen optimization solution, more and more attention is paid to the user experience", and the word segmentation result is: "Currently", "of", "large screen", "optimization", "solution", "in", "more and more", "pay attention to", "user", "experience". For candidate_text[3], a word segmentation in the text of the screen image corresponding to the speech stream is "large screen", so the weight of the candidate word "large screen" in candidate_text[3] can be set to 100. For a word segmentation parameter, if the text screen_text in the screen image does not contain a word that is the same as any of the candidate words of the word segmentation parameter, then compare the pronunciation of each word segmentation in screen_text with the pronunciation of the candidate words of the word segmentation parameter, calculate the similarity, and in the word segmentation parameter, update the weight of the word segmentation in the screen image with the highest similarity. For example, for candidate_text[1], a word segmentation in the text of the screen image corresponding to the speech stream is "currently", and the similarity of the pronunciation with the candidate words "tight line", "near first", "advanced", and "trapped" is relatively close. "Currently" can be added to candidate_text[1], and the weight of "currently" is set to 100.

[0121] After candidate_text[1] to candidate_text[6] are all corrected, the candidate word with the highest weight among candidate_text[1] to candidate_text[6] can be taken out as the target word of each word to be corrected. The target word of each word to be corrected is combined into a subtitle.

[0122] See also Figure 4 , when the speaker’s speech content is “Today’s large-screen optimization solution”, the subtitle can be generated: “Today’s large-screen optimization solution”.

[0123] It can be seen that by using the subtitle generation method of the above embodiment, the accuracy of the subtitles can be improved after the semantic text obtained by speech recognition is corrected by the screen image text.

[0124] In order to solve the problem that subtitles block the display content that users need to see, the embodiment of the present application provides a method for setting the position of subtitles, see Figure 5 , the method may include the following steps:

[0125] Step S210: receiving a subtitle request.

[0126] Step S220: In response to the subtitle request, when a voice stream is received, voice recognition is performed on the voice stream to obtain subtitles.

[0127] In some embodiments, the semantic text obtained by speech recognition can be directly used as subtitles.

[0128] In some embodiments, the Figure 2 The subtitle generation method shown obtains subtitles.

[0129] Step S230: Calculating the matching degree between the subtitle and the text in each image area, wherein the image area is a local display area of ​​a target image corresponding to the voice stream, and the target image includes multiple image areas.

[0130] In some embodiments, a target image corresponding to the voice stream may be obtained. A method for obtaining the target image may refer to Figure 2 Description.

[0131] In some embodiments, the target text in the target image may be recognized by an optical character recognition method, and the coordinates of the target text in the target image may be obtained.

[0132] In some embodiments, the target image can be divided into fixed image areas, such as two upper and lower image areas, which are respectively located in the upper and lower half of the display device, or two left and right image areas, which are respectively located in the left and right half of the display device. In such a fixed image area, there may be text on the boundary line. If the text is located on the boundary line of two image areas, the text can be set to belong to one of the image areas. For example, the text can be set to be located in the image area of ​​the foreword, where the foreword refers to the text before the boundary line, and the text located after the boundary line can be called the postword. In some embodiments, the image area can also be divided according to the text coordinates in the target image. For example, according to the fact that the text in the target image is concentrated at the top and bottom of the target image, and there is less text in the middle, the target image can be divided into three image areas at the top, middle and bottom. This method of dividing the image area according to the text coordinates in the target image can avoid the situation where the text in the target image is located at the dividing line of the two image areas.

[0133] In some embodiments, in each image area, a local display area can be divided as a subtitle display area for displaying subtitles. For example, in the upper half of the screen, the left half area can be set as the subtitle display area, and in the lower half of the screen, the left half area can also be set as the subtitle display area.

[0134] In some embodiments, after the target image is divided into a plurality of image regions, the text contained in each image region may be set according to the coordinates of the target text. In some embodiments, after the target image is divided into a plurality of image regions, text recognition may be performed in each image region to obtain the text contained in each image region.

[0135] In some embodiments, after obtaining the text contained in each image region, the matching degree between the subtitle and the text in each image region may be calculated.

[0136] An exemplary matching degree calculation method may be: performing word segmentation processing on the text on the target image to obtain multiple target word segments; performing word segmentation processing on the subtitles to obtain multiple subtitle word segments; calculating the matching degree of each subtitle word segment with the corresponding target word segment in each image area; adding or weighting all matching degrees in each image area to obtain the matching degree of the subtitles and the text in each image area.

[0137] For example, if the image region contains words that are consistent with the segmented text, the matching degree is 1.

[0138] If the image area does not contain words consistent with the segmented text, but contains similar segmented words, the matching degree is set to 0.1 to 0.9 according to the similarity, wherein the similarity can be determined according to some commonly used confusion sets. For example, in a confusion set, the similarities of texts A, B, and C are 0.8 and 0.6 respectively. If a segmented word obtained after speech recognition is segmented word A, the target image is divided into two image areas, neither of which contains segmented word A, the first image area contains text B, and the second image area contains text C, then the matching degree of segmented word A with the image area containing segmented word B is 0.8, and the matching degree with the image area containing segmented word C is 0.6.

[0139] If the image area does not contain words that are consistent with the segmented text, nor does it contain similar segmented words, the matching degree is 0.

[0140] Step S240: setting the coordinate area of ​​the subtitle within the image area where the matching degree is less than the maximum value.

[0141] In some embodiments, in the target image, if the matching degree of an image area is large, it indicates that the content of the voice stream is more relevant to the image area. Conversely, if the matching degree of an image area is small, it indicates that the content of the voice stream may not be relevant to the image area. Therefore, setting the coordinate area of ​​the subtitles in the image area with the smallest matching degree will minimize the impact on the user's viewing of the target image.

[0142] According to the above subtitle position setting method, an example of subtitle position setting is:

[0143] Exemplarily, the subtitles converted from the voice streams received at times t0, t1, t2, t20, t21, and t22 are:

[0144] subtitle(t0)="xxxxxxyyyyyyyzzzzaaabbbbbcccoosdkckkeffadkasdl";

[0145] subtitle(t1)="mmmnnnnnnwwwyyxxxxxuuu";

[0146] subtitle(t2)="ccdddddeeeeeeffffffgggg";

[0147] subtitle(t20)="Asdfkckweffa 1234kasdfkk 5678llldsf 0000";

[0148] subtitle(t21)="Cckkkwwdfaaaaa456 dkkasdf";

[0149] subtitle(t22)="1111hhhh kkkkk".

[0150] Among them, the segmentation result of subtitle(t0) is:

[0151] SEGMENT(subtitle(t0))=["xxxxxx","yyyyyy","zzzz","aaa","bbbbb","ccc","oosdkckkeffadkas dl"]

[0152] See also Figure 6 , the screen image is divided into two image areas: a first area 201 and a second area 202, wherein the first area 201 is the display area of ​​the upper half of the screen, and the second area 202 is the display area of ​​the lower half of the screen.

[0153] The target texts for the two image regions are:

[0154] SEGMENT(screen_text[1][1])=["xxx","zzzz","bbbb","ccc"],

[0155] SEGMENT(screen_text[1][2])=["mmm","nn","www","yy","xxxxx","uuu"],

[0156] SEGMENT(screen_text[1][3])=...,

[0157] SEGMENT(screen_text[1][4])=...,

[0158] SEGMENT(screen_text[2][1])=...,

[0159] SEGMENT(screen_text[2][2])=...,

[0160] SEGMENT(screen_text[2][3])=...

[0161] Here, SEGMENT(screen_text[1][1]) represents the target text in the first line of the first region 201, SEGMENT(screen_text[2][1]) represents the target text in the first line of the second region 202, and so on.

[0162] Calculate the matching degree p between each word in SEGMENT(screen_text[1][1]) and the word in SEGMENT(subtitle(t0)). According to the calculation method shown in step S260, the following calculation results are obtained:

[0163] p("xxx")=0.5; p("zzzz")=1; p("bbbb")=1; p("ccc")=1,...

[0164] Add the word matching degrees to get the similarity index between subtitle(t0) and screen_text[1][1]

[0165] P(screen_text[1][1],subtitle(t0))=3.5;

[0166] The same method is used to calculate: P(screen_text[1][2]subtitle(t0))=0;

[0167] P(screen_text[1][3],subtitle(t0))=0;

[0168] P(screen_text[1][4],subtitle(t0))=0;

[0169] P(screen_text[2][1],subtitle(t0))=0;

[0170] P(screen_text[2][2],subtitle(t0))=0;

[0171] P(screen_text[2][3],subtitle(t0))=0.

[0172] Based on this calculation result, it is determined that the matching degree of subtitle(t0) with screen_text[2] is less than that with screen_text[1], and the display position of subtitle(t0) screen_text[2] is sent to the video conferencing app of display device 2, so that display device 2 can display subtitles at the position of screen_text[2]. Alternatively, the server can also send the screen area screen_text[1] with the highest matching degree to the video conferencing app of display device 2, so that display device 2 can avoid displaying subtitles at the position of screen_text[1].

[0173] Similarly, the display positions of subtitle(t1) and subtitle(t2) are also the positions corresponding to screen_text[2], and the display positions of subtitle(t20), subtitle(t21), and subtitle(t22) are the positions corresponding to screen_text[1].

[0174] See Figure 7 , the display positions of subtitle(t0), subtitle(t1), and subtitle(t2) are in the second area 202, and the content that the audience needs to watch is in the first area 201. Therefore, the subtitles will not block the content that the audience needs to watch. Further, a subtitle area 204 can be set in the second area 202, and the display area of this subtitle area 204 is smaller than the area of the second area 202, reducing the occlusion of the second area 202.

[0175] See Figure 8 , the display positions of subtitle(t20), subtitle(t21), and subtitle(t22) are in the first area 201, and the content that the audience needs to watch is in the second area 202. Therefore, the subtitles will not block the content that the audience needs to watch. Further, a subtitle area 203 can be set in the first area 201, and the display area of this subtitle area 203 is smaller than the area of the first area 201, reducing the occlusion of the first area 201.

[0176] To further illustrate the subtitle generation method and the subtitle position setting method provided by the embodiments of the present application, the process of subtitle generation and display will be described below starting from when a user joins a video conference.

[0177] In some embodiments, a process of sharing a desktop can be seen in Fig. 9 , which is a timing diagram of sharing a desktop.

[0178] As Fig. 9 shown, the speaker can enter the conference number in the conference application on the display device 1. After receiving the conference number, the display device 1 can obtain its own device ID and send a conference joining request including the device ID of the display device 1 and the conference number to the server.

[0179] In some embodiments, after receiving a request to join a meeting from display device 1, the server can detect whether the meeting corresponding to the meeting number has been started. If not, the meeting can be started and the default meeting interface data can be returned to display device 1 so that display device 1 displays the default meeting interface. If it has been started and no participant has turned on the shared desktop function, the default meeting interface data can be returned to display device 1. If a participant has turned on the desktop sharing function, the current desktop data of the participant who has turned on the desktop sharing function can be sent to display device 1 so that display device 1 can display the current desktop of the participant who has turned on the desktop sharing function.

[0180] Fig. 9 In the example, the speaker is the first user to enter the conference corresponding to the conference number. The data returned by the server to the display device 1 according to the request to join the conference is the default conference interface data. After the display device 1 receives the default conference interface data, it can display the default conference interface corresponding to the default conference interface data.

[0181] In some embodiments, the default conference interface may be provided with a shared desktop control, a microphone control, and a subtitle control.

[0182] like Fig. 9 As shown, the process for the audience to join the conference corresponding to the above conference number is the same as the process for the speaker to join the conference.

[0183] In some embodiments, after joining the conference, the audience can operate the subtitle control on the display device 2 to enable the subtitle function of the display device 2, or the audience can operate the subtitle control after the speaker starts speaking. In response to the subtitle control being triggered, the display device 2 obtains its own device ID, generates a subtitle request including the device ID and the conference number, and sends the subtitle request to the server.

[0184] In some embodiments, after receiving a subtitle request, the server may initiate a subtitle generation task, wherein the subtitle generation task is configured to generate subtitles according to the subtitle generation method and subtitle position setting method introduced in the embodiments of the present application.

[0185] In some embodiments, after the audience joins the conference, the speaker can operate the shared desktop control on the display device 1 so that the audience can see the content displayed on the display device 1. In response to the shared desktop control being triggered, the display device 1 generates a shared desktop request including the conference number and the device ID of the display device 1, and sends the shared desktop request and the current screen image of the display device 1 to the server, or sets the current screen image of the display device 1 in the shared desktop request, so that only the shared desktop request needs to be sent to the server.

[0186] In some embodiments, after receiving a request to share the desktop and the current screen image of display device 1, the server may transmit the current screen image of display device 1 to display device 2. After receiving the screen image, display device 2 may display the screen image, thereby enabling display device 2 to share the desktop of display device 1.

[0187] After sharing the desktop, the operations performed by the speaker, display device 1, server, and display device 2 can be seen in Fig.10 , which is a timing diagram of subtitle generation according to some embodiments.

[0188] like Fig.10 As shown, after sharing the desktop, if the shared file has multiple pages, the speaker can operate the page turning control on the display device 1, and then operate the microphone control and input voice to explain the current page through voice. Of course, if the file shared by the speaker has only one page, there is no need to operate the page turning control, only the microphone control and then input voice.

[0189] Taking the example of a file shared by a speaker having multiple pages, after the speaker jumps to a certain page using a page turning control, the display device 1 can display the screen image after the page is turned, and send the screen image after the page is turned and a page turning message to the server.

[0190] In some embodiments, after receiving the screen image sent by display device 1, the server sends the screen image to display device 2, and display device 2 replaces the currently displayed image with the screen image sent by the server.

[0191] In some embodiments, after receiving the page turning message, the server obtains the text in the screen image after the page turning, and caches the text in the screen image after the page turning in blocks according to the partitioning method. Taking the pre-set partitioning method of dividing the screen image into two upper and lower image areas as an example, the text in the upper half of the screen is stored as a group of target texts in screen_text[1], and the text in the lower half of the screen is stored as another group of target texts in screen_text[2].

[0192] In some embodiments, to ensure the timeliness of subtitle display, the display device sends the acquired voice stream to the server for voice recognition each time the speaker pauses in inputting voice, and sends the acquired voice stream to the server for voice recognition after the next speaker's speech, thereby realizing cyclic voice recognition and improving the efficiency of subtitle display.

[0193] Usually, a pause in the speaker's input voice indicates that the speaker has finished a sentence. The conference application is pre-configured to upload the voice stream to the server if a pause interval is reached after receiving the voice. Exemplarily, the pause interval can be 0.4 seconds, that is, when receiving the sound, if no voice is received for 0.4 seconds after the last voice was received, the voice stream corresponding to the voice received this time will be sent to the server.

[0194] In some embodiments, after receiving the voice stream sent by the display device 1, the server performs voice recognition on the voice stream to obtain semantic text, and the semantic text includes multiple word segments.

[0195] In some embodiments, the server may modify each group of words in the semantic text according to multiple groups of target texts to obtain subtitles.

[0196] In some embodiments, the server may set the subtitle display area to the screen area where the least mapped target text is located based on the mapping relationship between the subtitle and each group of target texts. For example, the subtitle display area is set to the screen area corresponding to screen_text[2].

[0197] After obtaining the subtitles and the display area of ​​the subtitles, the server may send the subtitles and the display area to the display device 2, so that the display device 2 displays the subtitles in the display area.

[0198] To further describe the process of generating subtitles on the server, Fig.11 FIG. 4 shows a timing diagram of subtitle generation according to some embodiments. Fig.11 As shown, the server may be provided with the following functional modules: a video cache module, an image-text conversion module and a speech recognition module, wherein the video cache module is used to store screen images sent by the display device, the image-text conversion module is used to recognize texts in the screen images, and the speech recognition module is used to perform speech recognition on the speech stream.

[0199] The screen image after page turning sent by the display device 1 can be stored in the video buffer module. The page turning message can be transmitted to the image text conversion module and the voice recognition module in sequence.

[0200] After receiving the page turning message, the image text conversion module can obtain the latest screen image from the video cache module, divide the screen image into multiple image areas according to the text layout in the screen image, and then recognize the text in each image area and perform word segmentation on the recognized text.

[0201] After receiving the page turning message, the speech recognition module can start the speech recognition task. The speech recognition task can perform speech recognition on the speech stream sent by the display device to obtain word segmentation, and then correct the word segmentation obtained by speech recognition according to the word segmentation recognized by the screen image to obtain subtitles, and calculate the matching degree of the subtitles and the text in each image area, set the image area with the smallest matching degree as the display area of ​​the subtitles, and then send the subtitles and the display area of ​​the subtitles to the display device 2, so that the display device 2 displays the subtitles in the display area.

[0202] It can be seen from the above embodiments that the embodiments of the present application can improve the accuracy of subtitles by acquiring a target image corresponding to the voice stream and correcting the word segmentation obtained by voice recognition according to the text on the target image, so that the corrected target word segmentation corresponds to the text on the target image; further, by calculating the matching degree between the subtitles and the text in different image areas in the target image, the subtitles are set in the image area with the smallest matching degree, thereby reducing the impact of the subtitles on the understanding of the voice stream caused by the occlusion of the target image, thereby improving the user experience of viewing subtitles.

[0203] Since the above embodiments are all described by citing and combining with other embodiments, different embodiments have the same parts, and the same and similar parts between the embodiments in this specification can be referred to each other. No further detailed description is given here.

[0204] It should be noted that, in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such circuit structure, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the circuit structure, article or device including the element.

[0205] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the disclosure of the invention herein. The present application is intended to cover any modification, use or adaptation of the present invention, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the content of the claims.

[0206] The above implementation modes of the present application do not constitute a limitation on the protection scope of the present application.

Claims

1. A method for setting a subtitle position, It is characterized in that include: Receive subtitle requests; In response to the subtitle request, when receiving a voice stream, performing voice recognition on the voice stream to obtain subtitles; Calculating the matching degree between the subtitle and the text in each image area, wherein the image area is a local display area of ​​a target image corresponding to the voice stream, the target image includes a plurality of image areas, and the image area division method includes dividing the target image according to the distribution of the text in the target image; The coordinate area of ​​the subtitle is set within the image area where the matching degree is less than the maximum value.

2. The setting method according to claim 1, It is characterized in that The text of the target image is the text recognized by the target image through an optical character recognition method.

3. The setting method according to claim 1, It is characterized in that Calculating the matching degree between the subtitle and the text in each image region, including: Calculating the matching degree of each subtitle segmentation word with the corresponding target segmentation word in each image area, wherein the subtitle segmentation word is the segmentation word obtained by segmenting the subtitle word, and the target segmentation word is the segmentation word obtained by segmenting the text in the image area; All matching degrees in each image region are weighted to obtain the matching degree between the subtitle and the text in each image region.

4. The setting method according to claim 1, It is characterized in that Each image region includes a subtitle region whose area is smaller than that of the image region, and the subtitle coordinate region is one of the subtitle regions.

5. The setting method according to claim 1, It is characterized in that The target image is the screen image obtained in response to each time receiving a message that the target display device triggering the shared desktop control sends a screen image and a message indicating a preset screen change to the server.

6. The setting method according to claim 5, It is characterized in that The message indicating the change of the preset screen includes a page turning message.

7. The setting method according to claim 1, It is characterized in that Setting the coordinate area of ​​the subtitle within the image area where the matching degree is less than the maximum value includes: The coordinate area of ​​the subtitle is set within the image area with the minimum matching degree.

8. The setting method according to claim 1, It is characterized in that Setting the coordinate area of ​​the subtitle within the image area where the matching degree is less than the maximum value includes: If the image area with a matching degree less than the maximum value includes a historical image area, the historical image area is set as the coordinate area of ​​the subtitles, wherein the historical image area is a display area of ​​the subtitles corresponding to the previous voice stream.

9. A server, It is characterized in that For implementing the method according to any one of claims 1 to 8, the server comprises: A video cache module, used for storing screen images sent by a display device; An image-to-text conversion module, used for recognizing text in the screen image; The speech recognition module is used to perform speech recognition on the speech stream.

Citation Information

Patent Citations

  • Method and device for prompting in video meeting

    CN102036051A

  • Method and device for detecting and extracting subtitle

    CN107480670A