Comment display method, device, electronic device and storage medium

By using the voice decoding network on the server side to identify the voice data of the target video and comparing it with the video content to generate sensitive words, the problem of manual settings in the prior art is solved, and the automatic generation of sensitive words with high accuracy and comment data filtering is achieved.

CN114512130BActive Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011279992.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-16
Publication Date
2025-06-06
Estimated Expiration
2040-11-16

AI Technical Summary

Technical Problem

Sensitive words used in the prior art for commenting data filtering rely on manual implementation, resulting in low accuracy and cumbersome addition process.

Method used

By obtaining the voice data related to the target video on the server side, using the built voice decoding network for voice recognition, combining the video content data comparison to generate sensitive words, and sending filtered comment data when the terminal initiates a blocking request.

Benefits of technology

The automatic generation of sensitive words is realized, the accuracy of sensitive words is improved, the tedious process of manual addition is avoided, and the problem of relying on manual settings for comment data filtering is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114512130B_ABST
    Figure CN114512130B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a comment display method, device, electronic device and storage medium, which relates to the field of computer technology. Among them, the comment display method is applied to a server, and the method includes: obtaining voice data related to a target video; based on a constructed voice decoding network, performing voice recognition on the voice data related to the target video to obtain a voice recognition result; comparing the video content data of the target video with the voice recognition result, and generating sensitive words of the target video according to the comparison result; upon receiving a sensitive word shielding request initiated by a terminal, sending comment data filtered according to the sensitive words of the target video to the terminal, so that the terminal displays the filtered comment data during the playback of the target video. The embodiment of the present application uses automatic voice recognition technology to solve the problem that sensitive words used for comment data filtering in the prior art rely on manual implementation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to a comment display method, device, electronic device and storage medium. Background Art

[0002] With the widespread use of the Internet, people can watch videos through the Internet and comment on them while watching them. Specifically, during the video playback process, the terminal can obtain video-related comment data from the server and display it in the video playback interface, so that people can read comments about the video while watching it.

[0003] It is understandable that people often do not want to be spoiled in advance when watching a video. For this reason, before sending the comment data related to the video to the terminal, the server usually uses the sensitive words of the video to filter the comment data so that the terminal no longer displays the comment data that may spoil the content to people, thereby reducing the possibility of people being spoiled in advance.

[0004] At present, sensitive words used for comment data filtering mainly rely on manual settings by staff of video publishing platforms. Not only are they easily affected by subjective factors of staff, resulting in low accuracy of sensitive words, but the adding process is also too cumbersome. Summary of the invention

[0005] The embodiments of the present application provide a comment display method, device, electronic device and storage medium, which can solve the problem that sensitive words used for comment data filtering in related technologies rely on manual implementation. The technical solution is as follows:

[0006] According to one aspect of an embodiment of the present application, a comment display method is applied to a server, and the method includes: obtaining voice data related to a target video; based on a constructed voice decoding network, performing voice recognition on the voice data related to the target video to obtain a voice recognition result; comparing the video content data of the target video with the voice recognition result, and generating sensitive words of the target video according to the comparison result; upon receiving a sensitive word shielding request initiated by a terminal, sending comment data filtered according to the sensitive words of the target video to the terminal, so that the terminal displays the filtered comment data during the playback of the target video.

[0007] According to one aspect of an embodiment of the present application, a comment display device is applied to a server, and the device includes: a data acquisition module, which is used to acquire voice data related to a target video; a voice recognition module, which is used to perform voice recognition on the voice data related to the target video based on a constructed voice decoding network to obtain a voice recognition result; a data comparison module, which is used to compare the video content data of the target video with the voice recognition result, and generate sensitive words of the target video according to the comparison result; a data sending module, which is used to send comment data filtered according to the sensitive words of the target video to the terminal when a shielding request sent by the terminal is received, so that the terminal displays the filtered comment data during the playback of the target video.

[0008] According to one aspect of an embodiment of the present application, an electronic device includes: at least one processor, at least one memory, and at least one communication bus, wherein the memory stores computer-readable instructions, and the processor reads the computer-readable instructions in the memory through the communication bus; when the computer-readable instructions are executed by the processor, the comment display method as described above is implemented.

[0009] According to one aspect of an embodiment of the present application, a storage medium stores a computer program thereon, and when the computer program is executed by a processor, the comment display method as described above is implemented.

[0010] The beneficial effects of the technical solution provided by this application are:

[0011] In the above technical solution, the server obtains relevant voice data of the target video, and performs voice recognition on the voice data related to the target video through the constructed voice decoding network to obtain the voice recognition result, and generates the sensitive words of the target video according to the comparison result by comparing the video content data of the target video with the voice recognition result. In this way, when the terminal initiates a sensitive word shielding request, the server can send the comment data filtered according to the sensitive words of the target video to the terminal, thereby ensuring that the terminal displays the filtered comment data during the process of playing the target video. Therefore, the automatic generation of sensitive words is realized by using automatic voice recognition technology, which effectively solves the problem that the sensitive words used for comment data filtering in the prior art rely on manual implementation. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.

[0013] Figure 1 It is a schematic diagram of the implementation environment involved in this application.

[0014] Figure 2The figure is a flowchart of a comment display method according to an exemplary embodiment.

[0015] Figure 3 for Figure 2 A schematic diagram of speech recognition based on an acoustic model and a language model according to the embodiment.

[0016] Figure 4 yes Figure 2 The corresponding embodiment is a flowchart of step 35 in an embodiment.

[0017] Figure 5 yes Figure 2 The corresponding embodiment is a flowchart of step 37 in an embodiment.

[0018] Figure 6 for Figure 5 A schematic diagram of a video playback interface according to a corresponding embodiment.

[0019] Figure 7 yes Figure 5 A flowchart of an embodiment corresponding to step 375 in an embodiment.

[0020] Figure 8 yes Figure 2 The corresponding embodiment is a flowchart of step 33 in an embodiment.

[0021] Fig. 9 for Figure 8 A schematic diagram of a feature vector of a speech frame involved in the corresponding embodiment.

[0022] Fig.10 for Figure 8 A schematic diagram of speech recognition based on speech frames, states, phonemes, and words involved in the corresponding embodiments.

[0023] Fig.11 yes Figure 8 The corresponding embodiment is a flowchart of step 331 in an embodiment.

[0024] Fig.12 for Fig.11 A schematic diagram of the frame processing involved in the corresponding embodiment.

[0025] Fig.13 yes Fig.11 The corresponding embodiment is a flowchart of step 333 in an embodiment.

[0026] Fig.14 for Fig.13 A schematic diagram of a state network constructed based on a hidden Markov model according to a corresponding embodiment.

[0027] Fig.15 yes Fig.13The flowchart of step 3331 in one embodiment corresponds to the embodiment.

[0028] Fig.16 for Fig.15 A schematic diagram of implementing path search based on a stateful network according to a corresponding embodiment.

[0029] Fig.17 for Fig.15 A schematic diagram of the state probabilities of speech frames belonging to different states involved in the corresponding embodiments.

[0030] Fig.18 yes Fig.11 The corresponding embodiment is a flowchart of step 337 in one embodiment.

[0031] Fig.19 The figure is a flowchart of another comment display method according to an exemplary embodiment.

[0032] Fig. 20 The diagram is a schematic structural diagram of a comment display device according to an exemplary embodiment.

[0033] Fig.21 The figure is a schematic diagram of a server structure according to an exemplary embodiment.

[0034] Fig. 22 The diagram is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0035] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present invention.

[0036] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0038] The following is an introduction and explanation of several terms involved in this application:

[0039] Automatic Speech Recognition (ASR) technology refers to the use of machines (such as servers) to automatically convert speech into text.

[0040] Phonemes, according to the pronunciation characteristics of Chinese, the initial consonants and final vowels of words are defined as phonemes. For English, it is composed of 39 phonemes.

[0041] State, a more basic phonetic unit than phoneme. Usually a phoneme can be divided into three states.

[0042] Mel Frequency Cepstral Coefficents (MFCCs) are the coefficients that make up the Mel frequency cepstrum. The frequency bands of the Mel frequency cepstrum are equally divided on the Mel scale. It is more approximate to the human auditory system than the linearly spaced frequency bands used in the normal logarithmic cepstrum. Therefore, the Mel frequency cepstrum coefficient feature (MFCC feature) is an acoustic feature widely used in automatic speech recognition technology.

[0043] Voice Activity Detection (VAD) is generally used to identify the presence and disappearance of voice in audio signals.

[0044] As mentioned earlier, the sensitive words used for comment data filtering mainly rely on manual settings by the staff of the video publishing platform.

[0045] Specifically, before a video is released, the staff will browse the video to be released so that they can formulate corresponding sensitive words for the video to be released based on their own understanding of the video content of the video to be released, and follow the steps for adding sensitive words to gradually complete the manual addition of sensitive words to the video to be released in the relevant interface provided by the video publishing platform.

[0046] On the one hand, since the staff are faced with a massive amount of videos to be released, it is impossible for them to watch each video carefully. They can only quickly browse the complete video to be released, or even quickly browse several video clips of the video to be released. This may result in the staff's understanding of the video content of the video to be released being inaccurate. In addition, different staff members may have different understandings of the video content of the same video to be released. Therefore, the sensitive words formulated by the staff of the video publishing platform are easily affected by the subjective factors of the staff and are not accurate enough.

[0047] On the other hand, in the process of adding sensitive words, staff are usually required to manually open several interfaces and trigger relevant operations in each interface in order to add sensitive words. Especially when staff are faced with a large number of videos to be released, they need to manually add sensitive words for each video to be released, and the adding process is quite cumbersome.

[0048] From the above, we can see that how to avoid sensitive words from relying on manual implementation remains to be solved.

[0049] To this end, the present application provides a comment display method, device, electronic device and storage medium based on automatic speech recognition technology, aiming to solve the above technical problems of the prior art.

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the automatic speech recognition technology involved in the implementation mode of the present application will be further described in detail below with reference to the accompanying drawings.

[0051] Figure 1 The schematic diagram of an implementation environment involved in a comment display method includes a terminal 100 and a server 200 .

[0052] Specifically, the terminal 100 can be operated by a client with a video playback function, and can be an electronic device such as a desktop computer, a laptop computer, a tablet computer, a smart phone, etc., which is not limited here.

[0053] Among them, the client has a video playback function, for example, a media player, a browser, etc., which can be in the form of an application or a web page. Accordingly, the user interface for video playback on the client, that is, the video playback interface, can be in the form of a program window or a web page, which is not limited here.

[0054] The server end 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. For example, in this implementation environment, the server end 200 may provide the terminal 100 with cloud storage services for videos and comment data and / or comment data filtering services.

[0055] The terminal 100 and the server 200 establish a communication connection through wireless or wired means, so as to realize data transmission between the terminal 100 and the server 200 through the communication connection. For example, the transmitted data includes but is not limited to comment data.

[0056] Through the interaction between the terminal 100 and the server 200, the terminal 100 initiates a sensitive word shielding request to the server 200, thereby informing the server 200 of the target video for which the terminal 100 requests sensitive word shielding.

[0057] In this way, after receiving the sensitive word shielding request, the server 200 will filter the comment data related to the target video according to the sensitive words of the target video, and send the filtered comment data to the terminal 100, so that the terminal 100 can display the filtered comment data during the playback of the target video.

[0058] In the above process, the sensitive words of the target video are automatically generated by the server side 200 using automatic speech recognition technology, that is, based on the constructed speech decoding network, speech recognition is performed on the speech data related to the target video to obtain a speech recognition result, and the sensitive words of the target video are generated based on the comparison result of the speech recognition result with the video content data of the target video, thereby avoiding the sensitive words from relying on manual implementation.

[0059] See also Figure 2 The present application embodiment provides a comment display method, which is applicable to Figure 1 The server side 200 of the implementation environment is shown.

[0060] The method can be executed by the server side, or it can be understood that it is executed by an application program running in the server side. In the following method embodiment, for the convenience of description, the execution subject of each step is described as the server side, but this is not limited to this.

[0061] like Figure 2 As shown, the method may include the following steps:

[0062] Step 31, obtaining voice data related to the target video.

[0063] The target video refers to the video provided by the video publishing platform to the user. It can be a complete video or a collection of several video clips in the complete video. The types of the target video include movies, TV series, variety shows, game live broadcasts, etc. The voice data related to the target video refers to the voice in the video. Based on the type of the target video, the voice data includes dialogues and narrations between actors in movies or TV series; the host's opening remarks, speeches, and dialogues between participants in variety shows; dialogues between virtual characters in game live broadcasts, dialogues between players, the host user's opening remarks, speeches, and dialogues between the host user and the live broadcast user, etc.

[0064] It should be understood that the video publishing platform will deploy a server, which can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, to provide users with cloud storage services for massive videos. It can also be understood that the server is a "video source".

[0065] Of course, according to actual operational needs, the server side of the cloud storage service that provides massive videos and the server side of the cloud storage service and comment data filtering service that provide comment data can also be deployed on the same physical server, the same server cluster or the same distributed system. This embodiment does not constitute a specific limitation here.

[0066] Based on this, regarding the acquisition of the voice data related to the target video, in a possible implementation, the voice data related to the target video can be acquired from the video source. In a possible implementation, the voice data related to the target video can be read from a local memory.

[0067] Step 33, based on the constructed speech decoding network, speech recognition is performed on the speech data related to the target video to obtain a speech recognition result.

[0068] Among them, the speech decoding network is constructed based on the acoustic model and language model. Figure 3 As shown, the acoustic model is modeled and trained using phonemes as modeling units, and the language model is modeled and trained based on the probability distribution of the language itself.

[0069] It is explained here that according to the contextual relationship of phonemes, phonemes can be divided into monophones, diphones and triphones. Among them, monophones only consider themselves when in use, diphones only consider the contextual relationship between them and the previous phoneme or the following phoneme when in use, and triphones consider both the contextual relationship between them and the previous phoneme and the contextual relationship between them and the following phoneme when in use. Therefore, the monophone acoustic model obtained by modeling and training with the monophone as the modeling unit, the diphone acoustic model obtained by modeling and training with the diphone as the modeling unit, and the triphone acoustic model obtained by modeling and training with the triphone as the modeling unit, which is not limited in this embodiment.

[0070] After obtaining the language model of the acoustic model, a speech decoding network for automatic speech recognition can be formed. Figure 3 , the speech data related to the target video is used as audio input, and features are first extracted and then decoded, that is, based on the acoustic model, the speech data related to the target video is recognized as phonemes; based on the language model, the phonemes recognized by the acoustic model are further recognized as words, thereby obtaining the speech recognition result. It can also be understood that the speech recognition result is actually a text containing several words formed by converting the speech in the target video.

[0071] In one possible implementation, the acoustic model is a hidden Markov model.

[0072] Step 35, compare the video content data of the target video with the speech recognition result, and generate sensitive words of the target video according to the comparison result.

[0073] Among them, the video content data of the target video is used to describe the video content of the target video. Based on the type of the target video, the video content data includes but is not limited to: plot synopsis of movies and TV series, cast list and related introductions to the characters played by the actors, plot introduction of TV series, etc.; content synopsis of variety shows, list of participants, etc.; content synopsis of games, list of virtual characters and related introductions, list of players, etc.

[0074] In a possible implementation, the video content data includes a number of reference words, which are used to describe the core content of the target video. For example, the target video is Tom and Jerry, and the reference words in the video content data of the target video include at least Tom and Jimmy. It should be noted that the reference words in the video content data can be pre-extracted by the staff of the video publishing platform, or can be pre-provided by the video provider, and then stored in the server side to facilitate the generation of sensitive words, which is not limited here.

[0075] Based on this, the comparison between the video content data and the speech recognition results is essentially to compare the reference words in the video content data with the words in the speech recognition results to generate sensitive words for the target video.

[0076] In a possible implementation, the comparison result is used to indicate whether the speech recognition result contains a word that matches the reference word. Figure 4 As shown, step 35 may include the following steps: step 351, searching for words matching the reference words in the speech recognition results according to the reference words in the video content data. Step 353, using the searched words as sensitive words of the target video.

[0077] For example, assuming that word a matching reference word a is searched in the speech recognition results, the comparison result indicates that the speech recognition results contain word a matching reference word a. Then, based on the comparison result, the sensitive word of the target video can be generated, that is, the word a indicated by the comparison result is used as the sensitive word a of the target video.

[0078] In other words, the sensitive words of the target video reflect the core content of the target video. If the comment data contains the sensitive words, it can be considered that the comment data may reflect the core content of the target video, and thus there is a possibility of early spoilers, and should be blocked.

[0079] It is worth mentioning that please refer to Figure 3 On the one hand, the automatic speech recognition based on the speech decoding network can use the speech recognition results as the basis for generating sensitive words. On the other hand, the sensitive words generated based on the speech recognition results can, in turn, be used as training data for training the speech decoding network. Thus, the self-learning of the speech decoding network can be realized, which can effectively expand the application scenarios actually covered by speech recognition, and thus help improve the accuracy of speech recognition.

[0080] For example, the target video is a TV series. Assume that the speech recognition results based on which sensitive words are generated come from the dialogues of different actors in the TV series. Accordingly, the training data can come from people who speak very fast, or from people with obvious accents, or from people who speak unclearly, thereby achieving the purpose of enriching the application scenarios actually covered by speech recognition.

[0081] Step 37, when receiving the sensitive word shielding request initiated by the terminal, sending the comment data filtered according to the sensitive words of the target video to the terminal.

[0082] Specifically, in a possible implementation, such as Figure 5 As shown, step 37 may include the following steps:

[0083] Step 371: Receive a sensitive word shielding request, and determine the target video currently being played by the terminal according to the sensitive word shielding request.

[0084] For the terminal, a sensitive word blocking entrance will be provided to the user, so that the user can trigger related operations through the sensitive word blocking entrance, and then the terminal detects the related operations and initiates a sensitive word blocking request to the server, thereby enabling the anti-spoiler function for the user.

[0085] For example, if Figure 6 As shown, when the target video is played in full screen on the video playback interface 401, if the user wants to turn on the anti-spoiler function, the user can click anywhere in the video playback interface 401. At this time, after detecting the click operation, the terminal will display a shielding control interface 402 accordingly. The shielding control interface 402 is provided with a clickable sensitive word shielding option 403, namely "shielding on".

[0086] When the user clicks on the sensitive word shielding option 403, the terminal responds to the detected click operation, adds the video identifier of the target video to the sensitive word shielding request, and then initiates the sensitive word shielding request to the server. The sensitive word shielding option 403 is regarded as a sensitive word shielding entry, and the click operation is regarded as a related operation triggered at the sensitive word shielding entry.

[0087] Of course, in other possible implementations, the sensitive word blocking option can also be displayed independently in the video playback interface, which does not constitute a specific limitation here.

[0088] Then, for the server side, it can extract the video identifier of the target video from the received sensitive word blocking request, and use it to determine the target video currently played by the terminal and requesting sensitive word blocking, and then provide the terminal with filtering services for relevant comment data about the target video.

[0089] It should be noted that, depending on the input components configured in the terminal (e.g., the touch layer covering the display screen, the mouse, the keyboard, etc.), the specific behaviors of the related operations triggered by the user at the control entrance may also be different. For example, for a smartphone inputted by a touch layer, the operation may be a gesture operation such as clicking and sliding, while for a laptop equipped with a mouse, the operation may be a mechanical operation such as dragging, single-clicking, double-clicking, etc., which are not specifically limited here.

[0090] Step 373, obtaining comment data related to the target video after the current playback time point.

[0091] That is to say, the sensitive word filtering is performed on the comment data related to the target video that has not yet been sent to the terminal. The comment data includes: the historical comment data uploaded by the user's terminal for the target video after the current playback time point, and the new comment data uploaded by the user's terminal for the target video after the current playback time point.

[0092] For example, the target video has a playback duration of 1 hour. Assuming that the user initiates a sensitive word blocking request to the server when watching the target video for 40 minutes, the current playback time point is 40 minutes. At this time, the server obtains the comment data for the target video between the 40th minute and the 60th minute.

[0093] Step 375: filter the acquired comment data according to the sensitive words of the target video.

[0094] like Figure 7 As shown, in a possible implementation, step 375 further includes the following steps:

[0095] Step 3751, matching the sensitive words of the target video with the words in the comment data to obtain the matching degree of the sensitive words in the comment data.

[0096] The sensitive word matching degree of the comment data = the number of words matching the sensitive words in the comment data / the total number of words in the comment data.

[0097] Step 3753: If the sensitive word matching degree of the comment data exceeds the matching degree threshold, the comment data is filtered.

[0098] For example, the target video is Tom and Jerry, and the sensitive words of the target video include Tom and Jimmy. Assuming the comment data is "Jimmy / is / about / to / be / caught / by / Tom", the sensitive word matching degree of the comment data = 2 / 6. " / " represents the separator between words in the comment data.

[0099] Then, if the sensitive word matching degree 2 / 6 of the comment data exceeds the matching degree threshold, it means that the comment data is consistent with the core content of the target video or the matching degree is too high, and the comment data will be filtered by the server.

[0100] Here, the matching degree threshold can be flexibly set according to the actual needs of the application scenario, and this embodiment does not limit this.

[0101] Step 377, sending the filtered comment data to the terminal.

[0102] As a result, the terminal can display the filtered comment data during the playback of the target video, thereby achieving the effect of shielding the comment data that may spoil the video in advance.

[0103] It should be added that in one possible implementation method, the comment data can be displayed in the form of "bullet screen" in the terminal. Therefore, the comment data is also called barrage data. The user can use the terminal to set the relevant parameters for displaying the barrage data. For example, the relevant parameters include transparency, font size, display area, barrage speed, etc., which are not limited here.

[0104] Through the above process, the automatic generation of sensitive words is realized by using automatic speech recognition technology. On the one hand, the formulation of sensitive words is no longer affected by the subjective factors of the staff of the video publishing platform, which improves the accuracy of sensitive words. On the other hand, the addition of sensitive words no longer depends on the manual completion of the staff, avoiding the problem of the manual addition of sensitive words being too cumbersome.

[0105] In a possible implementation, the complete video can be divided into multiple video clips, which are connected to each other in time, and each video clip has corresponding comment data (i.e., the comment message displayed during the playback of the video clip). In step 371, multiple video clips after the currently played video clip can be determined as the target video, that is, the target video is essentially a collection of several video clips in the complete video, and it can also be understood that the target video is a part of the complete video. At this time, voice data related to the multiple video clips after the currently played video clip are obtained, and voice recognition is performed on the voice data to obtain a voice recognition result, and the voice recognition result is used in step 35 to compare with the video content data of the target video, thereby reducing the resource consumption of voice recognition and improving the recognition efficiency of spoiler comment data. For example, the complete video has multiple video clips A, B, C and D that are connected to each other in time. If the current video clip B is played, it is determined that the target video includes video clips C and D, so that the voice recognition is performed on the voice data related to video clips C and D, rather than the voice data related to the complete video, so as to reduce the resource consumption of voice recognition.

[0106] In one possible implementation, the target video can be divided into multiple target sub-videos, and voice data related to the multiple target sub-videos are obtained respectively, and the voice data related to the multiple target sub-videos are compared with the video content data of the target video respectively to obtain sensitive words corresponding to the multiple target sub-videos, and the comment data corresponding to each target sub-video is filtered for sensitive words, wherein the comment data corresponding to each target sub-video in the multiple target sub-videos is filtered according to the sum of the sensitive words of each target sub-video and the subsequent multiple target sub-videos as in the above step 37. For example, a target video has multiple target sub-videos A, B, C and D that are connected to each other in time, and the multiple target sub-videos have sensitive words A1, B1, C1 and D1 respectively, and the multiple target sub-videos have corresponding comment data A2, B2, C2 and D2 respectively. Then, when filtering A2, sensitive words A1, B1, C1 and D1 can be used, when filtering B2, sensitive words B1, C1 and D1 can be used, when filtering C2, sensitive words C1 and D1 can be used, and when filtering D2, sensitive word D1 can be used. In this way, the accuracy of sensitive word filtering can be further improved.

[0107] See also Figure 8 , a possible implementation method is provided in the embodiment of the present application, step 33 includes the following steps:

[0108] Step 331 , extracting acoustic features from the speech frames included in the speech data to obtain feature vectors of the speech frames.

[0109] It should be understood that since each person has different speech speed and intonation, for example, women speak faster and men speak slower, even if the same sentence is spoken, the generated speech has different frequencies, which can also be understood as the speech data is an indefinite length time sequence. Here, the inventors realize that the indefinite length time sequence of speech data causes speech recognition to be unable to determine where to start from the speech data, and will form feature vectors with different processing dimensions, which is not conducive to the subsequent processing of the acoustic model.

[0110] Therefore, in this embodiment, on the one hand, the acoustic features are extracted from the speech frames in the speech data, that is, the acoustic features can be extracted only after the speech data is divided into a number of speech frames with a fixed frame length.

[0111] In a possible implementation, segmentation is implemented according to a set frame length, for example, the frame length is set to 25ms. In a possible implementation, segmentation is implemented using a moving window function. In a possible implementation, segmentation is implemented by stacking speech frames. In a possible implementation, segmentation is implemented by introducing the difference between adjacent speech frames as a new speech frame.

[0112] On the other hand, since the speech frame exists in the form of a waveform, and the waveform has almost no descriptive ability in the time domain, the waveform must be transformed, that is, the waveform must be transformed into a feature vector. Here, acoustic feature extraction actually refers to transforming the waveform so that the speech frame in the form of a waveform is transformed into a feature vector, and then the speech frame is described by the feature vector. Fig. 9 As shown, the speech frame is described as a 12-dimensional feature vector. Fig. 9 The depth of the color of the middle color block indicates the size of the vector value.

[0113] In a possible implementation, the acoustic feature is a Mel-frequency cepstral coefficient feature. In a possible implementation, the acoustic feature is a Mel-frequency cepstral coefficient (MFCC) feature and a fundamental frequency (PITCH) feature.

[0114] Step 333, based on the acoustic model, identify the states divided by phonemes for the feature vector of the speech frame, determine the state to which the speech frame belongs, and obtain the state sequence corresponding to the speech data according to the state to which the speech frame belongs.

[0115] As mentioned above, words are composed of phonemes, and states are more basic speech units than phonemes. A phoneme is usually divided into three states. Fig.10 As shown, several speech frames are first combined into a state, for example, the 1st speech frame to the 6th speech frame form state S1029, and then several states are combined into a phoneme, for example, states S1029, S124, and S561 form the phoneme ay, and finally several phonemes are combined into a word.

[0116] In this embodiment, the acoustic model actually establishes the optimal correspondence between speech frames and states. Then, based on the acoustic model, the state of the speech frame described by the feature vector can be determined according to the optimal correspondence between the speech frames and states. Then, based on the state of each speech frame in the speech data, the state sequence corresponding to the speech data can be obtained.

[0117] Step 335 , generating a phoneme set corresponding to the speech data from a state sequence corresponding to the speech data according to the states divided by the phonemes.

[0118] For example, the word "Tom" is composed of phonemes t, ang, m, and u, where t represents the initial consonant of "Soup", ang represents the final consonant of "Soup", m represents the initial consonant of "Mu", and u represents the final consonant of "Mu". Assume that the phoneme t is divided into states S1, S2, and S3, the phoneme ang is divided into states S4, S5, and S6, the phoneme m is divided into states S7, S8, and S9, and the phoneme u is divided into states S10, S1, and S2.

[0119] Then, for the word "Tom", the state sequence corresponding to the speech data is {S1, S2, S3; S4, S5, S6; S7, S8, S9; S10, S1, S2}. Correspondingly, according to the state divided by phonemes, the phoneme set corresponding to the speech data is {t, ang, m, u}.

[0120] Step 337, input the phoneme set corresponding to the speech data into the language model, predict the word to which the phoneme belongs, and obtain the speech recognition result.

[0121] As mentioned above, the language model is modeled and trained based on the probability distribution of the language itself. It is understandable that if there are many phonemes in the phoneme set, if the probability distribution of the language itself is not considered, that is, if the contextual relationship of the phonemes is not considered, then there may be many ways to combine the phonemes, which will inevitably lead to incorrect speech recognition results.

[0122] In this embodiment, the language model essentially predicts the words to which the phonemes belong based on the probability distribution of the language itself. It can also be considered that the more it conforms to the probability distribution of the language itself, the higher the accuracy of the predicted words to which the phonemes belong. Therefore, based on the language model, the correct speech recognition results can be predicted.

[0123] Under the effect of the above-mentioned embodiments, automatic speech recognition is realized, so that the speech recognition result can be used as the basis for generating sensitive words, which is conducive to realizing the automatic generation of sensitive words.

[0124] See also Fig.11 , a possible implementation method is provided in the embodiment of the present application, step 331 includes the following steps:

[0125] Step 3311, perform frame processing on the voice data to obtain a plurality of voice frames contained in the voice data.

[0126] In a possible implementation, the frame processing refers to dividing the voice data into a plurality of voice frames of a set frame length. Here, the set frame length can be flexibly adjusted according to the actual needs of the application scenario, and this embodiment does not limit this.

[0127] Furthermore, in order to improve the reliability of the frame processing according to the set frame length, in a possible implementation, a moving window function is used to implement the frame processing so that two adjacent speech frames to be divided overlap each other. Fig.12 As shown, based on the moving window function with a frame length of 25 ms and a frame shift of 10 ms, there is an overlap of 15 ms between two adjacent segmented speech frames.

[0128] Therefore, based on the frame processing, each speech frame has a relatively short and fixed duration, which is not only sufficient for speech recognition, for example, the 100th to 105th speech frames can be recognized as the initial consonant t, and the 106th to 115th speech frames can be recognized as the final ang, but also facilitates the subsequent processing of the acoustic model.

[0129] Furthermore, before step 3311, in a possible implementation, the method may further include the following steps:

[0130] Perform voice endpoint detection on voice data.

[0131] Accordingly, step 3311 may include the following steps:

[0132] The voice data after voice endpoint detection is framed.

[0133] It should be understood that in daily life, people can quickly understand each other's language and understand what they want to express, but for the server side, automatic speech recognition technology is a complex process and is easily interfered by external factors such as background noise and environmental reverberation. For example, the accuracy of speech recognition in the subway is significantly reduced.

[0134] To this end, before performing frame processing, voice endpoint detection is performed on the voice data to extract the silent parts at the beginning and end of the voice data, thereby eliminating the silent parts in the voice data, reducing interference with subsequent voice recognition, and helping to improve the accuracy of voice recognition.

[0135] Step 3313: For each speech frame included in the speech data, extract the Mel-frequency cepstral coefficient feature (MFCC feature) of the speech frame.

[0136] Step 3315, calculating the feature vector of the speech frame according to the extracted Mel-frequency cepstral coefficient features.

[0137] The calculation process includes but is not limited to: calculating the mean of the MFCC features of all speech frames, calculating the difference between the MFCC features of all speech frames and the mean, concatenating the MFCC features of adjacent speech frames, dimensionality reduction, correlation elimination, etc., which are not described in detail here.

[0138] Therefore, speech frames that exist in the form of waveforms and have no descriptive ability can be represented as MFCC features with descriptive ability, and used as feature input for speech recognition, thereby enabling speech recognition.

[0139] See also Fig.13 , a possible implementation method is provided in the embodiment of the present application, and step 333 may include the following steps:

[0140] Step 3331, based on the feature vector of the speech frame, search for the state path that best matches the speech data in the state network constructed based on the hidden Markov model.

[0141] In this embodiment, the acoustic model is a hidden Markov model. Accordingly, the recognition of the phoneme division state of the feature vector of the speech frame based on the acoustic model is essentially realized by a state network constructed based on the hidden Markov model.

[0142] Specifically, the state network is pre-constructed according to the states divided by the phonemes, and the state network includes several state branches, each of which is composed of several path nodes and state paths connected between adjacent path nodes. Among them, each path node stores a state, that is, the path node is used to indicate the state divided by the phonemes; the state path has directionality, for example, the state path points from path node L1 to path node L2, which means that the state indicated by path node L1 is transferred to the state indicated by path node L2.

[0143] For example, if Fig.14 As shown, the state on the path node 501 is S1, the state on the path node 502 is S2, and the state path 503 points from the path node 501 to the path node 502, which is expressed as 501->502, that is, the state is transferred from S1 to S2.

[0144] Based on the above, after a state network is constructed based on a hidden Markov model, a state path that best matches the speech data can be searched for each state path in the state network.

[0145] Step 3333, generating a state sequence corresponding to the voice data according to the state of the voice frame indicated by each path node in the searched state path.

[0146] Still using the above example to illustrate, for the word "Tom", the state path that best matches the voice data is 501->502->503->504->505->506->507->508->509->510->501->502. Then, according to the state of the voice frame indicated by each path node in the state path, the state sequence corresponding to the voice data is {S1, S2, S3; S4, S5, S6; S7, S8, S9; S10, S1, S2}.

[0147] Under the effect of the above-mentioned embodiment, state recognition based on the acoustic model is realized through path search in the state network constructed based on the hidden Markov model, which is used as the input of the subsequent language model, so that automatic speech recognition can be realized and the correctness of speech recognition is fully guaranteed.

[0148] See also Fig.15 , in an embodiment of the present application, a possible implementation is provided. Step 3331 may include the following steps:

[0149] Step 3331a, based on the states indicated by each path node in the state network, calculate the observation probabilities of the speech frame corresponding to different states according to the feature vector of the speech frame, and calculate the transition probabilities of the speech frame transferring between the corresponding different states.

[0150] Step 3331c, calculate the state probabilities of the speech frame belonging to different states according to the observation probabilities and the transition probabilities.

[0151] Step 3331e, determine the state path in the state network that best matches the speech data according to the state probabilities of the speech frame belonging to different states.

[0152] As Fig.16 shown, 601 represents speech data containing several speech frames, existing in the form of a waveform; 602-603 represent the feature vectors of the speech frames in the speech data. Among them, 602 Fig.16 uses the depth of the color blocks to represent the magnitude of the vector value, and 603 represents each speech frame through each rectangular block, and O i (0 < i < 30) represents the feature vector of the i-th speech frame; 604 represents the state indicated by the path node in the state network.

[0153] Continuing to refer to Fig.16 , for example, based on the observation probabilities of each speech frame corresponding to different states, the 1st to 4th speech frames correspond to state S1, the 5th to 7th speech frames correspond to state S2, and so on. And for state S1, the transition probability from state 1 to state 2 is 0.4, and the transition probability from state 1 to its own state 1 is 0.6, and so on. Thus, the probabilities of the speech frame belonging to different states can be calculated.

[0154] Combined with Fig.17 shown, the state probabilities P(o|s i (i = 1, 2..., 5) of the speech frame belonging to state S i are 0.45, 0.25, 0.75, 0.45, 0.25 respectively. Among them, P(o|s 3 ) = 0.75 is the largest, indicating that the speech frame belongs to state S3. Furthermore, it is determined that the state path in the state network that best matches the speech data includes the path node for storing state S3.

[0155] Based on the above, the state path that best matches the speech data is searched in the state network, that is, according to the observation probability and the transition probability, a state path that maximizes the state probability of the state to which the speech frame belongs is searched in the state network as the state path that best matches the speech data.

[0156] Under the effect of the above-mentioned embodiments, path search in the state network is realized by state probability calculation, which can effectively improve the accuracy of path search, thereby facilitating the improvement of the accuracy of speech recognition.

[0157] See also Fig.18 , a possible implementation method is provided in the embodiment of the present application, and step 337 may include the following steps:

[0158] Step 3371, based on the language model, respectively calculate the language probability that the phonemes in the phoneme set belong to different words.

[0159] Step 3373, the phonemes in the phoneme set are combined into words according to the calculated language probabilities as the speech recognition result.

[0160] In this embodiment, prediction is to calculate the language probability that the phonemes in the phoneme set belong to different words.

[0161] Continuing with the above example, assuming that the language probability of phoneme t in the phoneme set belonging to the word "soup" is 0.75, the language probability of belonging to the word "he" is 0.5, and the language probability of belonging to the word "special" is 0.1, then, according to the maximum value of the language probability, it is determined that phoneme t belongs to the word "soup".

[0162] Similarly, assuming that the language probability of the phoneme ang in the phoneme set belonging to the word "tang" is 0.6, the language probability of belonging to the word "tang" is 0.5, and the language probability of belonging to the word "ang" is 0.4, then according to the maximum value of the language probability, it is determined that the phoneme ang belongs to the word "tang".

[0163] Thus, the phoneme t and the phoneme ang in the phoneme set form the word "soup" as a speech recognition result.

[0164] Through the above process, word recognition based on the language model is realized, so that the speech recognition results conform to the probability distribution of the language itself, further improving the accuracy of speech recognition, which is conducive to ensuring the correctness of the automatic generation of sensitive words.

[0165] See also Fig.19 , a possible implementation method is provided in an embodiment of the present application, and the method may further include the following steps:

[0166] Step 61, in response to the sensitive word synchronization request, obtain the sensitive words to be synchronized uploaded for the target video.

[0167] For example, see Figure 6 A clickable shielding synchronization option 404, namely "shielding synchronization", is provided in the shielding control interface 402. When the user clicks the shielding synchronization option 404, the terminal initiates a sensitive word synchronization request to the server in response to the detected click operation.

[0168] For the server, after receiving the sensitive word synchronization request, the sensitive words to be synchronized for the target video are added to the sensitive words of the target video.

[0169] The sensitive words to be synchronized are pre-stored on the server side and uploaded to the server side by the terminals of other users. For example, the staff of the video publishing platform can manually add sensitive words with the help of other terminals and push them to the server side as the sensitive words to be synchronized.

[0170] Step 63, adding the sensitive words to be synchronized to the sensitive words of the target video.

[0171] In other words, the sensitive words of the target video can not only be automatically generated based on the automatic speech recognition results, but can also be manually added by the user with the help of the terminal.

[0172] Through the cooperation of the above-mentioned embodiments, the synchronous collection of sensitive words based on massive terminals is realized, which serves as a supplement to the automatic speech recognition results to fully ensure the correctness of sensitive words, making the filtering of comment data more accurate, and thus more effectively preventing users from being spoiled in advance, thereby bringing users a better video viewing experience.

[0173] The following is an embodiment of the device of the present application, which can be used to execute the comment display method involved in the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the comment display method involved in the present application.

[0174] See also Fig. 20 In an embodiment of the present application, a comment display device 900 is provided, including but not limited to: a data acquisition module 901, a speech recognition module 903, a data comparison module 905 and a data sending module 906.

[0175] The data acquisition module 901 is used to acquire voice data related to the target video.

[0176] The speech recognition module 903 is used to perform speech recognition on speech data related to the target video based on the constructed speech decoding network to obtain a speech recognition result.

[0177] The data comparison module 905 is used to compare the video content data of the target video with the speech recognition result, and generate sensitive words of the target video according to the comparison result.

[0178] The data sending module 907 is used to send the comment data filtered according to the sensitive words of the target video to the terminal when receiving the blocking request sent by the terminal, so that the terminal displays the filtered comment data during the process of playing the target video.

[0179] In an embodiment of the present application, a comment display device 900 is provided, wherein a speech recognition module 903 includes but is not limited to: a feature extraction unit, a state recognition unit, a phoneme recognition unit, and a word recognition unit.

[0180] The feature extraction unit is used to extract acoustic features from the speech frames contained in the speech data to obtain feature vectors of the speech frames.

[0181] The state recognition unit is used to recognize the states divided by phonemes for the feature vector of the speech frame based on the acoustic model, determine the state to which the speech frame belongs, and obtain the state sequence corresponding to the speech data according to the state to which the speech frame belongs.

[0182] The phoneme recognition unit is used to generate a phoneme set corresponding to the voice data from a state sequence corresponding to the voice data according to the states divided by the phonemes.

[0183] The word recognition unit inputs the phoneme set corresponding to the speech data into the language model, predicts the word to which the phoneme belongs, and obtains the speech recognition result.

[0184] Among them, the speech decoding network is constructed based on the acoustic model and language model.

[0185] In an embodiment of the present application, a comment display device 900 is provided, wherein a feature extraction unit includes but is not limited to: a frame processing subunit, a feature extraction subunit, and a feature calculation subunit.

[0186] The frame processing subunit is used to perform frame processing on the voice data to obtain a number of voice frames contained in the voice data.

[0187] The feature extraction subunit is used to extract the Mel-frequency cepstral coefficient features of each speech frame contained in the speech data.

[0188] The feature calculation subunit is used to calculate the feature vector of the speech frame according to the extracted Mel-frequency cepstral coefficient features.

[0189] The embodiment of the present application provides a comment display device 900, wherein the frame processing subunit includes but is not limited to: an endpoint detection subunit.

[0190] Among them, the endpoint detection subunit is used to perform voice endpoint detection on voice data.

[0191] Correspondingly, performing frame processing on the voice data includes: performing frame processing on the voice data after voice endpoint detection.

[0192] In an embodiment of the present application, a comment display device 900 is provided, wherein a state identification unit includes but is not limited to: a path search subunit and a state generation subunit.

[0193] The path search subunit is used to search for the state path that best matches the speech data in the state network constructed based on the hidden Markov model according to the feature vector of the speech frame.

[0194] The state generation subunit is used to generate a state sequence corresponding to the voice data according to the state to which the voice frame belongs indicated by each path node in the searched state path.

[0195] In an embodiment of the present application, a comment display device 900 is provided, wherein the path search subunit includes but is not limited to: a first probability calculation subunit, a second probability calculation subunit and a path determination subunit.

[0196] Among them, the first probability calculation subunit is used to calculate the observation probability of the speech frame corresponding to different states based on the states indicated by each path node in the state network and according to the feature vector of the speech frame, and calculate the transition probability of the speech frame transferring between the corresponding different states.

[0197] The second probability calculation subunit is used to calculate the state probability of the speech frame belonging to different states according to the observation probability and the transition probability.

[0198] The path determination subunit determines the state path that best matches the speech data in the state network according to the state probabilities that the speech frames belong to different states.

[0199] In an embodiment of the present application, a comment display device 900 is provided, wherein a word recognition unit includes but is not limited to: a language probability calculation subunit and a word generation subunit.

[0200] The language probability calculation subunit is used to calculate the language probability that the phonemes in the phoneme set belong to different words based on the language model.

[0201] The word generation subunit is used to combine the phonemes in the phoneme set into words according to the calculated language probabilities as the speech recognition result.

[0202] In an embodiment of the present application, a comment display device 900 is provided, wherein a data comparison module 905 includes but is not limited to: a word search unit and a sensitive word definition unit.

[0203] The word search unit is used to search for words matching the reference words in the speech recognition results according to the reference words in the video content data.

[0204] The sensitive word definition unit is used to use the searched words as sensitive words of the target video.

[0205] The embodiment of the present application provides a comment display device 900, which also includes but is not limited to: a word acquisition module and a word synchronization module.

[0206] Among them, the word acquisition module is used to respond to the sensitive word synchronization request and obtain the sensitive words to be synchronized for the target video upload.

[0207] The word synchronization module is used to add the sensitive words to be synchronized to the sensitive words of the target video.

[0208] The embodiment of the present application provides a comment display device 900, wherein the data sending module 907 includes but is not limited to: a request receiving unit, a data acquiring unit and a data filtering unit.

[0209] Among them, the request receiving unit is used to receive a sensitive word shielding request, and determine the target video currently played by the terminal according to the sensitive word shielding request.

[0210] The data acquisition unit is used to acquire comment data related to the target video after the current playback time point.

[0211] The data filtering unit is used to filter the obtained comment data according to the sensitive words of the target video.

[0212] The data sending unit is used to send the filtered comment data to the terminal.

[0213] An embodiment of the present application provides a comment display device 900, wherein the data filtering unit includes but is not limited to: a word matching subunit and a word filtering subunit.

[0214] Among them, the word matching subunit is used to match the sensitive words of the target video with the words in the comment data to obtain the sensitive word matching degree of the comment data.

[0215] The word filtering subunit is used to filter the comment data if the matching degree of the sensitive words in the comment data exceeds the matching degree threshold.

[0216] It should be noted that the comment display device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when displaying comments. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the comment display device will be divided into different functional modules to complete all or part of the functions described above.

[0217] In addition, the comment display device and the comment display method provided in the above embodiments belong to the same concept, wherein the specific manner in which each module performs operations has been described in detail in the method embodiments and will not be repeated here.

[0218] In the above process, the automatic generation of sensitive words is realized by using automatic speech recognition technology, which effectively solves the problem that the sensitive words used for comment data filtering in the prior art rely on manual implementation.

[0219] Fig.21 A schematic diagram of the structure of a server is shown according to an exemplary embodiment. The server is suitable for Figure 1 The server side 200 of the implementation environment is shown.

[0220] It should be noted that the server is only an example adapted to the present application and cannot be considered to provide any limitation on the scope of use of the present application. The server cannot be interpreted as requiring dependence on or having Fig.21 One or more components of exemplary server 2000 are shown.

[0221] The hardware structure of the server 2000 may vary greatly due to different configurations or performances, such as Fig.21 As shown, the server 2000 includes: a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .

[0222] Specifically, the power supply 210 is used to provide operating voltage for each hardware device on the server 200 .

[0223] The interface 230 includes at least one wired or wireless network interface for interacting with external devices. Figure 1 The interaction between the terminal 100 and the server 200 in the implementation environment is shown.

[0224] Of course, in other examples adapted by the present application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input-output interface 235, and at least one USB interface 237, etc. Fig.21 As shown, this is not intended to be a specific limitation.

[0225] The memory 250 is a carrier for storing resources, which may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 251, an application 253 and data 255, etc. The storage method may be temporary storage or permanent storage.

[0226] Among them, the operating system 251 is used to manage and control the hardware devices and application programs 253 on the server 200 to enable the central processor 270 to calculate and process the massive data 255 in the memory 250. It can be WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0227] The application program 253 is a computer program that performs at least one specific task based on the operating system 251, and may include at least one module ( Fig.21 (not shown), each module may respectively include a series of computer-readable instructions for the server 200. For example, the comment display device may be regarded as an application 253 deployed on the server 200.

[0228] The data 255 may be photos, pictures, etc. stored in a disk, or may be comment data, etc., stored in the memory 250 .

[0229] The central processor 270 may include one or more processors and is configured to communicate with the memory 250 through at least one communication bus to read the computer-readable instructions stored in the memory 250, thereby realizing the operation and processing of the mass data 255 in the memory 250. For example, the comment display method is completed by the central processor 270 reading a series of computer-readable instructions stored in the memory 250.

[0230] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.

[0231] See also Fig. 22 In an embodiment of the present application, an electronic device 400 is provided, including at least one processor 4001, at least one communication bus 4002, and at least one memory 4003. For example, the electronic device 400 can be a desktop computer, a laptop computer, a server, and the like.

[0232] The processor 4001 and the memory 4003 are connected, such as through a communication bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0233] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0234] The communication bus 4002 may include a path to transmit information between the above components. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 22 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0235] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.

[0236] The memory 4003 stores computer-readable instructions, and the processor 4001 reads the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .

[0237] When the computer-readable instructions are executed by the processor 4001, the comment display method in the above-mentioned embodiments is implemented.

[0238] A storage medium is provided in an embodiment of the present application. A computer program is stored on the storage medium. When the computer program is executed by a processor, the comment display method in the above embodiments is implemented.

[0239] A computer program product is provided in an embodiment of the present application, the computer program product includes computer-readable instructions, the computer-readable instructions are stored in a storage medium. A processor of a computer device reads the computer-readable instructions from the storage medium, and the processor executes the computer-readable instructions, so that the computer device executes the comment display method in each of the above embodiments.

[0240] Compared with the existing technology, the automatic generation of sensitive words is realized by using automatic speech recognition technology. On the one hand, the formulation of sensitive words is no longer affected by the subjective factors of the staff of the video publishing platform, which improves the accuracy of sensitive words. On the other hand, the addition of sensitive words no longer depends on the manual completion of the staff, avoiding the problem that the manual addition process of sensitive words is too cumbersome, and effectively solves the problem that the sensitive words used for comment data filtering in the existing technology rely on manual implementation.

[0241] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0242] The above descriptions are only some embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for displaying comments, It is characterized in that Applied to a server, the method comprises: Acquire voice data related to the target video; divide the target video into multiple sub-videos, and obtain voice data related to each sub-video; Based on the constructed speech decoding network, speech recognition is performed on the speech data related to each sub-video to obtain the speech recognition result of each sub-video; Acquire video content data of a target video; wherein the video content data of the target video includes reference words for describing the core content of the target video; and search for words matching the reference words in the speech recognition results of each sub-video according to the reference words for describing the core content of the target video; For each sub-video, using the words searched from the speech recognition result of the sub-video and the words searched from the sub-video after the sub-video as the sensitive words of the sub-video; Upon receiving a sensitive word blocking request for the target video initiated by the terminal, determine the target sub-video currently being played by the terminal; filter the acquired comment data of the target sub-video according to the sensitive words of the target sub-video; and send the filtered comment data of the target sub-video to the terminal, so that the terminal displays the filtered comment data of the target sub-video during the playing of the target sub-video.

2. The method according to claim 1, It is characterized in that The speech decoding network is constructed according to the acoustic model and the language model; The speech decoding network constructed based on the speech recognition is performed on the speech data related to each sub-video to obtain the speech recognition result of each sub-video, including: For the speech frames included in the speech data, acoustic feature extraction is performed to obtain feature vectors of the speech frames; Based on the acoustic model, identifying the states divided by phonemes for the feature vector of the speech frame, determining the state to which the speech frame belongs, and obtaining a state sequence corresponding to the speech data according to the state to which the speech frame belongs; According to the states of the phoneme division, a phoneme set corresponding to the voice data is generated from a state sequence corresponding to the voice data; The phoneme set corresponding to the speech data is input into the language model, and the word to which the phoneme belongs is predicted to obtain the speech recognition result.

3. The method according to claim 2, It is characterized in that The step of extracting acoustic features from the speech frames contained in the speech data to obtain feature vectors of the speech frames includes: Performing frame processing on the voice data to obtain a plurality of voice frames contained in the voice data; For each of the speech frames included in the speech data, extracting Mel-frequency cepstral coefficient features of the speech frame; The feature vector of the speech frame is calculated according to the extracted Mel-frequency cepstral coefficient features.

4. The method according to claim 3, It is characterized in that Before the voice data is subjected to frame processing, the method further comprises: Performing voice endpoint detection on the voice data; The step of performing frame processing on the voice data comprises: The voice data after voice endpoint detection is framed.

5. The method according to claim 2, It is characterized in that The acoustic model is a hidden Markov model; The step of identifying the states of the feature vectors of the speech frames divided by phonemes based on the acoustic model, determining the states to which the speech frames belong, and obtaining the state sequence corresponding to the speech data according to the states to which the speech frames belong, comprises: According to the feature vector of the speech frame, searching for a state path that best matches the speech data in a state network constructed based on the hidden Markov model; A state sequence corresponding to the voice data is generated according to the state to which the voice frame belongs indicated by each path node in the searched state path.

6. The method according to claim 5, It is characterized in that The step of searching, according to the feature vector of the speech frame, a state path that best matches the speech data in a state network constructed based on the hidden Markov model comprises: Based on the states indicated by the path nodes in the state network, and according to the feature vector of the speech frame, calculating the observation probability of the speech frame corresponding to different states; and calculating the transition probability of the speech frame transitioning between the corresponding different states; Calculating the state probability of the speech frame belonging to different states according to the observation probability and the transition probability; According to the state probabilities that the speech frame belongs to different states, a state path in the state network that best matches the speech data is determined.

7. The method according to claim 2, It is characterized in that The step of inputting the phoneme set corresponding to the speech data into the language model, predicting the word to which the phoneme belongs, and obtaining the speech recognition result includes: Based on the language model, respectively calculating the language probabilities that the phonemes in the phoneme set belong to different words; The phonemes in the phoneme set are combined into words according to the calculated language probabilities as the speech recognition result.

8. The method according to claim 1, It is characterized in that The method further comprises: In response to the sensitive word synchronization request, obtaining the sensitive words to be synchronized uploaded for the target video; The sensitive words to be synchronized are added to the sensitive words of the target video.

9. The method according to claim 1, It is characterized in that The filtering the acquired comment data of the target sub-video according to the sensitive words of the target sub-video includes: Matching the sensitive words of the target sub-video with the words in the comment data of the target sub-video to obtain the sensitive word matching degree of the comment data of the target sub-video; If the sensitive word matching degree of the comment data of the target sub-video exceeds the matching degree threshold, the comment data of the target sub-video is filtered.

10. A comment display device, It is characterized in that Applied to a server, the device comprises: The data acquisition module is used to acquire the voice data related to the target video; divide the target video into multiple sub-videos, and acquire the voice data related to each sub-video; A speech recognition module is used to perform speech recognition on the speech data related to each sub-video based on the constructed speech decoding network to obtain a speech recognition result for each sub-video; A data comparison module is used to obtain video content data of a target video; wherein the video content data of the target video includes reference words for describing the core content of the target video; based on the reference words for describing the core content of the target video, a word matching the reference words is searched in the speech recognition results of each sub-video; for each sub-video, a matching word searched from the speech recognition results of the sub-video and a word searched from a sub-video after the sub-video are used as sensitive words of the sub-video; A data sending module is used to determine the target sub-video currently played by the terminal when receiving a sensitive word blocking request for the target video sent by the terminal; filter the acquired comment data of the target sub-video according to the sensitive words of the target sub-video; and send the filtered comment data of the target sub-video to the terminal, so that the terminal displays the filtered comment data of the target sub-video during the process of playing the target sub-video.

11. The device according to claim 10, It is characterized in that The speech decoding network is constructed according to the acoustic model and the language model; The speech recognition module comprises: A feature extraction unit, configured to extract acoustic features from a speech frame included in the speech data to obtain a feature vector of the speech frame; A state recognition unit, configured to recognize states divided by phonemes on the feature vector of the speech frame based on the acoustic model, and obtain a state sequence corresponding to the speech frame; A phoneme recognition unit, configured to generate a phoneme set corresponding to the speech frame from a state sequence corresponding to the speech frame according to the states divided by the phonemes; The word recognition unit inputs the phoneme set corresponding to the speech frame into the language model, predicts the word to which the phoneme belongs, and obtains the speech recognition result.

12. The device according to claim 11, It is characterized in that The feature extraction unit includes a frame processing subunit, a feature extraction subunit and a feature calculation subunit; The frame processing subunit is used to perform frame processing on the voice data to obtain a plurality of voice frames contained in the voice data; A feature extraction subunit, configured to extract Mel-frequency cepstral coefficient features of each speech frame contained in the speech data; The feature calculation subunit is used to calculate the feature vector of the speech frame according to the extracted Mel-frequency cepstral coefficient features.

13. The device according to claim 12, It is characterized in that The framing processing subunit includes: an endpoint detection subunit; Among them, the endpoint detection subunit is used to perform voice endpoint detection on the voice data; the frame processing of the voice data includes: performing frame processing on the voice data after the voice endpoint detection.

14. The device according to claim 11, It is characterized in that The acoustic model is a hidden Markov model; The state identification unit includes: a path search subunit and a state generation subunit; The path search subunit is used to search for a state path that best matches the speech data in a state network constructed based on the hidden Markov model according to the feature vector of the speech frame; The state generation subunit is used to generate a state sequence corresponding to the voice data according to the state to which the voice frame belongs indicated by each path node in the searched state path.

15. The device according to claim 14, It is characterized in that The path search subunit includes: a first probability calculation subunit, a second probability calculation subunit and a path determination subunit; The first probability calculation subunit is used to calculate the observation probability of the speech frame corresponding to different states based on the states indicated by each path node in the state network and according to the feature vector of the speech frame; and calculate the transition probability of the speech frame transitioning between the corresponding different states; A second probability calculation subunit, used for calculating the state probability that the speech frame belongs to different states according to the observation probability and the transition probability; The path determination subunit determines the state path in the state network that best matches the voice data according to the state probabilities that the voice frame belongs to different states.

16. The device according to claim 11, It is characterized in that The word recognition unit includes: a language probability calculation subunit and a word generation subunit; The language probability calculation subunit is used to calculate the language probability of the phonemes in the phoneme set belonging to different words based on the language model; The word generation subunit is used to combine the phonemes in the phoneme set into words according to the calculated language probabilities as the speech recognition result.

17. The device according to claim 10, It is characterized in that The device also includes: The sub-acquisition module is used to respond to the sensitive word synchronization request and acquire the sensitive words to be synchronized uploaded for the target video; A word synchronization module is used to add the sensitive words to be synchronized to the sensitive words of the target video.

18. The device according to claim 10, It is characterized in that The data sending module includes a data filtering unit, and the data filtering unit includes: a word matching subunit and a word filtering subunit; The word matching subunit is used to match the sensitive words of the target sub-video with the words in the comment data to obtain the sensitive word matching degree of the comment data of the target sub-video; The word filtering subunit is used to filter the comment data of the target sub-video if the sensitive word matching degree of the comment data of the target sub-video exceeds a matching degree threshold.

19. An electronic device, It is characterized in that include: at least one processor, at least one memory, and at least one communication bus, wherein: The memory stores computer-readable instructions, and the processor reads the computer-readable instructions in the memory through the communication bus; When the computer-readable instructions are executed by the processor, the comment display method according to any one of claims 1 to 9 is implemented.

20. A storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the comment display method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Information processing method and apparatus, terminal device and storage medium

    CN107613392A