Bullet screen detection method and device, electronic equipment and computer readable medium
By introducing a large language model for bullet screen detection and utilizing video clip parsing and a sensitive word database, the problem of low accuracy in bullet screen detection in existing technologies is solved, achieving more efficient identification of illegal bullet screens.
Patent Information
- Application Number
- CN202510964234.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-28
AI Technical Summary
In existing technologies, the accuracy of bullet screen detection is low, and it is easy to have false positives and false negatives. In particular, it cannot effectively handle bullet screens with complex expressions, metaphors, homophones, special characters, and confusing word order.
A large language model is introduced for bullet screen detection. By acquiring video segments associated with the bullet screen to be tested, parsing the video content information, generating a summary, and combining it with a pre-set sensitive word library for violation detection, the natural language processing and semantic understanding capabilities of the large language model are used to deeply understand the intent of the bullet screen and video content.
It reduces the false positive and false negative rates of bullet screen detection, improves the accuracy of bullet screen detection, and can more accurately identify illegal bullet screens.
Smart Images

Figure CN120856945A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to methods, apparatus, electronic devices, and computer-readable media for detecting bullet comments. Background Technology
[0002] "Bullet comments" refer to the commentary text that pops up when watching videos online. Users can send bullet comments in real time while watching videos, interacting with other viewers and sharing their opinions and feelings. Among the massive number of bullet comments, there are usually some containing inappropriate content, so bullet comments need to be detected.
[0003] In existing technologies, bullet screen detection is typically performed using preset rules, such as keyword matching. However, due to the rich variety of bullet screen content and diverse language styles, this method cannot handle bullet screens with complex expressions, metaphors, homophones, special characters, or confusing word order, which easily leads to false positives and false negatives, resulting in low accuracy in bullet screen detection. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and computer-readable medium for detecting bullet comments, which reduces the false detection rate and false negative rate of bullet comment detection and improves the accuracy of bullet comment detection.
[0005] In a first aspect, embodiments of this application provide a method for detecting bullet comments, the method comprising: acquiring a video segment associated with a bullet comment to be tested; parsing the video segment to obtain video content information; extracting a summary of the video content information using a large language model, wherein the large language model is trained based on a sample set containing bullet comment samples, and the bullet comment samples are labeled with indications of whether they violate regulations; and generating a bullet comment detection result indicating whether the bullet comment to be tested violates regulations based on the bullet comment to be tested, the summary, and a preset sensitive word library, using the large language model.
[0006] Secondly, embodiments of this application provide a bullet screen detection device, which includes: an acquisition unit for acquiring a video segment associated with a bullet screen to be tested; a parsing unit for parsing the video segment to obtain video content information; a determination unit for extracting a summary of the video content information using a large language model, wherein the large language model is trained based on a sample set containing bullet screen samples, and the bullet screen samples have annotations indicating whether they violate regulations; and a detection unit for generating a bullet screen detection result indicating whether the bullet screen to be tested violates regulations based on the bullet screen to be tested, the summary, and a preset sensitive word library, using the large language model.
[0007] Thirdly, embodiments of this application provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any embodiment of the first aspect.
[0008] Fourthly, embodiments of this application provide a computer-readable medium having a computer program stored thereon that, when executed by a processor, implements the method as described in any embodiment of the first aspect.
[0009] The bullet screen detection method, apparatus, electronic device, and computer-readable medium provided in this application have the following advantages. Firstly, during bullet screen detection, a large language model pre-trained based on bullet screen samples is introduced. Because this model is trained specifically for the bullet screen moderation scenario and possesses excellent natural language processing and semantic understanding capabilities, it can deeply understand the intent and semantics of the bullet screen and video content being tested, reducing the false positive and false negative rates and improving the accuracy of bullet screen detection. Secondly, during bullet screen detection, not only is the bullet screen itself considered, but also the video content information of the associated video segments is incorporated. Since the video content information reflects the video content context and bullet screen sending scenario, it can combine the bullet screen content with its specific context to perform bullet screen violation detection. This allows the large language model to accurately understand the overall intent of the bullet screen, further reducing the false positive and false negative rates and thus further improving the accuracy of bullet screen detection. Attached Figure Description
[0010] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0011] Figure 1 This is a flowchart illustrating an embodiment of the bullet screen detection method of this application;
[0012] Figure 2 This is a schematic diagram of the processing procedure of the bullet screen detection method of this application;
[0013] Figure 3 This is a schematic diagram of the structure of an embodiment of the bullet screen detection device of this application;
[0014] Figure 4 This is a schematic diagram of the structure of an electronic device used to implement the embodiments of this application. Detailed Implementation
[0015] All actions involving the acquisition of signals, information, or data in this application are carried out in accordance with the relevant data protection laws and policies of the country where the application is located, and with the authorization of the owner of the relevant device.
[0016] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] Please refer to Figure 1 This document illustrates a flowchart 100 of an embodiment of the bullet comment detection method according to this application. This bullet comment detection method can be applied to various electronic devices with data processing capabilities. For example, the aforementioned electronic devices may include, but are not limited to, physical servers, cloud servers, server clusters, etc. The executing entity of the bullet comment detection method may be a processor in the aforementioned electronic device. The application scenarios of this bullet comment detection method can be scenarios where users send bullet comments, specifically including but not limited to short video playback scenarios, long video playback scenarios, game live streaming scenarios, etc.
[0019] This bullet screen detection method includes the following steps:
[0020] Step 101: Obtain the video segment associated with the bullet comments to be tested.
[0021] In this embodiment, the bullet comment to be tested can be any bullet comment sent by any user while watching any video. When any user sends a bullet comment while watching any video, that bullet comment can be used as the bullet comment to be tested and associated with that video. Then, a video segment of a set duration before and after the time the bullet comment to be tested was sent is used as the video segment associated with the bullet comment to be tested. For example, the video segment associated with the bullet comment to be tested can be a video segment of 30 seconds before and after the time the bullet comment to be tested was sent in the aforementioned video.
[0022] Step 102: Analyze the video clips to obtain video content information.
[0023] In this embodiment, the video content information can be information used to characterize the video content. In practice, video segments can be parsed using various methods. Accordingly, depending on the parsing method, the video content information can be represented in different forms.
[0024] As an example, the subtitle segments corresponding to the above video clips can be extracted from the subtitle file corresponding to the above video as video content information.
[0025] As another example, the audio segments corresponding to the video clips can be extracted from the audio files corresponding to the video above, and used as video content information.
[0026] As another example, text information can be extracted from the video frames contained in the above video clip as video content information.
[0027] As another example, keyframes or video segments containing key content can be extracted from the aforementioned video clips as video content information. Keyframes are image frames in the video that contain crucial information or significant changes, reflecting the main content of a shot.
[0028] It should be noted that the method of parsing video clips and the format of video content information can be set as needed, and no specific limitations are made here. Furthermore, the format of video content information can be one or more, and is not limited to those listed above.
[0029] In some optional implementations, the video content information can be multimodal information. For example, the video content information includes at least one of the following: text information, audio information, images, and video sub-segments from the aforementioned video clip. Specifically, the text information can be dialogue, subtitles, or other information from the aforementioned video clip; the audio information can be the audio file corresponding to the aforementioned video clip; the images can be keyframes from the aforementioned video clip; and the video sub-segments can be sub-segments from the aforementioned video clip that contain key content. In this case, the large language model in the following steps can be a multimodal large language model that supports multimodal content input.
[0030] Step 103: Extract a summary of the video content information using a large language model.
[0031] In this embodiment, the Large Language Model (LLM) is a deep learning model trained on massive amounts of text data. It can not only generate natural language text but also deeply understand the meaning of text and handle various natural language tasks, such as text summarization, question answering, and translation. In practical applications, the Large Language Model used in this embodiment can be either an open-source model or a self-developed model; no specific limitation is made here.
[0032] In this embodiment, the large language model can be trained based on a sample set containing bullet screen samples, which are labeled with indications of whether they violate regulations. These bullet screen samples are collected in advance from the internet by technical personnel. In practice, the above sample set can be used to further train a general large language model using machine learning methods to adjust its parameters, resulting in a large language model for bullet screen detection. This general large language model is the one trained using a common dataset. It is understood that the sample set used to train the large language model in this embodiment may include other text samples besides bullet screen samples, enabling the large language model to detect other types of text.
[0033] In practice, during the training process of the large language model, bullet screen samples can be input one by one to obtain the violation detection results output by the large language model. Then, based on the violation detection results and the corresponding annotation information of the input bullet screen samples, the loss value of the large language model can be determined. This loss value is the value of the loss function, a non-negative real-valued function used to characterize the difference between the detection result and the true result. Generally, the smaller the loss value, the better the robustness of the model. The loss function can be set according to actual needs. Afterwards, the parameters of the large language model can be updated using this loss value. Thus, each time a bullet screen sample is input, the parameters of the large language model can be updated based on the loss value corresponding to that bullet screen sample until training is complete.
[0034] In this embodiment, the aforementioned execution entity can extract a summary of video content information using the large language model. Specifically, a prompt can first be generated based on the video content information. Then, this prompt can be input into the large language model to obtain a summary of the video content information output by the model. The prompt here can instruct the large language model to output a summary of the video content information based on the video content information. For example, the prompt could be, "Assuming you are a professional comment moderator, the current video content information is: XXXX. Please extract the key content from the above video content information to form a summary." This summary is a concise overview of the video content information, indicating the core content within the video content information; it can be in text form.
[0035] Step 104: Based on the bullet comments to be tested, the summary, and the preset sensitive word library, generate bullet comment detection results to indicate whether the bullet comments to be tested violate regulations through a large language model.
[0036] In this embodiment, the sensitive word library may include a large number of sensitive words. Sensitive words are words that may cause controversy, offense, or violation of laws and regulations.
[0037] In this embodiment, the aforementioned execution entity can first generate another prompt word based on the bullet comment to be tested, the summary, and a preset sensitive word library. Then, this prompt word can be input into a large language model to obtain the bullet comment detection result output by the large language model. The prompt word here can be used to instruct the large language model to perform violation detection on the rewritten bullet comment based on the sensitive word library and the summary. Specifically, it can instruct the large language model to extract keywords from the bullet comment to be tested based on the content of the summary, and match these keywords with sensitive words in the sensitive word library to generate a bullet comment detection result based on the matching result. The bullet comment detection result can indicate whether the bullet comment to be tested violates regulations.
[0038] In practice, a sensitive word database can be provided as part of the input information to a large language model, enabling the model to detect violations in modified bullet comments based on the sensitive word database and the summary. Alternatively, the basic large language model can be fine-tuned in advance using a large amount of example data containing sensitive word databases and compliance judgment results, allowing the model to learn about the sensitive word database and thus detect violations in modified bullet comments based on the sensitive word database and the summary.
[0039] As an example, the prompt here could be: "Imagine you are a professional comment moderator. There is a comment with the content XX. The current video content summary is: XXX. Please carefully read the above content, extract the violating keywords from the above comment, match these keywords with sensitive words in the sensitive word database, and based on the matching results, conclude whether the comment being tested contains violating content."
[0040] In practice, large language models can use string matching to match keywords with sensitive words. If a keyword is the same as a sensitive word, they are considered a match. Alternatively, semantic similarity calculation can be used to match keywords with sensitive words. If the similarity between the semantic features of a keyword and the semantic features of a sensitive word is greater than a preset threshold, they are considered a match. It is understandable that when determining a match between a keyword and a sensitive word based on semantic similarity, the keyword and the sensitive word may be the same or different.
[0041] The method provided in the above embodiments of this application first obtains the video segment associated with the bullet comment to be tested; then, it parses the video segment to obtain video content information; subsequently, it extracts a summary of the video content information using a large language model, which is trained based on a sample set containing bullet comment samples, and the bullet comment samples are labeled with indications of whether they violate regulations; finally, based on the bullet comment to be tested, the summary, and a preset sensitive word library, the large language model generates a bullet comment detection result indicating whether the bullet comment to be tested violates regulations. On the one hand, in the bullet comment detection process, a large language model pre-trained based on bullet comment samples is introduced. Since this model is trained for the specific scenario of bullet comment review and has excellent natural language processing and semantic understanding capabilities, it can deeply understand the intent and semantics of the bullet comment to be tested and the video content, reducing the false detection rate and false negative rate of bullet comment detection and improving the accuracy of bullet comment detection. On the other hand, in the process of bullet screen detection, not only is the bullet screen itself being tested considered, but also the video content information of the video segments associated with the bullet screen is introduced. The video content information can reflect the video content context and bullet screen sending scenario when the bullet screen is sent. Therefore, it is possible to combine the bullet screen content and its specific scenario to detect bullet screen violations, so that the large language model can accurately understand the overall intent of the bullet screen, further reducing the false detection rate and false negative rate of bullet screen detection, thereby further improving the accuracy of bullet screen detection.
[0042] In some optional embodiments, step 101 above may further include the following sub-steps:
[0043] Sub-step 1011: Determine the sending time of the bullet screen to be tested.
[0044] Sub-step 1012: Determine the start and end times of the video segment to be extracted based on the sending time.
[0045] The start and end times include the start time and the end time. The start time can be a point in time before the sent time of the comment to be tested, and the end time can be a point in time after the sent time of the comment to be tested. For example, the start and end times can be the points in time corresponding to 30 seconds before the sent time of the comment to be tested. The end time can be the points in time corresponding to 30 seconds after the sent time of the comment to be tested.
[0046] Sub-step 1013: Based on the start and end times, extract video segments from the videos associated with the bullet comments to be tested. The extracted video segments are the video segments associated with the bullet comments to be tested, that is, the video segments from the aforementioned start time to the aforementioned end time.
[0047] The video clips obtained through the above process can reflect the video content context when the barrage under test was sent. This video content context reflects the video scene when the barrage under test was sent. Combining the video clips with the detection of the barrage under test helps to reduce the false detection rate and false negative rate of barrage detection, thereby improving the accuracy of barrage detection.
[0048] In some optional embodiments, step 102 above may further include the following sub-steps:
[0049] Sub-step 1021: Based on optical character recognition technology, extract the first text information from the video frames contained in the video segment.
[0050] Specifically, the aforementioned executing entity can first extract video frames from the aforementioned video segment, then use optical character recognition (OCR) technology to recognize the characters in the video frames, and then deduplicate the recognized characters to obtain the first text information. The first text information may include, but is not limited to, the dialogue in each video frame of the video segment.
[0051] Optical Character Recognition (OCR) is a technology that converts text content in images into editable and searchable text. It acquires images using optical scanning equipment, then uses computer algorithms to recognize characters within the image and convert them into a computer-processable text format. The principle is as follows: first, brightness detection is performed on the image, detecting the dark and light patterns in multiple regions to determine the character shapes. Then, character recognition methods (such as Euclidean space alignment, dynamic program alignment, neural network-based character alignment, etc.) are used to translate the character shapes into computer text.
[0052] Sub-step 1022: Based on automatic speech recognition technology, the speech corresponding to the video segment is converted into text to obtain the second text information.
[0053] Specifically, the aforementioned executing entity may first extract the audio segment corresponding to the aforementioned video segment, and then use automatic speech recognition technology to recognize the aforementioned audio segment, and use the text recognition result obtained after recognition as the second text information.
[0054] Automatic Speech Recognition (ASR) is a technology that converts human speech into text. Its core lies in using computer algorithms to identify and analyze the acoustic features in speech signals and convert them into recognizable text or instructions. The principle is as follows: First, the acquired audio undergoes preprocessing such as noise reduction and frame segmentation to improve audio quality. Then, audio features such as Mel-frequency cepstral coefficients are extracted from the preprocessed audio. Next, the extracted audio features are compared with a pre-trained acoustic model, mapping the audio features to individual speech sounds or phonemes. Then, a statistical language model is used to assemble the identified phonemes into words and phrases, predicting the most likely word sequences. Finally, the outputs of the acoustic and language models are combined to search for the best text recognition result.
[0055] Sub-step 1023: Generate video content information based on the first text information and the second text information.
[0056] Here, the first and second text information can be directly summarized to obtain the video content information. Alternatively, the first and second text information can be deduplicated to obtain the video content information.
[0057] By using OCR technology to identify text information in video frames and ASR technology to convert speech dialogue in video clips into text, the context of the video content when the bullet comments are sent can be obtained. This video content context reflects the video scene when the bullet comments to be tested are sent. Combining this video clip with the detection of bullet comments helps to reduce the false detection rate and false negative rate of bullet comment detection, thereby improving the accuracy of bullet comment detection.
[0058] In some optional implementations of this embodiment, sub-step 1023 may further include the following steps: obtaining video cataloging information; generating video content information based on the first text information, the second text information, and the video cataloging information.
[0059] Video cataloging information describes the content and attributes of a video, aiming to efficiently organize, retrieve, manage, and utilize video resources. Video cataloging information may include, but is not limited to, video title, actor names, video description, and video tags. Here, the first text information, the second text information, and the video cataloging information can be combined to generate video content information.
[0060] By introducing video cataloging information, the video content information can be further supplemented and improved, making the video content information more accurate and richer, thereby further improving the accuracy of bullet screen detection.
[0061] In conjunction with the above optional embodiments, see further... Figure 2 , which shows a schematic diagram of the bullet screen detection process. First, a video segment associated with the bullet screen to be detected can be obtained. For example, a video segment of 30 seconds before and after the time point when the bullet screen to be detected is sent can be intercepted. Then, OCR recognition and ASR recognition can be performed on this video segment. After that, the obtained video content information can be input into a large language model to obtain a summary of the video content information. After that, this summary and the bullet screen to be detected can be input into the large language model to obtain the keywords of the bullet screen to be detected. After that, the large language model can match the keywords with a sensitive word library. If a keyword matches any sensitive word in the sensitive word library, it can be determined that the bullet screen to be detected is违规 (i.e., the bullet screen to be detected is a违规 bullet screen); on the contrary, if the keyword does not match any sensitive word in the sensitive word library, it can be determined that the bullet screen to be detected is合规 (i.e., the bullet screen to be detected is a normal bullet screen).
[0062] In some optional embodiments, the above step 104 may further include the following sub-steps:
[0063] Sub-step 1041, generating a first prompt word based on the bullet screen to be detected and inputting the first prompt word into the large language model to obtain a rewritten bullet screen. The first prompt word is used to instruct the large language model to rewrite the bullet screen to be detected.
[0064] In practical applications, a first prompt word can be constructed based on the bullet screen to be detected. Then, this prompt word can be input into the large language model to obtain the rewritten bullet screen output by the large language model. The first prompt word here can instruct the large language model to rewrite the bullet screen to be detected with metaphorical homophony, complex expressions, special characters, mixed word orders, etc. For example, the first prompt word here can be "You are a language error correction master. There is the following bullet screen to be detected: XXX. This bullet screen to be detected may have complex expressions, special characters, mixed word orders, etc. Please rewrite it into normal text."
[0065] Exemplarily, if the bullet screen to be detected is "duck不必", its rewritten bullet screen can be "大可不必". If the bullet screen to be detected is "XX*YY*ZZ", its rewritten bullet screen can be "XXYYZZ".
[0066] Sub-step 1042, generating a second prompt word based on the rewritten bullet screen, the summary, and a preset sensitive word library, and inputting the second prompt word into the large language model to obtain a bullet screen detection result. The second prompt word is used to instruct the large language model to perform a violation detection on the rewritten bullet screen based on the sensitive word library and the summary. The bullet screen detection result is used to indicate whether the bullet screen to be detected is in violation.
[0067] In practical applications, a second prompt word can first be generated based on the rewritten bullet comments, the summary, and a preset sensitive word library. Then, this second prompt word can be input into a large language model to obtain the bullet comment detection result output by the large language model. The second prompt word here can be used to instruct the large language model to perform violation detection on the rewritten bullet comments based on the sensitive word library and the summary. Specifically, the second prompt word can instruct the large language model to extract keywords from the rewritten bullet comments based on the content of the summary, and match these keywords with sensitive words in the sensitive word library to generate a bullet comment detection result based on the matching result. The bullet comment detection result indicates whether the bullet comment under test violates regulations. See the example in step 104; to avoid repetition, it will not be repeated here.
[0068] Before performing bullet screen detection using a large language model, the bullet screen to be tested is rewritten. This rewrites complex expressions, special characters, metaphors, homophones, and confusing word order in the bullet screen, thereby avoiding false detections and missed detections caused by the above problems, and further improving the accuracy of bullet screen detection.
[0069] In some optional embodiments, after step 104 is performed, the following steps may be further performed:
[0070] Step 105: Send the bullet comments to be tested, video clips, and bullet comment detection results to the terminal.
[0071] Step 106: Receive the review result of the bullet screen detection results returned by the terminal. The review result is used to indicate whether there is an error in the bullet screen detection results.
[0072] Here, the terminal can display the bullet comments to be tested, video clips, and bullet comment detection results. Technicians can then review the bullet comment detection results based on the video clips and bullet comments to determine if there are any errors, and then upload the review results.
[0073] Step 107: If the review result indicates that the bullet screen detection result is incorrect, the bullet screen to be tested is taken as a new bullet screen sample. Based on the review result, the annotation of the new bullet screen sample is generated, and the new bullet screen sample with annotation is stored in the sample set. The updated sample set is used to continue training the large language model.
[0074] Here, if the bullet screen detection result indicates that the bullet screen to be tested is in violation, and the review result indicates that the bullet screen detection result is incorrect, a label indicating that the bullet screen to be tested is compliant can be generated; if the bullet screen detection result indicates that the bullet screen to be tested is compliant, and the review result indicates that the bullet screen detection result is incorrect, a label indicating that the bullet screen to be tested is in violation can be generated.
[0075] By reviewing the results of bullet screen detection and constructing new bullet screen samples when the review results indicate that the detection results are incorrect, the sample set can be updated in real time. Therefore, as the content of bullet screens constantly changes, the large language model updated based on this sample set can quickly learn and adapt to new inappropriate content and expressions, resulting in more accurate detection capabilities.
[0076] In some optional embodiments, the bullet screen detection result generated in step 104 may include keywords from the bullet screen to be tested. After step 104, the following steps may also be performed: if the bullet screen detection result indicates that the bullet screen to be tested is compliant and the review result indicates that the bullet screen detection result is incorrect, then the keywords are stored as sensitive words in the sensitive word database.
[0077] Understandably, if the bullet comment detection result indicates that the bullet comment under test is compliant, while the review result indicates that the bullet comment detection result is incorrect, then the bullet comment under test is actually a violation. In this case, the keywords in the bullet comment under test can be considered as prohibited sensitive words, and storing them in a sensitive word database can enable dynamic updates to the sensitive word database, improving its richness and real-time performance. Based on this, the subsequent detection capabilities of the large language model can be further improved, further reducing false positives and false negatives.
[0078] Further reference Figure 3 As an implementation of the methods shown in the above figures, this application provides an embodiment of a bullet screen detection device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0079] like Figure 3 As shown, the bullet screen detection device 300 of this embodiment includes: an acquisition unit 301, used to acquire video segments associated with bullet screens to be tested; a parsing unit 302, used to parse the video segments to obtain video content information; a determination unit 303, used to extract a summary of the video content information through a large language model, wherein the large language model is trained based on a sample set containing bullet screen samples, and the bullet screen samples have annotations indicating whether they violate regulations; and a detection unit 304, used to generate a bullet screen detection result indicating whether the bullet screen to be tested violates regulations based on the bullet screen to be tested, the summary, and a preset sensitive word library, through the large language model.
[0080] In some optional implementations of this embodiment, the acquisition unit 301 is further configured to: determine the sending time of the bullet comment to be tested; determine the start and end times of the video segment to be extracted based on the sending time; and extract the video segment from the video associated with the bullet comment to be tested based on the start and end times.
[0081] In some optional implementations of this embodiment, the parsing unit 302 is further configured to: extract first text information from the video frames contained in the video segment based on optical character recognition technology; convert the speech corresponding to the video segment into text based on automatic speech recognition technology to obtain second text information; and generate video content information based on the first text information and the second text information.
[0082] In some optional implementations of this embodiment, the parsing unit 302 is further configured to: obtain video cataloging information, wherein the video cataloging information is information used to describe the content and attributes of the video; and generate video content information based on the first text information, the second text information and the video cataloging information.
[0083] In some optional implementations of this embodiment, the determining unit 303 is further configured to: generate a first prompt word based on the bullet screen to be tested, and input the first prompt word into the large language model to obtain a rewritten bullet screen, wherein the first prompt word is used to instruct the large language model to rewrite the bullet screen to be tested; generate a second prompt word based on the rewritten bullet screen, the summary, and a preset sensitive word library, and input the second prompt word into the large language model to obtain a bullet screen detection result, wherein the second prompt word is used to instruct the large language model to perform violation detection on the rewritten bullet screen based on the sensitive word library and the summary, and the bullet screen detection result is used to indicate whether the bullet screen to be tested violates regulations.
[0084] In some optional implementations of this embodiment, the device further includes a first storage unit, configured to: send the bullet comment to be tested, the video clip, and the bullet comment detection result to a terminal; receive an audit result returned by the terminal regarding the bullet comment detection result, the audit result indicating whether the bullet comment detection result is incorrect; if the audit result indicates that the bullet comment detection result is incorrect, then take the bullet comment to be tested as a new bullet comment sample, generate an annotation for the new bullet comment sample based on the audit result, store the new bullet comment sample with annotation in the sample set, and use the updated sample set to continue training the large language model.
[0085] In some optional implementations of this embodiment, the bullet screen detection result includes keywords in the bullet screen to be tested; the device further includes a second storage unit, used to: if the bullet screen detection result indicates that the bullet screen to be tested is compliant and the review result indicates that the bullet screen detection result is incorrect, then store the keywords as sensitive words in the sensitive word library.
[0086] In some optional implementations of this embodiment, the video content information includes at least one of the following: text information, voice information, images, and video sub-segments in the video segment.
[0087] The apparatus provided in the above embodiments of this application first acquires a video segment associated with the bullet comment to be tested; then, it parses the video segment to obtain video content information; subsequently, it extracts a summary of the video content information using a large language model, which is trained based on a sample set containing bullet comment samples, and the bullet comment samples are labeled with indications of whether they violate regulations; finally, based on the bullet comment to be tested, the summary, and a preset sensitive word library, the large language model generates a bullet comment detection result indicating whether the bullet comment to be tested violates regulations. On the one hand, in the bullet comment detection process, a large language model pre-trained based on bullet comment samples is introduced. Since this model is trained for the specific scenario of bullet comment review, and the model has excellent natural language processing and semantic understanding capabilities, it can deeply understand the intent and semantics of the bullet comment to be tested and the video content, reducing the false detection rate and false negative rate of bullet comment detection, and improving the accuracy of bullet comment detection. On the other hand, in the process of bullet screen detection, not only is the bullet screen itself being tested considered, but also the video content information of the video segments associated with the bullet screen is introduced. The video content information can reflect the video content context and bullet screen sending scenario when the bullet screen is sent. Therefore, it is possible to combine the bullet screen content and its specific scenario to detect bullet screen violations, so that the large language model can accurately understand the overall intent of the bullet screen, further reducing the false detection rate and false negative rate of bullet screen detection, thereby further improving the accuracy of bullet screen detection.
[0088] The following is for reference. Figure 4 It shows a schematic diagram of the structure of an electronic device used to implement some embodiments of this application. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this application.
[0089] like Figure 4 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0090] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, disks, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 4 Each box shown can represent a device or multiple devices as needed.
[0091] In particular, according to some embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined above in the methods of some embodiments of this application.
[0092] It should be noted that the computer-readable medium described in some embodiments of this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0093] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0094] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a video segment associated with the bullet comment to be tested; parse the video segment to obtain video content information; extract a summary of the video content information using a large language model, the large language model being trained based on a sample set containing bullet comment samples, the bullet comment samples having annotations indicating whether they violate regulations; and, based on the bullet comment to be tested, the summary, and a preset sensitive word library, generate a bullet comment detection result indicating whether the bullet comment to be tested violates regulations using the large language model.
[0095] Computer program code for performing operations of some embodiments of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++; and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, or it can be connected to an external computer (e.g., via the Internet using an Internet service provider), including local area networks (LANs) or wide area networks (WANs).
[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0097] The units described in some embodiments of this application can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a first determining unit, a second determining unit, a selecting unit, and a third determining unit. The names of these units do not necessarily limit the specific unit itself.
[0098] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0099] The above description is merely a selection of preferred embodiments of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this application.
Claims
1. A method for detecting bullet comments, characterized in that, The method includes: Obtain the video clip associated with the bullet comments to be tested; The video segment is parsed to obtain video content information; A summary of the video content information is extracted using a large language model, which is trained on a sample set containing bullet screen samples, and the bullet screen samples are labeled to indicate whether they violate regulations. Based on the bullet comments to be tested, the summary, and the preset sensitive word library, the large language model generates bullet comment detection results to indicate whether the bullet comments to be tested violate regulations.
2. The method according to claim 1, characterized in that, The acquisition of the video segment associated with the bullet comments to be tested includes: Determine the sending time of the bullet comments to be tested; Based on the sending time, determine the start and end times of the video segment to be extracted; Based on the start and end times, the video segment is extracted from the video associated with the bullet comments to be tested.
3. The method according to claim 1, characterized in that, The step of parsing the video segment to obtain video content information includes: Based on optical character recognition technology, first text information is extracted from the video frames contained in the video segment; Based on automatic speech recognition technology, the speech corresponding to the video segment is converted into text to obtain second text information; Based on the first text information and the second text information, video content information is generated.
4. The method according to claim 3, characterized in that, The step of generating video content information based on the first text information and the second text information includes: Obtain video cataloging information, which is information used to describe the content and attributes of the video; Based on the first text information, the second text information, and the video cataloging information, video content information is generated.
5. The method according to claim 1, characterized in that, The process of generating a bullet screen detection result based on the bullet screen to be tested, the summary, and a preset sensitive word library, using the large language model, to indicate whether the bullet screen to be tested violates regulations, includes: A first prompt word is generated based on the bullet screen to be tested, and the first prompt word is input into the large language model to obtain the rewritten bullet screen. The first prompt word is used to instruct the large language model to rewrite the bullet screen to be tested. A second prompt word is generated based on the rewritten bullet screen, the summary, and a preset sensitive word library. The second prompt word is then input into the large language model to obtain a bullet screen detection result. The second prompt word is used to instruct the large language model to perform violation detection on the rewritten bullet screen based on the sensitive word library and the summary. The bullet screen detection result is used to indicate whether the bullet screen to be tested violates regulations.
6. The method according to claim 1, characterized in that, After generating a bullet comment detection result using the large language model to indicate whether the bullet comment under test violates regulations, the method further includes: The bullet comment to be tested, the video clip, and the bullet comment detection results are sent to the terminal; The terminal returns an audit result for the bullet screen detection result, which indicates whether the bullet screen detection result is incorrect. If the review result indicates that the bullet screen detection result is incorrect, the bullet screen to be tested is taken as a new bullet screen sample, and an annotation is generated for the new bullet screen sample based on the review result. The new bullet screen sample with annotation is stored in the sample set, and the updated sample set is used to continue training the large language model.
7. The method according to claim 6, characterized in that, The bullet screen detection results include keywords from the bullet screen to be tested; After receiving the review result for the bullet screen detection result returned by the terminal, the method further includes: If the bullet screen detection result indicates that the bullet screen to be tested is compliant and the review result indicates that the bullet screen detection result is incorrect, then the keyword will be stored as a sensitive word in the sensitive word library.
8. The method according to claim 1, characterized in that, The video content information includes at least one of the following: text information, audio information, images, and video sub-segments within the video segment.
9. A bullet screen detection device, characterized in that, The device includes: The acquisition unit is used to acquire the video segment associated with the bullet comments to be tested. The parsing unit is used to parse the video segment to obtain video content information; The extraction unit is used to extract a summary of the video content information through a large language model, which is trained based on a sample set containing bullet screen samples, and the bullet screen samples are labeled with indications of whether they violate the rules. The detection unit is used to generate a bullet screen detection result indicating whether the bullet screen under test violates the rules, based on the bullet screen to be tested, the summary, and a preset sensitive word library, through the large language model.
10. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.
11. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.