Video graphic element display method and device
By extracting key information in video subtitles and performing object detection in video frames, and displaying matching video graphic elements using visual emphasis, the problem of limited video propagation effect in the prior art is solved, and more efficient video content attention and dissemination impact is achieved.
Patent Information
- Application Number
- CN202510230179.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-13
AI Technical Summary
When the prior art improves the video dissemination effect, it is impossible for viewers to pay efficient and precise attention to important video content, resulting in limited communication effect.
By extracting subtitle focus from the target video and performing object detection processing in the video frame associated with the subtitle focus, the video graphic elements matching the subtitle focus are displayed in a visual emphasis.
It has achieved the close integration and emphasis on subtitle focus with video graphic elements, which has improved the audience's attractiveness, attention, memory and acceptance of important video content, thereby enhancing the communication influence of video.
Smart Images

Figure CN119996772A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a method, apparatus, computer device, computer-readable storage medium, and computer program product for displaying video graphic elements. Background Art
[0002] In today's digital age, the amount of video information has grown massively. It is difficult for viewers to quickly and accurately obtain key information from complex video content within their limited time and energy, which poses a challenge to the effective dissemination of video information.
[0003] In order to cope with the above challenges, we usually start with video editing and video image quality. For example, by better editing or improving the video image quality, we can increase the audience's attention to the video content, thereby improving the dissemination effect of the video.
[0004] However, the inventors have found that the above method has limited effect on the dissemination of videos and cannot enable viewers to focus on important video content more efficiently and accurately.
[0005] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the invention
[0006] The embodiments of the present application provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for displaying video graphic elements to solve or alleviate one or more of the technical problems raised above.
[0007] An aspect of an embodiment of the present application provides a method for displaying a video graphic element, the method comprising: Extract subtitle highlights from the target video; Performing object detection processing on the video frames in the target video that are associated with the subtitle highlights; When it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target video graphic element is displayed in a visually emphasized manner.
[0008] Optionally, the method further comprises: In the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, generating a target video graphic element matching the subtitle emphasis; The target video graphic element is displayed in a corresponding area in the video frame using the visual emphasis method.
[0009] Optionally, in the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, generating a target video graphic element matching the subtitle emphasis includes: In the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, a pre-trained video graphic element generation model is used to generate a target video graphic element matching the subtitle emphasis.
[0010] Optionally, extract subtitle highlights from the target video, including: In the case where there is subtitle text in the target video, the subtitle text in the target video is analyzed using natural language processing technology to obtain the subtitle highlights; or, In the case where there is no subtitle text in the target video, performing speech-to-text processing on the audio in the target video to obtain subtitle text; The subtitle text is analyzed using natural language processing technology to obtain the subtitle highlights.
[0011] Optionally, performing object detection processing on the video frame associated with the subtitle focus in the target video includes: Obtaining a display timestamp corresponding to the subtitle highlights; Determining a video frame associated with the subtitle highlight according to the display timestamp; A pre-trained object detection model is used to perform object detection processing on the video frame.
[0012] Optionally, there are multiple visual emphasis methods, and when it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target video graphic element is displayed in a visual emphasis manner, including: In the case where a target video graphic element matching the subtitle emphasis is detected in the video frame, determining a target visual emphasis mode corresponding to the subtitle emphasis; The target video graphic element is displayed in the determined target visual emphasis manner.
[0013] Optionally, in the case where a target video graphic element matching the subtitle emphasis is detected in the video frame, determining a target visual emphasis mode corresponding to the subtitle emphasis includes: When it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target visual emphasis method is determined according to the viewer's visual preference.
[0014] Another aspect of an embodiment of the present application provides a video graphic element display device, the device comprising: An extraction module, used to extract subtitle highlights from the target video; A detection module, used for performing object detection processing on the video frames in the target video that are associated with the key points of the subtitles; The display module is used to display the target video graphic element in a visually emphasized manner when detecting that there is a target video graphic element matching the subtitle emphasis in the video frame.
[0015] Another aspect of an embodiment of the present application provides a computer device, including: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein: the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0016] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.
[0017] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the method described above when executed by a processor.
[0018] The embodiments of the present application adopt the above-mentioned technical solution, which may include the following advantages: by extracting the key points of subtitles from the target video, and based on the key points of subtitles, displaying the video graphic elements matching the key points of subtitles in a visually emphasized manner in the video frame associated with the key points of subtitles, the key points of subtitles are closely combined with the video graphic elements and emphasized, thereby establishing a more intelligent and intuitive connection between subtitles and video content, improving the audience's attraction, attention, memory and acceptance of important video content, and thus enhancing the dissemination influence of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0020] Figure 1 A diagram schematically shows an operating environment of a method for displaying video graphic elements according to Embodiment 1 of the present application; Figure 2A flowchart of a method for displaying video graphic elements according to Embodiment 1 of the present application is schematically shown; Figure 3 Schematically shows Figure 1 Flow chart of sub-steps in step S202; Figure 4 The newly added flow chart of the method for displaying video graphic elements according to the first embodiment of the present application is schematically shown; Figure 5 Schematically shows Figure 1 Flow chart of sub-steps in step S204; Figure 6 A block diagram schematically shows a video graphic element display device according to the second embodiment of the present application; and Figure 7 The hardware architecture diagram of the computer device according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0022] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0023] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.
[0024] First, the following terms are explained: Subtitle highlights: In the video subtitle text, words, phrases or sentence parts that contain key information, semantic core or require special attention are determined by specific algorithms or rules. For example, in a historical documentary, the names of historical events, key figures, important time nodes, etc. can all be used as subtitle highlights.
[0025] Video graphic elements: various visual components such as graphics, images, icons, lines, shapes, color areas, etc. that already exist in the video; or when there are no graphic elements corresponding to the subtitle focus in the video, new graphic elements are generated and inserted into the video based on the subtitle focus information. These graphic elements are used to assist in expressing information, directing the audience's attention, or enhancing the visual effect of the video.
[0026] Object detection: Use computer vision technology to identify and locate specific objects or areas in the video, such as people, vehicles, text, signs, etc.
[0027] To facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: In the prior art, in order to improve the dissemination effect of video information, video editing and video picture quality are generally taken into consideration. For example, by better editing of the video or improving the video picture quality, the audience's attention to the video content is increased, thereby improving the dissemination effect of the video.
[0028] However, the inventors have found that the above method has limited effect on the dissemination of videos and cannot enable viewers to focus on important video content more efficiently and accurately.
[0029] To this end, the embodiment of the present application provides a technical solution for displaying video graphic elements. In this technical solution, by extracting the key points of the subtitles from the target video, and based on the key points of the subtitles, displaying the video graphic elements matching the key points of the subtitles in a visually emphasized manner in the video frame associated with the key points of the subtitles, the key points of the subtitles are closely combined with the video graphic elements and emphasized, thereby establishing a more intelligent and intuitive association between the subtitles and the video content, improving the audience's attraction, attention, memory and acceptance of important video content, and thus enhancing the dissemination influence of the video. See below for details.
[0030] Finally, for ease of understanding, an exemplary operating environment is provided below.
[0031] like Figure 1 As shown, the environment diagram includes a service platform 2, a network 4, and a client 6, wherein: The service platform 2 may be composed of a single or multiple computing devices. The multiple computing devices may include virtualized computing instances. Virtualized computing instances may include virtual machines, such as simulations of computer systems, operating systems, servers, etc. The computing device may load a virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for simulation. As the demand for different types of processing services changes, different virtual machines may be loaded and / or terminated on one or more computing devices. A hypervisor may be implemented to manage the use of different virtual machines on the same computing device.
[0032] The service platform 2 may be configured to communicate with the client 6 or the like via a network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 may include physical links, such as coaxial cable links, twisted pair cable links, optical fiber links, combinations thereof, and the like, or wireless links, such as cellular links, satellite links, Wi-Fi links, and the like.
[0033] The service platform 2 can provide storage, reading, writing, querying, deleting and other services, such as providing web page access services for clients.
[0034] The client 6 may be an electronic device running an operating system such as Windows, Android™ or iOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, a vehicle terminal, or a smart TV. Based on the above operating system, various applications may be run, such as an application for displaying video graphic elements.
[0035] The client 6 may provide / configure a user access page for manipulating the service platform 2 or uploading an object, etc.
[0036] It should be noted that the above devices are exemplary, and the number and type of devices are adjustable in different scenarios or according to different needs.
[0037] The technical solutions of the present application are described below through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be construed as being limited to the embodiments described here.
[0038] Embodiment 1 Figure 2 The flowchart of the video graphic element display method according to the first embodiment of the present application is schematically shown.
[0039] like Figure 2 As shown, the video graphic element display method may include steps S200 to S204, wherein: Step S200, extracting subtitle highlights from the target video.
[0040] Step S202: performing object detection processing on the video frames in the target video that are associated with the subtitle highlights.
[0041] Step S204: when it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target video graphic element is displayed in a visually emphasized manner.
[0042] The video graphic element display method provided in this embodiment achieves the close integration and emphasis of the subtitle focus and the video graphic element by extracting the subtitle focus from the target video and displaying the video graphic element matching the subtitle focus in the video frame associated with the subtitle focus in a visual emphasis manner based on the subtitle focus, thereby establishing a more intelligent and intuitive connection between the subtitle and the video content, improving the audience's attraction, attention, memory and acceptance of important video content, and further enhancing the dissemination influence of the video.
[0043] The following combination Figure 2 , each step in steps S200~S204 and other optional steps are described in detail.
[0044] Step S200 , extract subtitle highlights from the target video.
[0045] The subtitle focus is the words, phrases or sentence parts in the video subtitle text that are determined by specific algorithms or rules to have key information, semantic core or require special attention.
[0046] For example, in a historical documentary, the names of historical events, key figures, important time nodes, etc. can all be used as subtitle highlights. In another documentary, the "Sahara Desert" in the subtitles can be used as the subtitle highlight.
[0047] As an example, in an educational video, "molecular structure of carbon dioxide" in the subtitles can be used as the subtitle focus.
[0048] As an example, in a commercial video, "the latest smartphone" in the subtitles can serve as the subtitle emphasis.
[0049] It should be noted that the video subtitle text mentioned above can be the subtitle text existing in the video itself, or can be the subtitle text obtained by converting the speech in the video into text.
[0050] In an optional implementation, extracting subtitle highlights from a target video includes: when there is subtitle text in the target video, using natural language processing technology to analyze the subtitle text in the target video to obtain the subtitle highlights.
[0051] Natural Language Processing (NLP) is an important research direction in the field of artificial intelligence, integrating knowledge from multiple disciplines such as linguistics, computer science, machine learning, mathematics, cognitive psychology, etc. Natural language processing technology includes two main aspects: natural language understanding and natural language generation. The research content of natural language processing technology includes multiple levels such as characters, words, phrases, sentences, paragraphs and chapters, and is a bridge for communication between machine language and human language.
[0052] In this embodiment, natural language processing technology can be used to first perform syntax analysis and semantic analysis on all subtitle texts in the target video, and then, based on the results of the syntax analysis and semantic analysis, the key points of the subtitles are extracted from the subtitle texts.
[0053] Among them, grammatical analysis is the process of analyzing sentence structure, including part-of-speech tagging, parsing, and dependency parsing, etc. Syntactic analysis helps to understand the structure and semantics of a sentence.
[0054] Among them, semantic analysis is the process of understanding the meaning of a sentence, including word sense disambiguation (Word Sense Disambiguation), named entity recognition (Named Entity Recognition, NER) and semantic role labeling (Semantic Role Labeling), etc. Semantic analysis helps computers understand the deep meaning of text.
[0055] In a specific example, natural language processing technology (e.g., BERT model, large language model) can be used to analyze the subtitle text in the target video, and key words used to describe scenes, characters, objects, important time nodes, etc. in the subtitle text can be extracted as subtitle highlights.
[0056] In an optional implementation, extracting subtitle highlights from a target video further includes: when no subtitle text exists in the target video, performing speech-to-text processing on the audio in the target video to obtain subtitle text; and using natural language processing technology to analyze the subtitle text to obtain the subtitle highlights.
[0057] In this embodiment, when there is no subtitle text in the target video and only audio is present, the audio will first be processed by speech-to-text conversion to obtain the corresponding subtitle text, and then the subtitle text will be analyzed using natural language processing technology to obtain the subtitle key points.
[0058] In this embodiment, speech-to-text processing is performed through the audio in the target video, so that the key points of the subtitles can be extracted from the target video without subtitle text.
[0059] Step S202 , performing object detection processing on the video frames in the target video that are associated with the subtitle focus.
[0060] The video frame associated with the subtitle highlight refers to the video frame that has an associated relationship with the subtitle highlight.
[0061] The object detection process is used to detect whether there is a target video graphic element matching the subtitle emphasis in the video frame.
[0062] In this embodiment, various object detection algorithms can be used to detect video frames. For example, an object detection model designed based on a decision tree or a random forest algorithm can be used to detect video frames. For another example, a region-based object detection model can be used to detect video frames, or a detection box-based object detection model can be used to detect video frames.
[0063] In an alternative embodiment, see Figure 3 , step S202 may include: Step S300, obtaining a display timestamp corresponding to the subtitle emphasis.
[0064] Step S302: determining a video frame associated with the subtitle highlight according to the display timestamp.
[0065] Step S304: Perform object detection processing on the video frame using a pre-trained object detection model.
[0066] In this embodiment, since the subtitle focus belongs to the words in the subtitle text in the target video, and each sentence of the subtitle text in the target video has a corresponding display timestamp, that is, a time node when the corresponding subtitle appears, the subtitle focus also has a corresponding display timestamp in the target video.
[0067] For example, if the time node of the appearance of a subtitle text is 20 minutes 15 seconds to 20 minutes 16 seconds, the display timestamp of the subtitle highlights extracted from the subtitle text is also 20 minutes 15 seconds to 20 minutes 16 seconds.
[0068] After the display timestamp is determined, it can be determined which video frames are associated with the subtitle highlight according to the display timestamp.
[0069] As an example, if the display timestamp of the subtitle highlight is 20 minutes 15 seconds to 20 minutes 16 seconds, all video frames displayed from 20 minutes 15 seconds to 20 minutes 16 seconds can be used as video frames associated with the subtitle highlight, or some video frames displayed from 20 minutes 15 seconds to 20 minutes 16 seconds can be used as video frames associated with the subtitle highlight.
[0070] After determining the video frame, if the determined video frame is multiple frames, a pre-trained object detection model can be used to perform object detection processing on the multiple video frames in sequence, so as to detect whether there is a target video graphic element matching the subtitle focus in each video frame.
[0071] The object detection model may be obtained after training based on network models such as YOLO and R-CNN.
[0072] In this embodiment, the video frame is acquired by combining the display timestamp of the subtitle highlight, so that the video frame associated with the subtitle highlight can be accurately located.
[0073] Step S204 When it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target video graphic element is displayed in a visually emphasized manner.
[0074] The target video graphic element refers to various visual components such as graphics, images, icons, lines, shapes or color areas in the video screen that match the subtitle emphasis. When there is no graphic element corresponding to the subtitle emphasis in the video, the target video graphic element can also be a graphic element newly generated and inserted into the video based on the subtitle emphasis information. These graphic elements are used to assist in expressing information, guide the audience's attention or enhance the visual effect of the video.
[0075] The visual emphasis method is used to emphasize the display of the target video graphic element. The visual emphasis method includes but is not limited to: arrow indication, frame highlighting, color change processing, special effect display, etc.
[0076] The arrow indicates that an arrow is added to the target video graphic element to point to a corresponding subtitle emphasis, or an arrow is added to the subtitle emphasis to point to the target video graphic element.
[0077] The frame highlighting refers to adding a highlighted frame to the periphery of the target video graphic element to highlight the target video graphic element.
[0078] The color change process refers to changing the color of the target video graphic element, for example, changing the color of the target video graphic element to a color with a highlight effect such as red or yellow.
[0079] The special effect display refers to adding preset special effects to the target video graphic elements, and the special effects include flickering, magnification, shadow and other special effects.
[0080] In this embodiment, after detecting the target video graphic element, the target video graphic element is displayed in a visually emphasized manner, so that the target video graphic element can be highlighted, thereby enhancing the user's attention and memory of the target video graphic element, thereby improving the dissemination influence of the video.
[0081] In an alternative embodiment, see Figure 4 , the method further comprises: Step S400: generating a target video graphic element matching the subtitle emphasis when it is detected that the target video graphic element matching the subtitle emphasis does not exist in the video frame.
[0082] Step S402: displaying the target video graphic element in a corresponding area in the video frame using the visual emphasis method.
[0083] When it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, a target video graphic element matching the subtitle emphasis is generated based on the subtitle emphasis, thereby ensuring semantic consistency between the generated target video graphic element and the subtitle content.
[0084] The generated target video graphic elements may be static graphic elements, such as characters, objects, scenes, etc. The generated target video graphic elements may be dynamic graphic elements, such as animated paths, annotation boxes, etc.
[0085] As an example, when the subtitle focuses on "carbon dioxide molecular structure", a "carbon dioxide molecular model schematic diagram" can be generated as the corresponding target video graphic element.
[0086] As an example, when the subtitle focuses on "Sahara Desert", "a schematic map of the Sahara Desert" may be generated as the corresponding target video graphic element.
[0087] As an example, when the subtitle focuses on "the latest smartphone", "3D model and detailed description icon of smartphone" can be generated as the corresponding target video graphic element.
[0088] It should be noted that the corresponding area in the video frame may be a background area in the video frame, or a blank area in the video frame, or may be any area in the video frame.
[0089] In this embodiment, after the target video graphic element is generated, the visual emphasis method is used to display the target video graphic element in the corresponding area of the video frame, so that the generated target video graphic element can be highlighted to enhance the visual appeal.
[0090] In this embodiment, when there is no matching target video graphic element in the video frame, the target video graphic element is generated, so that all kinds of videos can be applied, thereby enhancing the information dissemination effect of the video in different application scenarios.
[0091] In one embodiment, in order to improve the display effect and enhance the user experience, after the target video graphic element is generated, the style of the generated target video graphic element can be adaptively adjusted. For example, the shape, size, and color of the graphic element can be dynamically adjusted to match the visual style and background complexity of the video frame.
[0092] In an optional implementation, step S400 further includes: In the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, a pre-trained video graphic element generation model is used to generate a target video graphic element matching the subtitle emphasis.
[0093] The video graphic element generation model can be obtained by training a generative network based on GAN (Generative Adversarial Networks), or by fine-tuning an existing multimodal large language model.
[0094] In this embodiment, the target video graphic element is generated by using a pre-trained video graphic element generation model, thereby ensuring the semantic consistency between the generated target video graphic element and the subtitle content.
[0095] In an alternative embodiment, see Figure 5 There are multiple visual emphasis methods. The multiple visual emphasis methods can be selected later according to actual user settings or system default rules. Accordingly, step S204 may include: Step S500: when it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, a target visual emphasis mode corresponding to the subtitle emphasis is determined.
[0096] Step S502: displaying the target video graphic element in the determined target visual emphasis manner.
[0097] In this embodiment, when it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target visual emphasis method corresponding to the subtitle emphasis can be determined based on one or more dimensions such as the type of video, playback progress, characteristics of the subtitle emphasis (such as importance, frequency of occurrence, duration, etc.), and the audience's visual perception habits.
[0098] After the target visual emphasis mode is determined, the target video graphic element may be displayed using the determined target visual emphasis mode.
[0099] In this embodiment, by pre-setting a variety of target visual emphasis methods, the emphasis method of graphic elements can be adjusted in real time and flexibly later, while effectively highlighting key information and ensuring that the audience has a comfortable and smooth viewing experience and avoids visual fatigue or information overload.
[0100] In an optional implementation manner, when it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, determining the target visual emphasis mode corresponding to the subtitle emphasis includes: When it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target visual emphasis method is determined according to the viewer's visual preference.
[0101] The audience refers to the account owner of the application that plays the target video.
[0102] As an example, the visual preference may be determined based on the historical viewing data of the account owner. For example, the historical viewing data of viewers may be collected, and the historical viewing data includes the viewers' preferences for various visual elements in the video, including color, font, image style, scene transition, and other visual elements. By analyzing the frequency of use of these elements in the video and the audience's feedback, the audience's preferences for different visual elements may be understood, and the audience's visual preferences may be obtained.
[0103] As an example, the visual preference can also utilize AI visual technology to monitor the facial expressions of viewers in real time while watching videos. By analyzing the changes in facial expressions, the emotional reactions of viewers to different visual content can be understood, thereby inferring their visual preferences.
[0104] As an example, the visual preference can also be recorded by eye tracking technology, where the viewer's gaze movement track is recorded when watching a video. The duration of the viewer's gaze and the point of gaze can reflect their attention to different visual elements, thereby inferring their visual preference.
[0105] In this embodiment, the target visual emphasis method is determined according to the visual preference of the audience, so that the target visual emphasis method finally adopted can be more matched with the audience, thereby improving the user experience.
[0106] Embodiment 2 Figure 6 The block diagram of the video graphic element display device 600 according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Figure 6 As shown, the video graphic element display device 600 may include: an extraction module 610, a detection module 620 and a display module 630, wherein: An extraction module 610 is used to extract subtitle highlights from a target video; A detection module 620, configured to perform object detection processing on a video frame in the target video that is associated with the subtitle focus; The display module 630 is configured to display the target video graphic element in a visually emphasized manner when detecting that there is a target video graphic element matching the subtitle emphasis in the video frame.
[0107] As an optional embodiment, the video graphic element display device 600 further includes a generation module.
[0108] The generating module is used for generating a target video graphic element matching the subtitle emphasis when it is detected that the target video graphic element matching the subtitle emphasis does not exist in the video frame.
[0109] The display module 630 is further configured to display the target video graphic element in a corresponding area in the video frame by using the visual emphasis method.
[0110] As an optional embodiment, in the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, generating a target video graphic element matching the subtitle emphasis includes: In the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, a pre-trained video graphic element generation model is used to generate a target video graphic element matching the subtitle emphasis.
[0111] As an optional embodiment, extracting subtitle highlights from a target video includes: In the case where there is subtitle text in the target video, the subtitle text in the target video is analyzed using natural language processing technology to obtain the subtitle highlights; or, In the case where there is no subtitle text in the target video, performing speech-to-text processing on the audio in the target video to obtain subtitle text; The subtitle text is analyzed using natural language processing technology to obtain the subtitle highlights.
[0112] As an optional embodiment, the performing object detection processing on the video frame associated with the subtitle focus in the target video includes: Obtaining a display timestamp corresponding to the subtitle highlights; Determining a video frame associated with the subtitle highlight according to the display timestamp; A pre-trained object detection model is used to perform object detection processing on the video frame.
[0113] As an optional embodiment, there are multiple visual emphasis methods. When detecting that there is a target video graphic element matching the subtitle emphasis in the video frame, displaying the target video graphic element in a visual emphasis manner includes: In the case where a target video graphic element matching the subtitle emphasis is detected in the video frame, determining a target visual emphasis mode corresponding to the subtitle emphasis; The target video graphic element is displayed in the determined target visual emphasis manner.
[0114] As an optional embodiment, in the case where a target video graphic element matching the subtitle emphasis is detected in the video frame, determining a target visual emphasis mode corresponding to the subtitle emphasis includes: When it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target visual emphasis method is determined according to the viewer's visual preference.
[0115] Embodiment 3 Figure 7The schematic diagram of the hardware architecture of a computer device 10000 suitable for implementing the method for displaying video graphic elements according to the third embodiment of the present application is shown schematically. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. Figure 7 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as a hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk equipped on the computer device 10000, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 10010 can also include both the internal storage module of the computer device 10000 and its external storage device. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed in the computer device 10000, such as the program code of the video graphic element display method, etc. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.
[0116] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0117] The network interface 10030 may include a wireless network interface or a wired network interface, and the network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0118] It should be pointed out that Figure 7 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementation of all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.
[0119] In this embodiment, the video graphic element display method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as processor 10020) to complete the embodiment of the present application.
[0120] Embodiment 4 An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the video graphic element display method in the embodiment are implemented.
[0121] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the computer-readable storage medium can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store an operating system and various application software installed on the computer device, such as the program code of the video graphic element display method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or are to be output.
[0122] Embodiment 5 An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.
[0123] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0124] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for displaying video graphic elements, characterized in that: The method comprises: Extract subtitle highlights from the target video; Performing object detection processing on the video frames in the target video that are associated with the subtitle highlights; When it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target video graphic element is displayed in a visually emphasized manner.
2. The method according to claim 1, characterized in that The method further comprises: In the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, generating a target video graphic element matching the subtitle emphasis; The target video graphic element is displayed in a corresponding area in the video frame using the visual emphasis method.
3. The method according to claim 2, characterized in that The step of generating a target video graphic element matching the subtitle emphasis when detecting that the target video graphic element matching the subtitle emphasis does not exist in the video frame comprises: In the case where it is detected that there is no target video graphic element matching the subtitle emphasis in the video frame, a pre-trained video graphic element generation model is used to generate a target video graphic element matching the subtitle emphasis.
4. The method according to claim 1, characterized in that: Extract subtitle highlights from the target video, including: In the case where there is subtitle text in the target video, the subtitle text in the target video is analyzed using natural language processing technology to obtain the subtitle highlights; or, In the case where there is no subtitle text in the target video, performing speech-to-text processing on the audio in the target video to obtain subtitle text; The subtitle text is analyzed using natural language processing technology to obtain the subtitle highlights.
5. The method according to claim 1, characterized in that The performing object detection processing on the video frame associated with the subtitle focus in the target video includes: Obtaining a display timestamp corresponding to the subtitle highlights; Determining a video frame associated with the subtitle highlight according to the display timestamp; A pre-trained object detection model is used to perform object detection processing on the video frame.
6. The method according to any one of claims 1 to 5, characterized in that: There are multiple visual emphasis methods. When a target video graphic element matching the subtitle emphasis is detected in the video frame, the target video graphic element is displayed in a visual emphasis manner, including: In the case where a target video graphic element matching the subtitle emphasis is detected in the video frame, determining a target visual emphasis mode corresponding to the subtitle emphasis; The target video graphic element is displayed in the determined target visual emphasis manner.
7. The method according to claim 6, characterized in that In the case where a target video graphic element matching the subtitle emphasis is detected in the video frame, determining a target visual emphasis mode corresponding to the subtitle emphasis includes: When it is detected that there is a target video graphic element matching the subtitle emphasis in the video frame, the target visual emphasis method is determined according to the viewer's visual preference.
8. A video graphic element display device, characterized in that: The device comprises: An extraction module, used to extract subtitle highlights from the target video; A detection module, used for performing object detection processing on the video frames in the target video that are associated with the key points of the subtitles; The display module is used to display the target video graphic element in a visually emphasized manner when detecting that there is a target video graphic element matching the subtitle emphasis in the video frame.
9. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 7 are implemented.