Method and device for publishing video content and identifying chapters, equipment and storage medium
By identifying the image frame text information in the video content and integrating the text box position information, the video chapters are accurately identified, and the problem of inaccurate chapter identification in the prior art is solved, and the user's viewing experience is improved.
Patent Information
- Application Number
- CN202510180975.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art cannot accurately identify chapter information in video content, resulting in users being unable to accurately jump to expected chapters, affecting the viewing experience.
By identifying text recognition information in the image frame of the video content, integrating the position information of the text box, determining the chapter text box indicating the chapter information, and determining the chapter information of the video content based on the text content in the chapter text box.
Accurate chapter recognition of video content is achieved, the efficiency of setting chapter titles is improved, and the user's viewing experience is enhanced.
Smart Images

Figure CN120050447A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for publishing video content and identifying chapters. Background Art
[0002] With the rapid development of computer technology, short video social platforms have gradually become a social medium for people. People can create video content and publish it on short video social platforms to record and share their lives. When creating video content, people can add a progress bar to the video content to make it easier for others to watch the video content. Summary of the invention
[0003] In a first aspect of the present disclosure, a method for publishing video content is provided. The method includes: presenting a publishing interface for video content; in response to obtaining first chapter information identified based on the video content, presenting guide information in the publishing interface; in response to triggering the guide information, presenting a chapter setting interface for the video content, the chapter setting interface presenting a first group of chapter titles determined based on the first chapter information; and determining second chapter information of the video content via the chapter setting interface to publish the video content, the second chapter information at least indicating a second group of chapter titles of the video content.
[0004] In a second aspect of the present disclosure, a method for identifying chapters is provided. The method includes: determining text recognition information of an image frame of video content, the text recognition information indicating a first group of text boxes in the image frame; based on position information of the first group of text boxes, fusing at least one text box in the first group of text boxes to determine a second group of text boxes; based on trajectory information of the second group of text boxes in multiple image frames, determining a chapter text box indicating chapter information from the second group of text boxes; and determining chapter information of the video content based on text content in the chapter text box, the chapter information indicating multiple chapter titles of the video content.
[0005] In a third aspect of the present disclosure, a device for publishing video content is provided. The device includes: a first presentation module configured to present a publishing interface of the video content; a second presentation module configured to present guidance information in the publishing interface in response to obtaining first chapter information identified based on the video content; a third presentation module configured to present a chapter setting interface of the video content in response to triggering the guidance information, the chapter setting interface presenting a first group of chapter titles determined based on the first chapter information; and a first determination module configured to determine second chapter information of the video content via the chapter setting interface to publish the video content, the second chapter information at least indicating a second group of chapter titles of the video content.
[0006] In a fourth aspect of the present disclosure, a device for identifying chapters is provided. The device includes: a second determination module configured to determine text recognition information of an image frame of video content, the text recognition information indicating a first group of text boxes in the image frame; a fusion module configured to, in response to the first group of text boxes including a plurality of positionally associated text boxes, fuse a plurality of associated text boxes in the first group of text boxes to determine a second group of text boxes; a screening module configured to determine a chapter text box indicating chapter information from the second group of text boxes based on trajectory information in the second group of text boxes in a plurality of image frames; and a third determination module configured to determine chapter information of the video content based on text content in the chapter text box, the chapter information indicating a plurality of chapter titles of the video content.
[0007] In a fifth aspect of the present disclosure, a computing device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect or the second aspect.
[0008] In a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect or the second aspect.
[0009] It should be understood that the contents described in this content section are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure may be implemented;
[0012] Figure 2 A flowchart illustrating an example process for publishing video content according to some embodiments of the present disclosure;
[0013] FIG. 3A to FIG. 3C shows an example interface according to some embodiments of the present disclosure;
[0014] Figure 4 A flowchart illustrating an example process of identifying chapters according to some embodiments of the present disclosure;
[0015] Figure 5 A schematic diagram showing an example process of identifying a chapter according to some embodiments of the present disclosure;
[0016] Figure 6 shows an example image frame according to some embodiments of the present disclosure;
[0017] Figure 7 A schematic structural block diagram of an example device for publishing video content according to some embodiments of the present disclosure is shown;
[0018] Figure 8 A schematic structural block diagram showing an example apparatus for identifying chapters according to some embodiments of the present disclosure; and
[0019] Fig. 9 A block diagram of a computing device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0021] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this article, and any type of embodiment may be included under any section / subsection. In addition, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0022] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0023] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects are subject to the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user knows and confirms. Accordingly, when implementing each embodiment of the present disclosure, the type, scope of use, usage scenario, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0024] In this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or it is necessary to perform a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.
[0025] As mentioned above, currently the user-made progress bar cannot enable the viewer to accurately jump to the desired chapter, so that the viewer cannot obtain a good viewing experience.
[0026] In order to improve the viewing experience of viewers, short video social platforms need to identify the video content uploaded by users and re-create the progress bar based on the chapter content in the video content. However, the current methods for identifying chapters cannot accurately identify them.
[0027] The embodiment of the present disclosure proposes a scheme for publishing video content. The scheme includes: presenting a publishing interface for video content; in response to obtaining first chapter information identified based on the video content, presenting guide information in the publishing interface; in response to triggering the guide information, presenting a chapter setting interface for the video content, the chapter setting interface presenting a first group of chapter titles determined based on the first chapter information; and determining second chapter information of the video content via the chapter setting interface to publish the video content, the second chapter information at least indicating a second group of chapter titles of the video content.
[0028] In this way, the embodiments of the present disclosure can automatically add chapter titles by detecting the chapter information of the video (eg, a self-made progress bar), thereby improving the efficiency of setting chapter titles.
[0029] Various example implementations of the solution are described in detail below in conjunction with the accompanying drawings.
[0030] Example Environment
[0031] Figure 11 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, example environment 100 may include electronic device 110 .
[0032] In this example environment 100, the electronic device 110 may run an application 120 that supports publishing video content. The application 120 may be any suitable type of application for publishing video content, and examples thereof may include, but are not limited to, short video social platforms or other suitable applications. The user 140 may interact with the application 120 via the electronic device 110 and / or its attached devices.
[0033] exist Figure 1 In the environment 100 , if the application 120 is in an active state, the electronic device 110 may present an interface 150 for supporting the release of video content through the application 120 .
[0034] In some embodiments, the electronic device 110 communicates with the server 130 to provide services for the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio receiver, an e-book device, a game device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0035] The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The server 130 may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. The server 130 may provide background services for the application 120 in the electronic device 110 that supports publishing video content.
[0036] A communication connection may be established between the server 130 and the electronic device 110. The communication connection may be established in a wired manner or a wireless manner. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the electronic device 110 may implement signaling interaction through the communication connection between the two.
[0037] It should be understood that the structure and function of the various elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0038] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0039] Publish video content
[0040] Combine the following Figure 2 , FIG. 3A to FIG. 3C To describe the specific interactive process of publishing video content. Figure 2 A flow chart is shown of an example process 200 for publishing video content according to some embodiments of the present disclosure. FIG. 3A to FIG. 3C 2 shows example interfaces 300A to 300C according to some embodiments of the present disclosure. The process 200 may be implemented by a suitable electronic device, such as the electronic device 110. The interfaces 300A to 300C may be implemented by, for example, Figure 1 The electronic device 110 is provided as shown.
[0041] like Figure 2 As shown, in box 210, the electronic device 110 presents a publishing interface for video content.
[0042] In some embodiments, when the electronic device 110 receives the video content uploaded by the user 140, the electronic device 110 may present the following Figure 3A The interface 300A shown is used to facilitate the user 140 to edit the information related to the video content to complete the release of the video content.
[0043] like Figure 3AAs shown, in some embodiments, interface 300A may include multiple editing items and controls for publishing video content. Multiple editing items are used to edit multiple pieces of information related to the video content. The editing items may be, for example, editing items for editing the name of the video content, editing items for editing the introduction of the video content, editing items for editing the type of the video content, or editing items for editing the chapter information of the video content. After user 140 completes editing of all editing items, when electronic device 110 receives a click operation of user 140 on the control, the published video content may be presented in an interface such as a social platform. Additionally, interface 200A may also include a preview area for previewing the video content.
[0044] In box 220, the electronic device 110 presents guide information in the publishing interface in response to obtaining the first chapter information identified based on the video content.
[0045] The aforementioned editing item for editing chapter information can be used by the user 140 to edit the chapter information of the video content. It should be understood that the video content uploaded by the user 140 can be a video content containing the chapter information made by the user 140 or a progress bar related to the chapter information, or it can be a video content without the chapter information made by the user 140 or a progress bar related to the chapter information. For the former, the electronic device 110 can identify the chapter information in the video content. For the latter, the user 140 can add chapter information to the video content through the editing item.
[0046] The following further describes the case where the video content is video content including chapter information created by the user 140 or a progress bar related to the chapter information.
[0047] In some embodiments, the electronic device 110 can identify the screen content of the video content by calling a recognition algorithm to obtain the identified first chapter information. Wherein, the screen content can, for example, indicate the text elements included in at least one image frame of the video content. It should be understood that at least one image frame of the video content can present text elements such as subtitles, a progress bar related to the chapter information, and text content appearing in the scene corresponding to the image frame. By identifying the screen content containing these text elements, the first chapter information can be determined therefrom. As an example, the recognition algorithm can be implemented by a suitable device such as the electronic device 110 or the server 130.
[0048] In some embodiments, when the electronic device 110 identifies the first chapter information from the screen content of the video content, the electronic device 110 may present the guide information 305 in the publishing interface. Figure 3A As shown, the guidance information 305 may include "the platform has detected a self-made progress bar".
[0049] The first chapter information may indicate multiple chapter titles determined based on chapter content created by user 140 in the video content or a progress bar related to the chapter content. Additionally, the first chapter information may also indicate the start time points or end time points of multiple chapters, for example.
[0050] In block 230, the electronic device 110 presents a chapter setting interface of the video content in response to the triggering of the guide information. The chapter setting interface presents a first group of chapter titles determined based on the first chapter information.
[0051] like Figure 3B As shown, in some embodiments, when the electronic device 110 receives a trigger of the guidance information from the user 140, the electronic device 110 may present a chapter setting interface such as interface 300B. The chapter setting interface may include at least an editing panel and an auxiliary panel for editing chapter information. The user 140 may edit the chapter information by operating the editing panel and the auxiliary panel.
[0052] In some embodiments, the editing panel may include a plurality of unit panels corresponding to the first group of chapter titles. Each unit panel is used to edit the chapter information of a chapter title in the first group of chapter titles. Wherein, the first group of chapter titles may be determined based on the first chapter information, for example. For example, interface 300B may display a plurality of chapter titles extracted from a user-made progress bar, for example, chapter title 310.
[0053] In some embodiments, the unit panel may include, for example, a first area for editing the title name of the first group of chapter titles, a second area for editing the introduction of the first group of chapter titles, and a third area for editing the time points related to the first group of chapter titles. The first area may, for example, present the title name of the first group of chapter titles. The second area may, for example, present the introduction of the first group of chapter titles. The third area may, for example, present the time points related to the first group of chapter titles, such as the starting time point or the ending time point of the chapter title. As an example, when the electronic device 110 receives a click operation of the user 140 on any editing area on the unit panel, the electronic device 110 may cause the corresponding editing area on the chapter setting interface to present an editing mode, so that the user 140 can edit the corresponding information. In another example, the unit panel may also include an editing control. When the electronic device 110 receives a click operation of the user 140 on the editing control, the electronic device 110 may cause the corresponding unit panel on the chapter setting interface to present an editing mode.
[0054] In some embodiments, the electronic device 110 associates the first group of chapter titles with the first group of time points of the video content, and the time points presented in the third area can be obtained. The first group of time points can be, for example, preset time points. As an example, the first group of time points can be determined based on the number of the first group of chapter titles. For example, the electronic device 110 can preset the first group of time points based on the ratio of the duration of the video content to the number of the first group of chapter titles. Or, the electronic device 110 can preset the first group of time points as time points separated by a fixed duration.
[0055] In some embodiments, the auxiliary panel may include at least a video preview area and a timeline. Among them, the video preview area can be used to preview the video content. Since the video content contains a progress bar, and the progress bar can indicate the starting time point or the ending time point of multiple chapter titles, as well as the video playback progress, the user 140 can adjust the time point of the third area by referring to the video content of the video preview area. The timeline can be used to obtain the time point. When the electronic device 110 receives the selection operation of the user 140 on any time point on the timeline, the electronic device 110 can obtain the specific time of the time point selected by the user 140 through the timeline. Further, the electronic device 110 can synchronize the time point selected by the user 140 from the timeline to the third area by associating the timeline with the time point of the third area, thereby adjusting the time point of the third area.
[0056] Additionally, the chapter setting interface may also include controls for adding or deleting chapters.
[0057] It should be understood that the scenario in which the chapter setting interface is presented by triggering the guide information may be, for example, the first time that the user 140 uploads the video content containing the first chapter information. Based on the operation of the guide information, the user 140 may cause the electronic device 110 to present the setting interface to edit the chapter information of the video content. For another scenario, that is, the scenario in which the user 140 is not editing the chapter information of the video content for the first time. The user 140 may click on the editing item for editing the chapter information or the control corresponding thereto, so that the electronic device 110 presents the chapter setting interface. Specifically, when the electronic device 110 receives the click operation of the user 140 on the editing item or the corresponding control, the electronic device 110 may present the chapter setting interface.
[0058] In block 240 , the electronic device 110 determines second chapter information of the video content via the chapter setting interface to publish the video content.
[0059] In some embodiments, the second chapter information at least indicates a second set of chapter titles of the video content. As an example, the user can confirm the multiple chapter titles indicated by the first chapter information, and can publish the video to associate the video to the multiple chapter titles automatically identified.
[0060] In some other embodiments, the electronic device 110 may also receive an editing operation on the first set of chapter titles to add a new chapter title, modify an existing chapter title, or modify an existing chapter title. Thus, the electronic device 110 may determine the second set of chapter titles.
[0061] Additionally, the second chapter information may, for example, indicate a chapter introduction corresponding to the second group of chapter titles and a starting time point or an ending time point of the second group of chapter titles in the video content.
[0062] It should be understood that since the second chapter information can be divided into at least three parts: a second group of chapter titles, a chapter introduction, and a corresponding start time point or end time point, the implementation of the aforementioned step of "determining the second chapter information of the video content" can also be divided into three parts.
[0063] Regarding determining the second group of chapter titles, in some embodiments, the second group of chapter titles may be determined based on the first group of chapter titles. As an example, when the electronic device 110 receives an operation of adding, deleting, or modifying the first group of chapter titles from the user 140, the electronic device 110 may adjust the first group of chapter titles to the second group of chapter titles and present them in the editing panel.
[0064] In a specific example, when the electronic device 110 receives a click operation from the user 140 on the first area of a unit panel or an editing control, the electronic device 110 causes the first area to enter an editing mode, and then, when the electronic device 110 receives an input operation from the user in the first area, the electronic device 110 causes the first area to present the chapter title input by the user 140. Through the above process, the electronic device 110 can adjust the first group of chapter titles to the second group of chapter titles via the chapter setting interface. Based on the operation of adding a chapter or deleting a chapter, the process of adjusting the first group of chapter titles to the second group of chapter titles is similar to the above process, and will not be described in detail here.
[0065] For determining the chapter introductions corresponding to the second group of chapter titles, the implementation process can refer to the implementation process of determining the second group of chapter titles, which will not be elaborated here.
[0066] In some embodiments, the electronic device 110 can determine the start time point or the end time point corresponding to the second group of chapter titles based on the following process: through the chapter setting interface, the second group of chapter titles is associated with the second group of time points of the video content. Specifically, the second group of time points can be determined based on the editing operation of the user 140.
[0067] like Figure 3CAs shown, the electronic device 110 presents an indication element in the video area of the chapter setting interface in response to the selection of the chapter title 315 in the second group of chapter titles. In some embodiments, when the electronic device 110 receives the user's selection operation of the chapter title in the second group of chapter titles, the electronic device 110 can make the unit panel corresponding to the chapter title in the chapter setting interface present a selected state. As an example, the selected state can be represented by marking the color of the unit panel or making the outline of the unit panel bold.
[0068] In some embodiments, the indicator element 325 can represent the starting time point or the ending time point corresponding to the chapter title. When the time point in the third area is represented as the starting time point of the chapter title, the indicator element can represent the starting time point corresponding to the chapter title. Conversely, when the time point in the third area is represented as the ending time point of the chapter title, the indicator element can represent the ending time point corresponding to the chapter title. As an example, the indicator element 325 can be an auxiliary line.
[0069] Then, the electronic device 110 updates the position of the indicator element in the video preview area in response to receiving the adjustment operation 320 of the start time point or the end time point. When receiving the user's drag operation 320, the indicator element 325 can move accordingly to indicate the adjusted start time point or end time point.
[0070] In this way, the embodiments of the present disclosure can help the user to more accurately locate the boundary positions of different chapters by displaying indicator elements, thereby improving the efficiency of setting chapter information.
[0071] Identify video chapters
[0072] Figure 4 4 is a flowchart showing an example process 400 for identifying chapters according to some embodiments of the present disclosure. The process 400 may be implemented at the server 130. Figure 1 4. The process 400 is described below.
[0073] like Figure 4 As shown, at block 410 , the server 130 determines text recognition information of an image frame of video content, the text recognition information indicating a first group of text boxes within the image frame.
[0074] The following will refer to Figure 5 To describe process 400. Figure 5 As shown, the server 130 may obtain an input video 505 and may extract frames or samples the input video 505 at block 510. For example, the server 130 may extract image frames from the input video 505 at fixed intervals (eg, every N seconds).
[0075] Furthermore, in block 515, the server 130 may perform text recognition on the sampled image frame. For example, the server 130 may use OCR (Optical Character Recognition) to recognize text content in the image frame to determine a first group of text boxes in the image frame.
[0076] Figure 6 An example image frame 600 is shown according to some embodiments of the present disclosure. Figure 6 As shown, server 130 may utilize a text recognition module to determine a first group of text boxes from image frame 600 , for example, text box 605 , text box 610 , text box 615 , text box 620 , text box 625 , text box 630 , text box 635 , and text box 640 .
[0077] In block 420 , in response to the first group of text boxes including a plurality of text boxes associated with each other in position, the server 130 merges the plurality of associated text boxes in the first group of text boxes to determine a second group of text boxes.
[0078] Specifically, the server 130 may determine the position information of the first group of text boxes. Further, the server 130 may determine the plurality of text boxes associated with positions from the group of text boxes based on the position information of the first group of text boxes.
[0079] In block 520, the server 130 may perform clustering on the text boxes in a preset direction (eg, the Y-axis direction). Figure 6 As an example, the server 130 may determine the position information of the text boxes 605 to 640 in the image frame 600 to aggregate the text boxes in the Y direction.
[0080] As an example, the difference between the coordinate values of text boxes 615 to 640 in the preset direction is less than a threshold, and the position association can be determined. Further, the server 130 can aggregate text boxes 615 to 640 into one text box, so that the first group of text boxes can be updated to the second group of text boxes.
[0081] In block 430 , the server 130 determines a chapter text box indicating chapter information from the second group of text boxes based on the trajectory information of the second group of text boxes in the plurality of image frames.
[0082] Specifically, in block 525, the server 130 may determine the trajectory information of the second group of text boxes in the sampled multiple image frames. As an example, the server 130 may match the text boxes in different image frames to determine the trajectory information of the text boxes.
[0083] Specifically, the server 130 may determine whether two text boxes in different image frames match based on indicators such as similarity of text contents in the text boxes, positional relationship between text boxes, similarity of text box sizes, and continuity of frames in which text boxes appear.
[0084] Further, at 530 , the server 130 may determine the chapter text box that satisfies the at least one filtering condition from the second group of text boxes based on a comparison between the track information and the at least one filtering condition.
[0085] Specifically, the filtering condition may include detecting whether the duration of the text box is greater than a preset duration in box 535. For example, if the duration of the text box is less than or equal to the preset duration, it can be determined as text content that has not appeared in the video content for a long time, and therefore is not a user-made chapter text.
[0086] Additionally or alternatively, the filtering condition may include detecting whether the size of the text box is greater than a preset size at block 535. For example, if the size (e.g., width) of the text box is less than or equal to a preset width, the text box may be determined to be a non-user-generated chapter text.
[0087] Additionally or alternatively, the filter condition may include detecting whether the number of valid characters in the text box is greater than a preset number in box 540, wherein the valid characters are determined based on the type of characters in the text box. As an example, a valid character may represent a character of a non-symbol type in the text box. Further, if the number of valid characters in the text box is less than or equal to the preset number, the text box may be determined as a non-user-made chapter text.
[0088] Additionally or alternatively, the filtering condition may include detecting that the distance from the text box to the upper boundary or the lower boundary is less than a preset distance in block 540. For example, since the self-made chapter text is usually at the top or bottom of the video content, if the distance from the text box to the upper and lower boundaries is greater than or equal to a threshold, it can be determined as a non-user-made chapter text.
[0089] Further, in block 555 , in response to the second group of text boxes including a plurality of candidate text boxes satisfying the at least one filtering condition, the server 130 may determine the chapter text box from the plurality of candidate text boxes using a classification model.
[0090] As an example, the classification model may include a SER (Semantic Entity Recognition, text entity classification) model. In some embodiments, the SER model can be implemented using a multimodal transformer. Specifically, the server 130 can construct text features based on the text content and position information corresponding to the multiple candidate text boxes. For example, the server 130 can use a text encoder to encode the text content and coordinate information of each text box, thereby determining the text features corresponding to each text box.
[0091] In addition, the server 130 may construct visual features based on the image frames corresponding to the multiple candidate text boxes. Specifically, multiple sub-images of the image frame are encoded using an image encoder to construct the visual features, wherein the image encoder is trained using contrastive learning of sample images and sample texts. As an example, the image encoder may include an encoding unit in a CLIP (Contrastive Language-Image Pre-training) model.
[0092] Further, the server 130 may provide the text features and the visual features to the classification model to determine a classification result, wherein the classification result indicates whether the plurality of candidate text boxes are chapter text boxes. For example, the classification model may output a classification result of whether each candidate text box is a chapter text box.
[0093] In this way, embodiments of the present disclosure may be used to filter erroneously recalled non-section text regions.
[0094] Continue to refer Figure 4 In block 440 , the server 130 determines the chapter information of the video content based on the text content in the chapter text box, where the chapter information indicates a plurality of chapter titles of the video content.
[0095] like Figure 5 As shown, in box 560, the server 130 can segment the text content in the chapter text box. Specifically, the server 130 can detect a set of preset delimiters in the text content. As an example, such preset delimiters can include preset delimiters such as "|", "-".
[0096] Furthermore, the server 130 may divide the text content into a plurality of chapter titles corresponding to a plurality of chapters based at least on the set of preset delimiters, so as to serve as the chapter information of the video content. Figure 6As an example, server 130 may segment the contents of the aggregated text boxes corresponding to text boxes 615 to 640 to determine multiple chapter titles, namely, “Chapter 1”, “Chapter 2”, “Chapter 3”, “Chapter 4”, “Chapter 5”, and “Chapter 6”.
[0097] In some embodiments, in order to improve the recognition accuracy of the chapter text, the server 130 may also obtain a sub-image corresponding to the chapter text box, and may provide the sub-image to a text recognition model to recognize the text content in the chapter text box. The text recognition model may be, for example, a prediction based on CTC (Connectionist Temporal Classification).
[0098] In the CTC-based prediction process, the text recognition model can insert blank characters at each moment of the output to align the input sequence, and can use blank characters to perform chapter text content segmentation to avoid the problem of text stickiness.
[0099] In this way, the embodiments of the present disclosure can identify chapter text (eg, a self-made progress bar) in video content by tracking text boxes, thereby improving the accuracy of chapter identification.
[0100] Example devices and equipment
[0101] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 7 A schematic structural block diagram of an example apparatus 700 for publishing video content according to some embodiments of the present disclosure is shown. The apparatus 700 may be implemented as or included in the electronic device 110. Each module / component in the apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.
[0102] like Figure 7 As shown, the device 700 includes: a first presentation module 710, configured to present a publishing interface of video content; a second presentation module 720, configured to present guidance information in the publishing interface in response to obtaining first chapter information identified based on the video content; a third presentation module 730, configured to present a chapter setting interface of the video content in response to triggering the guidance information, the chapter setting interface presenting a first group of chapter titles determined based on the first chapter information; and a first determination module 740, configured to determine second chapter information of the video content via the chapter setting interface to publish the video content, the second chapter information at least indicating a second group of chapter titles of the video content.
[0103] In some embodiments, the second set of section headings is determined based on the first set of section headings.
[0104] In some embodiments, the first chapter information is determined based on a text element included in at least one image frame of the video content.
[0105] In some embodiments, the device 700 also includes an association module, which is configured to: associate a first group of chapter titles to a first group of time points in the video content, the first group of time points being determined based on the number of the first group of chapter titles; and associate a second group of chapter titles to a second group of time points in the video content via a chapter setting interface.
[0106] In some embodiments, the association module is further configured to: in response to selection of a chapter title in the second group of chapter titles, present an indicator element in the video preview area of the chapter setting interface, the indicator element representing the starting time point or the ending time point corresponding to the chapter title; in response to receiving an adjustment operation on the starting time point or the ending time point, update the position of the indicator element in the video preview area; and associate the chapter title to the adjusted starting time point or the ending time point.
[0107] In some embodiments, the chapter setting interface further includes a timeline corresponding to the video content, and the position of the indicating element is determined based on the position of the start time point or the end time point in the timeline.
[0108] Figure 8 A schematic structural block diagram of an example apparatus 800 for publishing video content according to some embodiments of the present disclosure is shown. The apparatus 800 may be implemented as or included in the electronic device 110. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0109] like Figure 8 As shown, the device 800 includes: a second presentation module 810, configured to determine text recognition information of an image frame of video content, the text recognition information indicating a first group of text boxes in the image frame; a fusion module 820, configured to fuse multiple associated text boxes in the first group of text boxes in response to the first group of text boxes including multiple text boxes associated with positions, so as to determine a second group of text boxes; a screening module 830, configured to determine a chapter text box indicating chapter information from the second group of text boxes based on trajectory information in the second group of text boxes in multiple image frames; and a third determination module 840, configured to determine chapter information of the video content based on text content in the chapter text box, the chapter information indicating multiple chapter titles of the video content.
[0110] In some embodiments, the apparatus 800 further includes a fourth determination module, which is configured to: determine position information of the first group of text boxes; and determine a plurality of positionally associated text boxes from a group of text boxes based on the position information of the first group of text boxes.
[0111] In some embodiments, a difference between coordinate values of the plurality of text boxes in a preset direction is smaller than a threshold.
[0112] In some embodiments, the filtering module 830 is further configured to: determine, based on a comparison between the track information and the at least one filtering condition, a chapter text box that satisfies the at least one filtering condition from the second group of text boxes.
[0113] In some embodiments, at least one filtering condition includes at least one of the following: the duration of the text box is greater than a preset duration; the size of the text box is greater than a preset size; the number of valid characters in the text box is greater than a preset number, and the valid characters are determined based on the type of characters in the text box; the distance from the text box to the upper or lower border is less than a preset distance.
[0114] In some embodiments, the screening module 830 is further configured to: in response to the second group of text boxes including multiple candidate text boxes that satisfy at least one filtering condition, determine a chapter text box from the multiple candidate text boxes using a classification model.
[0115] In some embodiments, the screening module 830 is further configured to: construct text features based on the text content and position information corresponding to multiple candidate text boxes; construct visual features based on the image frames corresponding to the multiple candidate text boxes; and provide text features and visual features to the classification model to determine the classification results, the classification results indicating whether the multiple candidate text boxes are chapter text boxes.
[0116] In some embodiments, the screening module 830 is further configured to: encode multiple sub-images of the image frame using an image encoder to construct visual features, wherein the image encoder is trained using comparative learning of sample images and sample texts.
[0117] In some embodiments, the third determination module 840 is further configured to: detect a set of preset separators in the text content; and divide the text content into multiple chapter titles corresponding to multiple chapters based at least on the set of preset separators as chapter information of the video content.
[0118] In some embodiments, the apparatus 800 further includes an acquisition module, which is configured to: acquire a sub-image corresponding to the chapter text box; and provide the sub-image to a text recognition model to recognize text content within the chapter text box.
[0119] The units included in the device 700 or the device 800 can be implemented in various ways, including software, hardware, firmware or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 700 or the device 800 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0120] like Fig. 9 As shown, the electronic device 900 is in the form of a general electronic device. The components of the electronic device 900 may include, but are not limited to, one or more processors or processing units 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 920. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 900.
[0121] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 930 can be a removable or non-removable medium, and can include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 900.
[0122] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Fig. 9As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. Memory 920 may include a computer program product 925 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.
[0123] The communication unit 940 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0124] The input device 950 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 960 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 900 may also communicate with one or more external devices (not shown) through the communication unit 940 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 900, or communicate with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0125] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0126] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0127] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0128] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0129] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0130] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. A method for publishing video content, comprising: A publishing interface for presenting video content; In response to acquiring first chapter information identified based on the video content, presenting guide information in the publishing interface; In response to triggering the guide information, presenting a chapter setting interface of the video content, wherein the chapter setting interface presents a first group of chapter titles determined based on the first chapter information; as well as The second chapter information of the video content is determined via the chapter setting interface to publish the video content, wherein the second chapter information at least indicates a second group of chapter titles of the video content.
2. The method of claim 1, wherein the second set of chapter headings is determined based on the first set of chapter headings. 3 . The method according to claim 1 , wherein the first chapter information is determined based on a text element included in at least one image frame of the video content.
4. The method according to claim 1, further comprising: Associating the first set of chapter titles to a first set of time points in the video content, the first set of time points being determined based on the number of the first set of chapter titles; as well as The second group of chapter titles is associated with a second group of time points of the video content via the chapter setting interface.
5. The method according to claim 4, wherein associating the second set of chapter titles to the second set of time points of the video content via the chapter setting interface comprises: In response to selection of a chapter title in the second group of chapter titles, presenting an indicator element in a video preview area of the chapter setting interface, the indicator element representing a start time point or an end time point corresponding to the chapter title; In response to receiving an adjustment operation on the start time point or the end time point, updating the position of the indicator element in the video preview area; and The chapter title is associated with the adjusted start time point or the end time point.
6. The method according to claim 5, wherein the chapter setting interface further comprises a timeline corresponding to the video content, and the position of the indicator element is determined based on the position of the start time point or the end time point in the timeline.
7. A method for identifying a chapter, comprising: Determining text recognition information of an image frame of video content, the text recognition information indicating a first group of text boxes within the image frame; In response to the first group of text boxes including a plurality of text boxes associated with each other in position, fusing the plurality of associated text boxes in the first group of text boxes to determine a second group of text boxes; determining a chapter text box indicating chapter information from the second group of text boxes based on the track information in the plurality of image frames in the second group of text boxes; as well as Based on the text content in the chapter text box, the chapter information of the video content is determined, and the chapter information indicates a plurality of chapter titles of the video content.
8. The method according to claim 7, further comprising: Determine position information of the first group of text boxes; Based on the position information of the first group of text boxes, the plurality of text boxes associated with each other in positions are determined from the group of text boxes. 9 . The method according to claim 7 , wherein a difference between coordinate values of the plurality of text boxes in a preset direction is less than a threshold value.
10. The method according to claim 7, wherein determining a chapter text box indicating chapter information from the second group of text boxes based on the track information of the second group of text boxes in the plurality of image frames comprises: Based on the comparison between the track information and at least one filtering condition, the chapter text boxes satisfying the at least one filtering condition are determined from the second group of text boxes.
11. The method according to claim 10, wherein the at least one filtering condition comprises at least one of the following: The duration of the text box is longer than the preset duration; The size of the text box is larger than the preset size; The number of valid characters in the text box is greater than a preset number, and the valid characters are determined based on the types of characters in the text box; The distance from the text box to the upper or lower border is smaller than the preset distance.
12. The method according to claim 10, wherein determining the chapter text box satisfying the at least one filtering condition from the second group of text boxes comprises: In response to the second group of text boxes including a plurality of candidate text boxes satisfying the at least one filtering condition, the chapter text box is determined from the plurality of candidate text boxes using a classification model.
13. The method according to claim 12, wherein determining the chapter text box from the plurality of candidate text boxes using a classification model comprises: Constructing text features based on the text content and position information corresponding to the multiple candidate text boxes; constructing visual features based on the image frames corresponding to the plurality of candidate text boxes; as well as The text features and the visual features are provided to the classification model to determine a classification result, the classification result indicating whether the plurality of candidate text boxes are chapter text boxes.
14. The method according to claim 13, wherein constructing visual features based on image frames corresponding to the plurality of candidate text boxes comprises: The plurality of sub-images of the image frame are encoded using an image encoder to construct the visual features, wherein the image encoder is trained using comparative learning of sample images and sample texts.
15. The method according to claim 7, wherein determining the chapter information of the video content based on the text content in the chapter text box comprises: Detecting a set of preset delimiters in the text content; as well as At least based on the set of preset separators, the text content is divided into a plurality of chapter titles corresponding to a plurality of chapters to serve as the chapter information of the video content.
16. The method according to claim 7, further comprising: Obtaining a sub-image corresponding to the chapter text box; as well as The sub-image is provided to a text recognition model to recognize the text content within the chapter text box.
17. A device for publishing video content, comprising: A first presentation module is configured to present a publishing interface of the video content; A second presentation module is configured to present guide information in the publishing interface in response to acquiring first chapter information identified based on the video content; A third presentation module is configured to present a chapter setting interface of the video content in response to triggering the guide information, wherein the chapter setting interface presents a first group of chapter titles determined based on the first chapter information; as well as The first determination module is configured to determine second chapter information of the video content via the chapter setting interface to publish the video content, wherein the second chapter information at least indicates a second group of chapter titles of the video content.
18. A device for publishing video content, comprising: A second determination module is configured to determine text recognition information of an image frame of video content, wherein the text recognition information indicates a first group of text boxes within the image frame; a fusion module configured to, in response to the first group of text boxes including a plurality of text boxes associated with each other in position, fuse the plurality of associated text boxes in the first group of text boxes to determine a second group of text boxes; A screening module configured to determine a chapter text box indicating chapter information from the second group of text boxes based on the trajectory information of the second group of text boxes in a plurality of image frames; as well as The third determination module is configured to determine the chapter information of the video content based on the text content in the chapter text box, where the chapter information indicates a plurality of chapter titles of the video content.
19. A computing device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform a method according to any one of claims 1 to 6 or 7 to 17.
20. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 6 or 7 to 17.