Video division method, device and equipment, readable storage medium and program product

By combining the video picture and semantic division information, mapping processing is carried out to correct the picture division, and the problem of video segments being too trivial and incoherent in the existing technology is solved, the consistency and semantic integrity of video segments are achieved, and the construction efficiency of video materials is improved.

CN120302123APending Publication Date: 2025-07-11SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510596798.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing video division methods mainly divide videos based on picture changes, resulting in video segments being too trivial, lacking content coherence, and possible audio screen truncation.

Method used

By obtaining the video picture and semantic division information, mapping is performed based on the matching degree between semantic segments and picture segments, the picture partition information is corrected, and the content consistency and semantic integrity of the video segments are ensured.

Benefits of technology

The content of video clips is consistent and semantic intact, avoids audio screen truncation, and improves the construction efficiency and quality of video materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302123A_ABST
    Figure CN120302123A_ABST
Patent Text Reader

Abstract

The invention provides a video division method and device, equipment, a readable storage medium and a program product. The video division method comprises the following steps: acquiring picture division information of a video, wherein the picture division information comprises a plurality of picture segments; obtaining semantic division information of the video, wherein the semantic division information comprises a plurality of semantic segments and description objects corresponding to the semantic segments; based on each semantic segment in the semantic division information and each picture segment in the picture division information, determining a description object matched with the picture segment for each picture segment; and based on the description object matched with the picture segment, carrying out correction processing on the picture division information of the video to obtain a plurality of video segments contained in the video. By adopting the embodiment of the invention, the image division information can be corrected by utilizing the semantic division information of the video, so that the corrected video is coherent in content and complete in semantics, and the problem that video division is too trivial can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and particularly to a video segmentation method, apparatus, device, readable storage medium, and program product. Background Art

[0002] When constructing a video material library, a large number of videos with single content are usually required as video materials. To save costs, videos containing multiple contents can usually be processed by video segmentation to obtain video segments with single content. Video segmentation refers to the process of dividing an original video to obtain one or more video segments.

[0003] However, existing video segmentation methods usually segment videos based on picture changes, which results in the obtained video segments being too trivial and the content of a single video segment lacking coherence and thus not being able to be used as video materials. Summary of the Invention

[0004] Embodiments of the present application provide a video segmentation method, apparatus, device, readable storage medium, and program product, which can use the semantic segmentation information of a video to correct the picture segmentation information of the video, so that the corrected video content is coherent and the semantics are complete, effectively solving the problem that video segmentation is too trivial.

[0005] On the one hand, embodiments of the present application provide a video segmentation method, which includes:

[0006] Obtain the picture segmentation information of the video, where the picture segmentation information is obtained by segmenting the picture content in the video, and the picture segmentation information includes multiple picture segments;

[0007] Obtain the semantic segmentation information of the video, where the semantic segmentation information is obtained by semantically segmenting the text content in the video, and the semantic segmentation information includes multiple semantic segments and the description objects corresponding to the semantic segments;

[0008] Based on each semantic segment in the semantic segmentation information and each picture segment in the picture segmentation information, determine a description object matching the picture segment for each picture segment;

[0009] Based on the description object matching the picture segment, perform a correction process on the picture segmentation information of the video to obtain multiple video segments included in the video.

[0010] Correspondingly, embodiments of the present application provide a video segmentation apparatus, which includes:

[0011] An acquisition unit, configured to acquire the screen division information of a video, where the screen division information is obtained by dividing the screen content in the video, and the screen division information includes a plurality of screen segments;

[0012] The acquisition unit is further configured to acquire the semantic division information of the video, where the semantic division information is obtained by semantically dividing the text content in the video, and the semantic division information includes a plurality of semantic segments and the description objects corresponding to the semantic segments;

[0013] A determination unit, configured to determine, for each of the screen segments, a description object matching the screen segment based on each of the semantic segments in the semantic division information and each of the screen segments in the screen division information;

[0014] A correction unit, configured to perform a correction process on the screen division information of the video based on the description object matching the screen segment, so as to obtain a plurality of video segments included in the video.

[0015] Correspondingly, an embodiment of the present application provides a computer device, which includes:

[0016] A processor, adapted to implement a computer program;

[0017] A computer-readable storage medium, storing a computer program, where the computer program is adapted to be loaded and executed by the processor to perform the above video division method.

[0018] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium, storing a computer program, which, when read and executed by the processor of a computer device, causes the computer device to perform the above video division method.

[0019] Correspondingly, an embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the above video division method.

[0020] In this application, based on each semantic segment in the semantic segmentation information of the video and each frame segment in the frame segmentation information of the video, a description object matching the frame segment can be determined for each frame segment; based on the description object matching the frame segment, the frame segmentation information of the video can be corrected to obtain multiple video segments included in the video, so that the content of the corrected multiple video segments is coherent and the semantics is complete, thereby effectively solving the problem of overly fragmented video segmentation; in addition, since this application corrects the frame segments based on semantic segments, the corrected frame segments (or videos) are not likely to have the situation of audio-visual truncation. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic diagram of the system architecture of a video segmentation system provided by an embodiment of the present application;

[0023] Figure 2 It is a schematic flowchart of a video segmentation method provided by an embodiment of the present application;

[0024] Figure 3 It is a schematic diagram of a matching degree determination method provided by an embodiment of the present application;

[0025] Figure 4 It is a schematic diagram of another matching degree determination method provided by an embodiment of the present application;

[0026] Figure 5 It is a schematic flowchart of another video segmentation method provided by an embodiment of the present application;

[0027] Figure 6 It is a schematic diagram of a video segmentation method provided by an embodiment of the present application;

[0028] Figure 7 It is a block diagram of the structure of a video segmentation device provided by an embodiment of the present application;

[0029] Figure 8 It is a block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed Embodiments

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0031] It should be noted that the descriptions such as "first" and "second" involved in the embodiments of the present application are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one of such features.

[0032] When constructing a video material library, in order to solve the cost, a method of video segmentation for an original video containing multiple contents can be adopted to obtain video materials containing a single content. The existing video segmentation methods usually segment the video into multiple picture segments according to the pictures of the video. However, this video segmentation method only considers the picture changes of the video and does not consider the semantic content of the video, which will result in the multiple picture segments obtained by segmentation being very trivial, and a single picture segment cannot contain complete semantics. Moreover, there may also be a situation of audio truncation.

[0033] The video segmentation method in the embodiments of the present application is applied to a video segmentation device. The video segmentation device is set at the segmentation end, and the segmentation end can run in a video segmentation device. One or more processors, memories, and one or more application programs are provided in the video segmentation device. One or more application programs are stored in the memory and are configured to be executed by the processor to implement the video segmentation method. Among them, the video segmentation device can be a server, such as a physical server, a cloud server, etc.

[0034] The embodiments of the present application provide a video segmentation system. Next, the architecture of the video segmentation system provided in the embodiments of the present application will be introduced in conjunction with the accompanying drawings.

[0035] Please refer to Figure 1 , which is a schematic diagram of the system architecture of a video segmentation system provided in the embodiments of the present application. The video segmentation system can include at least one video segmentation device 100 (a video segmentation device is integrated in any video segmentation device 100) and at least one video usage device 200. A computer-readable storage medium corresponding to the video segmentation method runs in the video segmentation device 100 to execute the video segmentation method. Any one video usage device 200 communicates with at least one video segmentation device 100, and the video usage device 200 is also used for storing and using picture segments, etc.

[0036] Any video partitioning device 100 refers to a server or terminal device that executes the video partitioning method provided in the embodiments of the present application, corrects the video's screen partitioning information, and obtains multiple video segments included in the video.

[0037] Any video usage device 200 refers to a server or terminal device that receives multiple video segments sent by the video partitioning device 100 and uses the multiple video segments for processing such as model training and database construction.

[0038] In this embodiment, the video partitioning device 100 can obtain the video to be partitioned, and the video to be partitioned contains text content. The video partitioning device 100 can obtain the screen partitioning information of the video. The screen partitioning information is obtained by the video partitioning device 100 partitioning the screen content in the video, and the screen partitioning information can include multiple screen segments. The video partitioning device 100 can also obtain the semantic partitioning information of the video. The semantic partitioning information is obtained by the video partitioning device 100 semantically partitioning the text content in the video, and the semantic partitioning information can include multiple semantic segments and the description objects corresponding to the semantic segments.

[0039] The video partitioning device 100 can determine, based on each semantic segment in the semantic partitioning information and each screen segment in the screen partitioning information, a description object that matches the screen segment for each screen segment, and based on the description object that matches the screen segment, correct the screen partitioning information of the video to obtain multiple video segments included in the video. The content included in the multiple video segments obtained by this correction is coherent and semantically complete.

[0040] The video partitioning device 100 can send the multiple video segments obtained by correction to the video usage device 200. The video usage device 200 can use the multiple video segments obtained by correction to build a video material library, or can also use the multiple video segments obtained by correction to train a model. Through the video partitioning method provided in the embodiments of the present application, the screen partitioning information of the video can be corrected using the semantic partitioning information of the video, so that the multiple video segments obtained by correction have coherent content, are semantically complete, and are not likely to have audio truncation.

[0041] Both the video partitioning device 100 and the video using device 200 in the video partitioning system can be servers. At this time, the video partitioning device 100 and the video using device 200 can be independent physical servers, or a server cluster or distributed system composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0042] Any video partitioning device 100 and any video using device 200 can be directly or indirectly communicatively connected through wired or wireless communication means, and this application does not limit this.

[0043] It can be understood that Figure 1 At least one video partitioning device 100 or at least one video using device 200 in the shown video partitioning system does not constitute a limitation to the embodiments of the present invention. That is, the number of video partitioning devices, the types of video partitioning devices, the number of video using devices, the types of video using devices, or the number and types of devices included in each video partitioning device or video using device do not affect the overall implementation of the technical solutions in the embodiments of the present invention, and can all be regarded as equivalent replacements or derivatives of the technical solutions required to be protected in the embodiments of the present invention.

[0044] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a video partitioning method provided by an embodiment of this application. This video partitioning method can be implemented by the above-mentioned video partitioning device 100, or by the above-mentioned video using device 200, or by other devices that can implement this video partitioning method. The following takes the implementation of this video partitioning method by the above-mentioned video partitioning device 100 as an example for description. The process of the video partitioning method provided in the embodiments of this application includes but is not limited to:

[0045] S201. Obtain the picture partitioning information of the video. The picture partitioning information is obtained by partitioning the picture content in the video, and the picture partitioning information includes multiple picture segments.

[0046] In the embodiments of the present application, a video partitioning device may obtain the screen partitioning information of a video, which is obtained by partitioning the screen content in the video. The screen partitioning information may include multiple screen segments. Specifically, the video partitioning device may partition according to the change of the screen content in the video to obtain the screen partitioning information. For example: If a video has 60 frames (the frames per second (FPS) of the video is 30, that is, 30 frames per second), and the screen content before and after the 30th frame of the video changes suddenly, then the video partitioning device may perform video partitioning at the 30th frame of the video, and partition the video into two screen segments. The first screen segment includes the 1st frame to the 30th frame (that is, the starting frame is the 1st frame, and the ending frame is the 30th frame), and the second screen segment includes the 31st frame to the 60th frame (that is, the starting frame is the 31st frame, and the ending frame is the 60th frame). In an actual application scenario, there may be multiple changes in the screen content in a video describing the same object (for example: a video introducing an umbrella may include images of a rainy environment, the umbrella surface, the umbrella ribs, the raw material composition of the umbrella, the production of the umbrella, etc.), and the audio and the screen in the video may not be completely corresponding (for example: the audio of the video is describing the raw material composition of the umbrella, while the screen of the video is the production screen of the umbrella). Therefore, the multiple screen segments obtained only based on the screen content are relatively trivial, with poor content coherence, and there may also be a situation of audio truncation. Through the method provided in the embodiments of the present application, the video can be initially partitioned in terms of the screen, so as to facilitate subsequent correction of the screen partitioning information according to the semantic partitioning information of the video, thereby ensuring the accuracy and coherence of the corrected screen partitioning information.

[0047] Optionally, the screen partitioning information may further include the time range of each screen segment. For example: Assume that the video partitioning device performs video partitioning at the 30th frame of the video and partitions the video into two screen segments. Then the time range of the first screen segment is from the 1st frame to the 30th frame (that is, the starting frame is the 1st frame, and the ending frame is the 30th frame), and the time range of the second screen segment is from the 31st frame to the 60th frame (that is, the starting frame is the 31st frame, and the ending frame is the 60th frame).

[0048] S202. Obtain the semantic partitioning information of the video. The semantic partitioning information is obtained by semantically partitioning the text content in the video, and the semantic partitioning information includes multiple semantic segments and the description objects corresponding to the semantic segments.

[0049] In the embodiments of the present application, the video includes text content for describing multiple description objects, that is, the video may include text content for describing objects of multiple categories. The description objects may include product categories (i.e., categories), domain categories, or scene categories, etc. For example: If the video is a video introducing three different categories of products (for example: the three products are an umbrella, a raincoat, and rain boots respectively), then the video may include text content for describing these three products. The video partitioning device can obtain the semantic partitioning information of the video, which is obtained by semantically partitioning the text content in the video according to the category, and the semantic partitioning information includes multiple semantic segments. One semantic segment usually corresponds to an object of one category. For example: The video has 180 frames (the FPS of the video is 30, that is, 30 frames per second), and the video includes text content for describing two different category objects (such as "An umbrella can block strong winds, and a raincoat is light and waterproof"), then the video partitioning device can semantically partition the text content in the video according to the category to obtain the semantic partitioning information, and the semantic partitioning information may include two semantic segments. The first semantic segment (content: "An umbrella can block strong winds") includes frames 20 to 110 (that is, the starting frame is frame 20 and the ending frame is frame 110), and the second semantic segment (content: "A raincoat is light and waterproof") includes frames 111 to 180 (that is, the starting frame is frame 111 and the ending frame is frame 180). Another example: If the video is a video introducing products in the clothing domain category, and this video introduces a coat (domain category: outerwear), a pair of pants (domain category: trousers), and a dress (domain category: dress) respectively, then the video may include text content for describing these three products. The video partitioning device can obtain the semantic partitioning information of the video, which is obtained by semantically partitioning the text content in the video according to the domain category, and the semantic partitioning information includes multiple semantic segments. One semantic segment usually corresponds to an object of one domain category. For example: The video has 150 frames (the FPS of the video is 30, that is, 30 frames per second), and the video includes text content for describing three different domain category objects, then the video partitioning device can semantically partition the text content in the video according to the domain category to obtain the semantic partitioning information, and the semantic partitioning information may include three semantic segments. The first semantic segment (content: "Outdoor windproof and warm windbreaker") includes frames 1 to 50 (that is, the starting frame is frame 1 and the ending frame is frame 50), the second semantic segment (content: "Work pants with cuffed legs") includes frames 51 to 80 (that is, the starting frame is frame 51 and the ending frame is frame 80), and the third semantic segment (content: "Printed vacation-style dress") includes frames 81 to 150 (that is, the starting frame is frame 81 and the ending frame is frame 150).

[0050] Through the method provided by the embodiments of the present application, the text content in the video can be semantically segmented to obtain semantically complete semantic segmentation information, which is conducive to subsequently correcting the picture segmentation information according to semantic segmentation and picture segmentation, so as to obtain multiple video segments with coherent content and complete semantics.

[0051] Optionally, the video and / or the obtained video segments may have description objects, where the video may describe the total object, and the obtained video segments may describe the sub-objects under the total object. For example, the total object may be the total category, and the obtained video segments may describe the sub-categories under the total category; another example is that the total object may be the total concept, and the obtained video segments may describe the sub-concepts under the total concept; another example is that the total object may be the total time interval (such as one day), and the obtained video segments may describe the sub-time intervals (such as morning, noon, and evening) within the total time interval.

[0052] Optionally, for any semantic segment included in the semantic segmentation information, there may be no corresponding description object for this semantic segment. Taking the category as the description object as an example, the video is a video introducing three commodities, and the categories of two of the commodities can be determined, for example, umbrellas and raincoats respectively, but the category of the other commodity cannot be determined. Then the category of this commodity is identified as "other", so the semantic segment for introducing this commodity has no corresponding category, that is, the semantic segment for introducing this commodity has no corresponding description object.

[0053] Among them, the text content in the video may include at least one of the following: the text content obtained by performing audio-to-text processing on the audio content in the video; the text content in each video frame included in the video; the subtitles associated with the video.

[0054] Optionally, the semantic segmentation information may further include the time range of each semantic segment. For example: Suppose the video segmentation device divides the video into three semantic segments. The first semantic segment includes frames 1 to 50, the second semantic segment includes frames 51 to 80, and the third semantic segment includes frames 81 to 150. Then the time range of the first semantic segment is from frame 1 to frame 50 (that is, the starting frame is frame 1 and the ending frame is frame 50), the time range of the second semantic segment is from frame 51 to frame 80 (that is, the starting frame is frame 51 and the ending frame is frame 80), and the time range of the third semantic segment is from frame 81 to frame 150 (that is, the starting frame is frame 81 and the ending frame is frame 150).

[0055] S203. Based on each semantic segment in the semantic segmentation information and each picture segment in the picture segmentation information, determine a description object that matches the picture segment for each picture segment.

[0056] Optionally, the screen division information may further include the time range of each screen segment, and the semantic division information may further include the time range of each semantic segment. Based on this, the video division device may perform a mapping process on the semantic segments and the screen segments based on the matching degree between the time ranges of the semantic segments in the semantic division information and the time ranges of the screen segments in the screen division information, and use the description object corresponding to the semantic segment mapped by each screen segment as the description object matching the screen segment.

[0057] In this embodiment, the video division device may perform a mapping process on the semantic segments and the screen segments based on the matching degree between the time ranges of the semantic segments in the semantic division information and the time ranges of the screen segments in the screen division information. Since the time range of the screen segment is usually short, while the time range of the semantic segment is usually long, after the mapping process, one semantic segment can usually be mapped to one or more screen segments. Through the method provided by the embodiments of the present application, the mapping relationship between the semantic segments and the screen segments can be determined based on the matching degree between the time ranges, and the description object matching the screen segment can be determined for each screen segment according to the result of the mapping process, which is beneficial to the subsequent correction process of the screen division information of the video based on the description object matching the screen segment, and ensures the semantic coherence of the multiple video segments obtained after correction.

[0058] In an embodiment, any semantic segment in the semantic division information is represented as the first semantic segment, and any screen segment in the screen division information is represented as the first screen segment. Then, based on the matching degree between the time ranges of the semantic segments in the semantic division information and the time ranges of the screen segments in the screen division information, the specific implementation manner of performing the mapping process on the semantic segments and the screen segments may be: obtaining the matching degree between the time range of the first semantic segment and the time range of the first screen segment; if the matching degree between the time range of the first semantic segment and the time range of the first screen segment is greater than or equal to the matching threshold, a mapping relationship is established between the first semantic segment and the first screen segment.

[0059] Specifically, any semantic segment in the semantic segmentation information can be represented as the first semantic segment, and any frame segment in the frame segmentation information can be represented as the first frame segment. When the video segmentation device performs mapping processing on the semantic segment and the frame segment, it can obtain the matching degree between the time range of the first semantic segment and the time range of the first frame segment. If the matching degree is greater than or equal to the matching threshold, the video segmentation device can establish a mapping relationship between the first semantic segment and the first frame segment. Conversely, if the matching degree is less than the matching threshold, the video segmentation device can refrain from establishing a mapping relationship between the first semantic segment and the first frame segment. In an actual application scenario, the matching threshold can be adaptively adjusted according to different application requirements. Through the method provided in the embodiments of this application, it is possible to determine whether to establish a mapping relationship between the first semantic segment and the first frame segment based on the matching degree and the matching threshold between the time range of the first semantic segment and the time range of the first frame segment, which can ensure the rationality and accuracy of establishing the mapping relationship, and thus facilitate the subsequent accurate correction processing of the frame segmentation information of the video according to the result of the mapping processing.

[0060] In one embodiment, the specific implementation method for obtaining the matching degree between the time range of the first semantic segment and the time range of the first frame segment can be as follows: obtain the time intersection between the time range of the first semantic segment and the time range of the first frame segment, where the time intersection includes the overlapping time points between the time range of the first semantic segment and the time range of the first frame segment; determine the matching degree between the time range of the first semantic segment and the time range of the first frame segment according to the number of overlapping time points included in the time intersection; where the matching degree between the time range of the first semantic segment and the time range of the first frame segment is: the number of overlapping time points in the time intersection, or the proportion of the number of overlapping time points in the time intersection; that the matching degree between the time range of the first semantic segment and the time range of the first frame segment is greater than or equal to the matching threshold includes: the number of overlapping time points in the time intersection is greater than the quantity threshold, or the proportion of the number of overlapping time points in the time intersection is greater than the proportion threshold.

[0061] Specifically, the video segmentation device can obtain the time intersection between the time range of the first semantic segment and the time range of the first picture segment. This time intersection includes the overlapping time points between the time range of the first semantic segment and the time range of the first picture segment. For example, if the time range of the first semantic segment is from frame 80 to frame 300, and the time range of the first picture segment is from frame 75 to frame 100, then the overlapping time points between the time range of the first semantic segment and the time range of the first picture segment are from frame 80 to frame 100. The video segmentation device can determine the matching degree between the time range of the first semantic segment and the time range of the first picture segment according to the number of overlapping time points included in the time intersection. This matching degree can be the number of overlapping time points in the time intersection, or the proportion of the number of overlapping time points in the time intersection. For example, if the number of overlapping time points in the time intersection is 20, then the matching degree between the time range of the first semantic segment and the time range of the first picture segment can also be 20. Another example, if the ratio of the number of overlapping time points in the time intersection to the number of time points included in the time range of the first picture segment is 20 / 25 = 0.8, then the matching degree between the time range of the first semantic segment and the time range of the first picture segment can also be 0.8.

[0062] Corresponding to the method for determining the matching degree, when the video segmentation device determines the relationship between the matching degree between the time range of the first semantic segment and the time range of the first picture segment and the matching threshold, that the matching degree between the time range of the first semantic segment and the time range of the first picture segment is greater than or equal to the matching threshold includes: the number of overlapping time points in the time intersection is greater than the quantity threshold, or the proportion of the number of overlapping time points in the time intersection is greater than the ratio threshold. That is, when the method for determining the matching degree is different, the matching threshold also needs to be adjusted adaptively. Through the method provided by the embodiments of the present application, the matching degree between each semantic segment and each picture segment can be accurately calculated, which is beneficial to accurately constructing the mapping relationship between the semantic segment and the picture segment according to the size relationship between the matching degree and the matching threshold, and further improving the accuracy of the correction process. In addition, the embodiments of the present application also provide a variety of different methods for determining the matching degree, which can be applied to a variety of different application scenarios.

[0063] Please refer to Figure 3 , which is a schematic diagram of a method for determining the matching degree provided by the embodiments of the present application. As Figure 3As shown, the semantic segmentation information of the video includes three semantic segments (Semantic Segment 1, Semantic Segment 2, and Semantic Segment 3) and the time range of each semantic segment. The frame segmentation information of the video includes five frame segments (Frame Segment A, Frame Segment B, Frame Segment C, Frame Segment D, and Frame Segment E) and the time range of each frame segment. The video segmentation device can obtain the time intersection between the time range of each semantic segment and each frame segment, and determine the matching degree between the time range of the corresponding semantic segment and the time range of the corresponding frame segment according to the number of overlapping time points included in the time intersection. The video segmentation device can also determine whether to construct a mapping relationship between the semantic segment and the frame segment according to the relationship between each matching degree and the matching threshold. Figure 3 Among them, the time intersection between the time range of Semantic Segment 1 and the time range of Frame Segment A is denoted as Time Intersection 1, and the time intersection between the time range of Semantic Segment 1 and the time range of Frame Segment B is denoted as Time Intersection 2, and so on. According to Time Intersection 1, the matching degree 1 between the time range of Semantic Segment 1 and the time range of Frame Segment A can be determined. Since this matching degree 1 is greater than the matching threshold, a mapping relationship between Semantic Segment 1 and Frame Segment A can be constructed. The matching degrees between the time range of Semantic Segment 1 and the time ranges of Frame Segment B - Frame Segment E are less than the matching threshold, so no mapping relationship is constructed. Using the same method, the time intersection between the time range of Semantic Segment 2 and the time range of Frame Segment B ( Figure 3 denoted as "the time intersection between Semantic Segment 2 and Frame Segment B" in it) can be determined, so that the matching degree between the time range of Semantic Segment 2 and the time range of Frame Segment B can be determined to be greater than the matching threshold, and a mapping relationship between Semantic Segment 2 and Frame Segment B can be constructed. According to the time intersection between the time range of Semantic Segment 2 and the time range of Frame Segment C ( Figure 3 denoted as "the time intersection between Semantic Segment 2 and Frame Segment C" in it), the matching degree between the time range of Semantic Segment 2 and the time range of Frame Segment C can be determined to be greater than the matching threshold, and a mapping relationship between Semantic Segment 2 and Frame Segment C can be constructed. According to the time intersection between the time range of Semantic Segment 3 and the time range of Frame Segment D ( Figure 3 denoted as "the time intersection between Semantic Segment 3 and Frame Segment D" in it), the matching degree between the time range of Semantic Segment 3 and the time range of Frame Segment D can be determined to be greater than the matching threshold, and a mapping relationship between Semantic Segment 3 and Frame Segment D can be constructed. According to the time intersection between the time range of Semantic Segment 3 and the time range of Frame Segment E ( Figure 3(expressed as "the time intersection between semantic segment 3 and video segment E"), it can be determined that the matching degree between the time range of semantic segment 3 and the time range of video segment E is greater than the matching threshold, and a mapping relationship between semantic segment 3 and video segment E can be constructed. Through the method provided by the embodiments of the present application, an accurate mapping relationship can be constructed based on the matching degrees between various semantic segments and video segments, which is conducive to subsequent correction processing of the video segmentation information according to the results of the mapping process.

[0064] In one embodiment, the specific implementation of obtaining the matching degree between the time range of the first semantic segment and the time range of the first video segment may be: constructing a timeline; marking corresponding video positions for each video segment on the timeline according to the time ranges of the video segments in the video segmentation information; marking corresponding semantic positions for each semantic segment on the timeline according to the time ranges of the semantic segments in the semantic segmentation information; and determining the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the timeline.

[0065] Specifically, the video segmentation device may construct a timeline, and the length of the timeline may be associated with the number of frames of the video. For example, if the total number of frames of the video is 100 frames, the length of the timeline may also be 100 units. The video segmentation device may mark corresponding video positions for each video segment on the timeline according to the time ranges of the video segments in the video segmentation information; and mark corresponding semantic positions for each semantic segment on the timeline according to the time ranges of the semantic segments in the semantic segmentation information. For example, if the time range of a video segment is from the 2nd frame to the 20th frame, the corresponding video position of this video segment may be marked at the corresponding position on the timeline, and if the time range of a semantic segment is from the 13th frame to the 50th frame, the corresponding semantic position of this semantic segment may be marked at the corresponding position on the timeline. If the first semantic segment is any semantic segment in the semantic segmentation information and the first video segment is any video segment in the video segmentation information, the video segmentation device may determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the timeline. Through the method provided by the embodiments of the present application, the time ranges of various semantic segments and video segments can be more intuitively displayed using the timeline, which is conducive to improving the determination efficiency of the matching degree between the time range of the semantic segment and the time range of the video segment.

[0066] In one embodiment, the specific implementation manner of determining the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment on the time axis and the first video position corresponding to the first video segment may be: obtaining the overlapping part between the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment; determining the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the overlapping part; wherein, if the overlapping part completely coincides with the time range of the first semantic segment, or the overlapping part completely coincides with the time range of the first video segment, then the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than the matching threshold; if the overlapping part does not completely coincide with the time range of the first semantic segment, and the overlapping part does not completely coincide with the time range of the first video segment, then the matching degree between the time range of the first semantic segment and the time range of the first video segment is: the time difference value between the start time of the first semantic segment and the start time of the first video segment, or the time difference value between the end time of the first semantic segment and the end time of the first video segment.

[0067] Specifically, the video segmentation device can obtain the overlapping part between the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment, and can determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the overlapping part. If the overlapping part completely coincides with the time range of the first semantic segment, or the overlapping part completely coincides with the time range of the first video segment, then it can be determined that the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than the matching threshold. For example: the time range of the first semantic segment is from the 20th frame to the 100th frame, the time range of the first video segment is from the 20th frame to the 60th frame, the overlapping part between the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment is the position from the 20th frame to the 60th frame on the time axis, and the overlapping part completely coincides with the time range of the first video segment. Therefore, it can be determined that the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than the matching threshold, that is, there is a mapping relationship between the first semantic segment and the first video segment.

[0068] If the overlapping part does not completely overlap with the time range of the first semantic segment, and the overlapping part also does not completely overlap with the time range of the first video segment, the matching degree between the time range of the first semantic segment and the time range of the first video segment can be: the time difference value between the start time of the first semantic segment and the start time of the first video segment, or the time difference value between the end time of the first semantic segment and the end time of the first video segment. This time difference value is determined according to a subtraction function of the time difference (for example: time difference value = -0.5 * time difference) (that is, the larger the time difference, the smaller the time difference value, and the smaller the matching degree). The time difference can be the difference between the start time of the first semantic segment and the start time of the first video segment, or the difference between the end time of the first semantic segment and the end time of the first video segment. It should be noted that this subtraction function can be adaptively adjusted according to different application requirements. Through the method provided by the embodiments of the present application, the matching degree between the time range of the first semantic segment and the time range of the first video segment can be intuitively and quickly determined by using the semantic position and video position on the time axis, which is beneficial to improving the overall processing efficiency of video segmentation.

[0069] In some cases, the above-mentioned first semantic segment can be any semantic segment in the semantic segmentation information, and the above-mentioned first video segment can be a video segment that has an overlapping part with the first semantic position corresponding to the first semantic segment on the time axis. At this time, the video segmentation device does not need to calculate the matching degree between each semantic segment in the semantic segmentation information and each video segment in the video segmentation information, but only needs to calculate the matching degree between the video segment that has an overlapping part with the first semantic position corresponding to the first semantic segment on the time axis and the first semantic segment. This method can effectively reduce the amount of calculation, save computing resources, and improve the efficiency of video segmentation processing at the same time.

[0070] Please refer to Figure 4 , which is a schematic diagram of another method for determining the matching degree provided by the embodiments of the present application. As Figure 4 shown, the semantic segmentation information of the video includes three semantic segments (semantic segment 1, semantic segment 2, and semantic segment 3 respectively) and the time range of each semantic segment, and the video segmentation information of the video includes five video segments (video segment A, video segment B, video segment C, video segment D, video segment E respectively) and the time range of each video segment. The video segmentation device can construct a time axis and mark the corresponding video position for each video segment on the time axis according to the time range of each video segment in the video segmentation information; mark the corresponding semantic position for each semantic segment on the time axis according to the time range of each semantic segment in the semantic segmentation information, and finally obtain as Figure 4The time axis shown. For semantic segment 1, if the picture segments that overlap with the semantic position corresponding to semantic segment 1 on the time axis are picture segment A and picture segment B, then the video segmentation device can determine the matching degree between the time range of semantic segment 1 and the time range of picture segment A, and the matching degree between the time range of semantic segment 1 and the time range of picture segment B based on the overlapping part. From the two matching degrees, it can be seen that a mapping relationship is established between semantic segment 1 and picture segment A, and no mapping relationship is established between semantic segment 1 and picture segment B. Therefore, the description object corresponding to semantic segment 1 can be used as the description object that matches picture segment A.

[0071] For semantic segment 2, if the picture segments that overlap with the semantic position corresponding to semantic segment 2 on the time axis are picture segment B and picture segment C, then the video segmentation device can determine the matching degree between the time range of semantic segment 2 and the time range of picture segment B, and the matching degree between the time range of semantic segment 2 and the time range of picture segment C based on the overlapping part. From the two matching degrees, it can be seen that a mapping relationship should be established between semantic segment 2 and picture segment B, and a mapping relationship should be established between semantic segment 2 and picture segment C. Therefore, the description object corresponding to semantic segment 2 can be used as the description object that matches picture segment B and picture segment C.

[0072] For semantic segment 3, if the picture segments that overlap with the semantic position corresponding to semantic segment 3 on the time axis are picture segment C, picture segment D, and picture segment E, then the video segmentation device can determine the matching degree between the time range of semantic segment 3 and the time range of picture segment C, the matching degree between the time range of semantic segment 3 and the time range of picture segment D, and the matching degree between the time range of semantic segment 3 and the time range of picture segment E based on the overlapping part. From the three matching degrees, it can be seen that no mapping relationship is established between semantic segment 3 and picture segment C, a mapping relationship should be established between semantic segment 3 and picture segment D, and a mapping relationship should be established between semantic segment 3 and picture segment E. Therefore, the description object corresponding to semantic segment 3 can be used as the description object that matches picture segment D and picture segment E.

[0073] Through the method provided by the embodiments of this application, the mapping relationship between semantic segments and picture segments can be determined intuitively and quickly, which is beneficial to improving the video segmentation efficiency and saving computing resources.

[0074] In one embodiment, when the overlapping part does not completely overlap with the time range of the first semantic segment and does not completely overlap with the time range of the first video segment, the matching degree between the time range of the first semantic segment and the time range of the first video segment is: the time difference value between the start time of the first semantic segment and the start time of the first video segment, or the time difference value between the end time of the first semantic segment and the end time of the first video segment. The time difference value is determined according to a decreasing function of the time difference, and the time difference can be the difference between the start time of the first semantic segment and the start time of the first video segment, or the difference between the end time of the first semantic segment and the end time of the first video segment. And there is a limit value for the time difference. If the time difference is greater than the limit value (for example: 10 frames), then it can be directly determined that the matching degree between the time range of the first semantic segment and the time range of the first video segment is less than the matching threshold (that is, no mapping relationship is established between the first semantic segment and the first video segment). For example: when the overlapping part does not completely overlap with the time range of the first semantic segment and does not completely overlap with the time range of the first video segment, if the difference between the end time of the first semantic segment and the end time of the first video segment is 20 frames (the limit value is 10 frames), then it can be determined that the matching degree between the time range of the first semantic segment and the time range of the first video segment is less than the matching threshold. Through the method provided by the embodiments of the present application, the matching degree between the time range of the semantic segment and the time range of the video segment can be directly determined by the magnitude of the time difference, which is beneficial to further improving the processing efficiency of video segmentation.

[0075] It should be noted that in some cases, there are video segments in the video segmentation information that are not mapped to any semantic segment. Such video segments may not contain any important semantic content. Therefore, such video segments can be discarded during the video segmentation process. Through the method provided by the embodiments of the present application, while realizing video semantic segmentation, redundant content or video segments that do not contain important semantic content in the video can be removed, realizing the streamlined processing of the video.

[0076] S204. Based on the description object that matches the video segment, correct the video segmentation information of the video to obtain multiple video segments included in the video.

[0077] Optionally, the video segmentation device may traverse multiple video segments included in the video segmentation information, merge at least one video sub-segment whose currently traversed matching description object is the same description object, to obtain a video segment. After the traversal is completed, multiple video segments included in the video are obtained. For example, the video segmentation information includes five video sub-segments, namely video sub-segment A, video sub-segment B, video sub-segment C, video sub-segment D, and video sub-segment E. The description object matching video sub-segment A is the description object corresponding to semantic segment 1. The description objects matching video sub-segment B and video sub-segment C are both the description objects corresponding to semantic segment 2. The description objects matching video sub-segment D and video sub-segment E are both the description objects corresponding to semantic segment 3. Therefore, video sub-segment A can be used as a video segment included in the video, video sub-segment B and video sub-segment C are merged to obtain another video segment included in the video, and video sub-segment D and video sub-segment E are merged to obtain yet another video segment included in the video.

[0078] In this embodiment, the video segmentation information of the video includes multiple video sub-segments, and these video sub-segments are obtained by segmenting the video according to the video content. Since only the video change situation is considered during the segmentation, the content in a single video sub-segment may be incoherent, the semantics may be incomplete, and there may be a situation of audio truncation. By using the method provided in this application, mapping processing can be performed on the semantic segments and the video sub-segments to obtain the result of the mapping processing. The video segmentation device may determine, according to the result of the mapping processing (i.e., the mapping relationship between each semantic segment and the video sub-segment), a description object matching the video sub-segment for each video sub-segment, and based on the description object matching the video sub-segment, perform correction processing on the video segmentation information of the video to obtain multiple video segments included in the video, that is, perform correction processing on multiple video sub-segments mapped to the same semantic segment, so that the content of the video segments obtained by the correction processing is coherent, the semantics is complete, and the situation of audio truncation is not likely to occur.

[0079] In an embodiment, any one semantic segment in the semantic segmentation information is represented as the first semantic segment; then the specific implementation manner of performing correction processing on the video segmentation information of the video may be: if the result of the mapping processing indicates that the first semantic segment is mapped to at least one video sub-segment, merge the at least one video sub-segment.

[0080] Specifically, since a semantic segment usually corresponds to a description object, multiple video segments mapped to the same semantic segment can be considered semantically related, that is, the description objects matched by multiple video segments mapped to the same semantic segment are the same, and they should be merged into one video segment to ensure the semantic coherence of the video segments. Therefore, if the result of the mapping process indicates that the first semantic segment is mapped to at least one video segment, the video partitioning device can merge at least one video segment in chronological order to obtain one video segment. The video segment obtained by this correction process has coherent content and complete semantics, and can correspond to an object of a category.

[0081] It should be noted that in this application, the mapping result between the semantic segment and the video segment can be used to correct the video partitioning information, and multiple video segments mapped to the same semantic segment are merged into one video segment. Since each video segment is obtained by partitioning the video content based on the video change, the video of each video segment has video integrity and usually there is no abnormal truncation of the video. Furthermore, it can be known that the video segment obtained by correction also has video integrity and usually there is no abnormal truncation of the video. Compared with the method of directly partitioning the video content based on the time range of the semantic segment, the method provided in this application can ensure the video integrity of the video segment obtained by correction, reduce the abnormal truncation of the video, and improve the video partitioning quality.

[0082] Through the method provided in the embodiments of this application, the video partitioning information and the semantic partitioning information of the video can be obtained, and based on the semantic partitioning information and the video partitioning information, a description object matched with the video segment can be determined for each video segment, and based on the description object matched with the video segment, the video partitioning information of the video is corrected to obtain multiple video segments included in the video, so that the video segments obtained by the correction process have coherent content, complete semantics, and are not prone to audio truncation; this application provides two different methods to determine the matching degree between the time range of the semantic segment and the time range of the video segment, which can accurately determine the matching degree, and then accurately determine whether a mapping relationship is established between the semantic segment and the video segment, and can also selectively calculate the matching degree, thereby saving computing resources and improving the processing efficiency of video partitioning.

[0083] Please refer to Figure 5 , Figure 5Schematic flowchart of another video segmentation method provided by an embodiment of this application. This video segmentation method can be implemented by the above-mentioned video segmentation device 100, or by the above-mentioned video usage device 200, or by other devices that can implement this video segmentation method. The following takes the implementation of this video segmentation method by the above-mentioned video segmentation device 100 as an example for illustration. The process of the video segmentation method provided by the embodiment of this application includes but is not limited to:

[0084] S501. Obtain the screen segmentation information of the video. The screen segmentation information is obtained by segmenting the screen content in the video, and the screen segmentation information includes multiple screen segments.

[0085] In an embodiment of this application, the video segmentation device can obtain the screen segmentation information of the video. The screen segmentation information is obtained by segmenting the screen content in the video. The screen segmentation information may include multiple screen segments.

[0086] In one embodiment, the specific implementation manner of obtaining the screen segmentation information of the video may be: obtaining the video frame sequence of the video. The video frame sequence includes multiple video frames; performing feature recognition processing on each video frame in the video frame sequence to obtain the screen change label of each video frame. The screen change label indicates whether the video frame is a screen change frame; if the screen change label of any video frame in the video frame sequence indicates that any video frame is a screen change frame, then perform segmentation processing on the video according to the time position of any video frame in the video to obtain the screen segmentation information.

[0087] Specifically, the video segmentation device can obtain the video frame sequence of the video, and the video frame sequence includes multiple video frames. Feature video processing can be performed on each video frame in the video frame sequence to obtain the screen change label of each video frame. The screen change label indicates whether the video frame is a screen change frame. If the screen change label of any video frame in the video frame sequence indicates that the video frame is a screen change frame, then the video segmentation device can perform segmentation processing on the video according to the time position of the video frame in the video to obtain multiple screen segments and the time range of each screen segment. For example: the video frame sequence contains 120 frames, and the screen change labels of the 20th frame, 50th frame, and 90th frame indicate that the video frames are screen change frames. Then the video segmentation device can segment the video to obtain 4 screen segments, and their time ranges are from the 1st frame to the 20th frame, from the 21st frame to the 50th frame, from the 51st frame to the 90th frame, and from the 91st frame to the 120th frame.

[0088] It should be noted that in some cases, the video segmentation device can also directly use video segmentation technologies or video segmentation models such as the shot transition detection neural network (TransNetV2) and the frame-based shot segmentation technology (Autoshot) to segment the video and obtain the video's frame segmentation information. The TransNetV2 method utilizes deep learning technology and performs well in the shot segmentation task. The core of TransNet V2 lies in its novel network structure, which includes 6 deep dilated convolutional (DDCNN) modules, as well as a processing method that combines histogram features and learnable similarity features, effectively reducing the risk of overfitting.

[0089] Optionally, the frame segmentation information may also include the time range of each frame segment. For example: Suppose the video segmentation device performs video segmentation at the 30th frame of the video and divides the video into two frame segments. Then, the time range of the first frame segment is from the 1st frame to the 30th frame (i.e., the starting frame is the 1st frame and the ending frame is the 30th frame), and the time range of the second frame segment is from the 31st frame to the 60th frame (i.e., the starting frame is the 31st frame and the ending frame is the 60th frame).

[0090] S502. Perform semantic recognition processing on the text content in the video to obtain the semantic text of the video and the timestamps of each word in the semantic text.

[0091] In the embodiments of the present application, the video contains audio content for describing multiple categories. In order to ensure the content coherence of the video segmentation result, the video segmentation device can perform semantic recognition processing on the text content in the video to obtain the semantic text of the video and the timestamps of each word in the semantic text. The text content in the video may include at least one of the following: the text content obtained by performing audio-to-text processing on the audio content in the video; the text content in each video frame included in the video; the subtitles associated with the video. In some cases, the video segmentation device can utilize the Automatic Speech Recognition (ASR) technology, such as the whisper model, to convert the spoken words in the video into text and record the timestamp of each word. The whisper model is an open-source automatic speech recognition model that supports functions such as multilingual speech recognition, speech translation, and language recognition; this model adopts a neural network architecture (Transformer) based on the self-attention mechanism, can handle multiple languages and tasks, and has high robustness and accuracy.

[0092] S503. Perform object information extraction processing on the content description information of the video to obtain video content object information.

[0093] In the embodiments of the present application, a database may store videos and content description information of the videos. The content description information may include the title, type tags, introduction text, review text, etc. of the videos. The content description information is used to indicate the general content of the videos, and the content description information may include the description objects to which the content in the videos belongs. The description objects may include product categories (i.e., categories), domain categories, or scenario categories, etc. For example: If a video introduces three different children's products (including children's chopsticks, children's laundry detergent, and children's spoons), the content description information of the video may be "Share some children's products". In some cases, the content description information may be input by the video publisher when publishing the video, or determined by the video detection object when detecting the video. The video partitioning device may perform object information extraction processing on the content description information of the video to obtain video content object information. The video content object information indicates the description objects to which the video content belongs. Through the method provided by the embodiments of the present application, the video content object information can be determined by using the content description information of the video, which is beneficial to subsequent semantic text partitioning of the video based on the video content object information. At the same time, this process can be completed by the video partitioning device without manual intervention, which can effectively save human resources.

[0094] S504. Perform semantic partitioning processing on the semantic text based on the video content object information to obtain multiple semantic segments, and determine the time range of each semantic segment according to the words included in each semantic segment and the timestamps of each word.

[0095] In the embodiments of the present application, the video partitioning device may perform semantic partitioning processing on the semantic text based on the video content object information to obtain multiple semantic segments, and may determine the time range of each semantic segment according to the words included in each semantic segment and the timestamps of each word. Each semantic segment and the time range of each semantic segment constitute the semantic partitioning information of the video. In some cases, the video partitioning device uses a large language model (LLM) to perform semantic partitioning processing on the semantic text of the video. Specifically, the video partitioning device may generate partitioning prompt words according to the video content category information, and input the semantic text of the video and the partitioning prompt words into the LLM for semantic partitioning processing to obtain multiple semantic segments. For example: If the video content object information is "baby products", the partitioning prompt words may be "Given a description text about baby products. The description text will contain one or more different baby products. Please divide the text into different text blocks according to the meaning of the text. Each text block describes a separate baby product description object. Note that the introductions at the beginning and end that are not related to baby products are divided into independent text blocks and labeled as other. The return format is a jsonarray object, where each item is a dictionary, in the format of <key> : <value>, replace with the divided content <key>, replace with the category name <value>。Directly return the content and value without adding unnecessary details. Do not change the description text of the input. Only make divisions.\nDescription text:\n\n”. It should be noted that when the video content category information and application requirements are different, the video segmentation device can generate different segmentation prompt words according to different video content object information and application requirements, and then different semantic segments can be obtained. Through the method provided in the embodiments of the present application, the semantic text of the video can be semantically segmented based on the video content object information to obtain semantic segments, which can effectively ensure the semantic integrity of the semantic segments and make one semantic segment correspond to one description object, thereby facilitating subsequent correction of the video segmentation information according to the mapping relationship between the semantic segments and the video segments.

[0096] S505. Perform mapping processing on the semantic segments and the video segments based on the matching degree between the time ranges of the semantic segments and the time ranges of the video segments in the video segmentation information.

[0097] In the embodiments of the present application, the video segmentation device can perform mapping processing on the semantic segments and the video segments based on the matching degree between the time ranges of the semantic segments in the semantic segmentation information and the time ranges of the video segments in the video segmentation information. Through the method provided in the embodiments of the present application, the mapping relationship between the semantic segments and the video segments can be determined according to the time ranges of the semantic segments and the video segments, and then multiple video segments that are semantically related can be determined, which is beneficial to subsequent correction of the video segmentation information according to the result of the mapping processing, so as to obtain video segments with complete semantics and coherent content.

[0098] S506. Use the description object corresponding to the semantic segment mapped by each video segment as the description object that matches the video segment.

[0099] In the embodiments of the present application, the video segmentation device can use the description object corresponding to the semantic segment mapped by each video segment as the description object that matches the video segment according to the result of the mapping processing (i.e., the mapping relationship between the semantic segments and the video segments).

[0100] S507. Perform correction processing on the video segmentation information of the video based on the description object that matches the video segment to obtain multiple video segments included in the video.

[0101] Optionally, the video segmentation device may merge the picture segments with the same description object as the matching description object to obtain a video segment included in the video, that is, it may perform correction processing on multiple picture segments mapped to the same semantic segment, so that the content of the video segment obtained by the correction processing is coherent, the semantics is complete, and the situation of audio truncation is not likely to occur. Since the same semantic segment corresponds to one description object, the video segment obtained by the correction processing also corresponds to one description object, that is, the content of the video segment obtained by the correction processing is single and can be used as video material. Through the method provided by the embodiments of the present application, video material can be constructed quickly, which is beneficial to improving the construction efficiency of the video material library.

[0102] In some scenarios, the video segmentation device may store the video segments obtained by the correction processing. The video segments may include the merged picture segments and the semantic segments corresponding to the picture segments, so that when performing model training later, the picture segments and semantic segments with single content can be directly used.

[0103] Please refer to Figure 6 , which is a schematic diagram of a video segmentation method provided by the embodiments of the present application. The video segmentation method provided by the present application mainly includes three stages: a segmentation stage, a fusion alignment stage, and a merging stage. Specifically, in the segmentation stage, the video segmentation device may segment the video based on the picture content to obtain the picture segmentation information of the video. The picture segmentation information includes multiple picture segments ( Figure 6 denoted as V1, V2,..., Vn in ). In the segmentation stage, the video segmentation device may also convert the text content in the video into semantic text, and input the segmentation prompt words and the semantic text into a large language model for semantic segmentation processing to obtain the semantic segmentation information of the video. The semantic segmentation information includes multiple semantic segments ( Figure 6 denoted as P1, P2,..., Pm in ) and the description objects corresponding to each semantic segment. In the fusion alignment stage, the video segmentation device may determine, based on each semantic segment in the semantic segmentation information and each picture segment in the picture segmentation information, a description object that matches the picture segment for each said picture segment. In the merging stage, the video segmentation device may perform correction processing on the picture segmentation information of the video based on the description object that matches the picture segment, to obtain multiple video segments included in the video, that is, merge the picture segments with the same description object as the matching description object to obtain a video segment included in the video. Through the method provided by the embodiments of the present application, the problems that may occur during video segmentation, such as the picture segments being too trivial, having no coherence, and single picture segments having no actual meaning, can be effectively solved. It can also solve the problem of incorrect segmentation caused by the semantic continuity but picture discontinuity between two picture segments, effectively ensuring the content coherence and semantic integrity of the corrected video segments.

[0104] By using the method provided in the embodiments of the present application, the screen division information and semantic division information of a video can be obtained. Based on the semantic division information and the screen division information, a description object matching the screen segment can be determined for each screen segment, and based on the description object matching the screen segment, the screen division information of the video can be corrected to obtain multiple video segments included in the video, so that the content of the video segments obtained by the correction process is coherent, the semantics is complete, and the situation of audio truncation is not likely to occur. The present application provides two different methods to determine the matching degree between the time range of semantic segmentation and the time range of screen segmentation, which can accurately determine the matching degree, and then accurately determine whether a mapping relationship is established between semantic segmentation and screen segmentation. It is also possible to selectively calculate the matching degree, thereby saving computing resources and improving the processing efficiency of video division. The video segments obtained by using the video division method provided in the present application can be used as video materials, that is, the method provided in the present application can effectively improve the construction efficiency of video materials and save construction costs.

[0105] Please refer to Figure 7 , Figure 7 which is a structural block diagram of a video division device provided in an embodiment of the present application. The video division device can be set in the computer device provided in the embodiment of the present application, and the computer device can be the video division device 100 in the video division system shown above Figure 1 . Figure 7 The video division device shown above can be a computer program running in a computer device. The video division device can be used to execute Figure 2 or Figure 5 some or all of the steps in the video division method embodiments shown above. Please refer to Figure 7 , and the video division device can include the following units:

[0106] An obtaining unit 701, configured to obtain the screen division information of the video. The screen division information is obtained by dividing the screen content in the video, and the screen division information includes multiple screen segments;

[0107] The obtaining unit 701 is further configured to obtain the semantic division information of the video. The semantic division information is obtained by semantically dividing the text content in the video, and the semantic division information includes multiple semantic segments and description objects corresponding to the semantic segments;

[0108] A determining unit 702, configured to determine, based on each semantic segment in the semantic division information and each screen segment in the screen division information, a description object matching the screen segment for each screen segment;

[0109] A correction unit 703, configured to correct the video frame division information based on the description object matching the video segment, so as to obtain multiple video segments included in the video.

[0110] In an embodiment, the frame division information further includes the time range of each video segment, and the semantic division information further includes the time range of each semantic segment.

[0111] When the determination unit 702 determines a description object matching each video segment based on each semantic segment in the semantic division information and each video segment in the frame division information, it is specifically configured to perform the following steps:

[0112] Perform mapping processing on the semantic segments and the video segments based on the matching degree between the time ranges of each semantic segment in the semantic division information and the time ranges of each video segment in the frame division information.

[0113] Use the description object corresponding to the semantic segment mapped by each video segment as the description object matching the video segment.

[0114] In an embodiment, any one semantic segment in the semantic division information is denoted as a first semantic segment, and any one video segment in the frame division information is denoted as a first video segment. When the determination unit 702 performs mapping processing on the semantic segments and the video segments based on the matching degree between the time ranges of each semantic segment in the semantic division information and the time ranges of each video segment in the frame division information, it is specifically configured to perform the following steps:

[0115] Obtain the matching degree between the time range of the first semantic segment and the time range of the first video segment.

[0116] If the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than or equal to a matching threshold, establish a mapping relationship between the first semantic segment and the first video segment.

[0117] In an embodiment, when the determination unit 702 is configured to obtain the matching degree between the time range of the first semantic segment and the time range of the first video segment, it is specifically configured to perform the following steps:

[0118] Obtain the time intersection between the time range of the first semantic segment and the time range of the first video segment, where the time intersection includes the overlapping time points between the time range of the first semantic segment and the time range of the first video segment.

[0119] Determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the number of overlapping time points included in the time intersection;

[0120] Among them, the matching degree between the time range of the first semantic segment and the time range of the first video segment is: the number of overlapping time points in the time intersection, or the proportion of the number of overlapping time points in the time intersection; that the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than or equal to the matching threshold includes: the number of overlapping time points in the time intersection is greater than the quantity threshold, or the proportion of the number of overlapping time points in the time intersection is greater than the ratio threshold.

[0121] In one embodiment, when the determining unit 702 is used to obtain the matching degree between the time range of the first semantic segment and the time range of the first video segment, it is specifically used to perform the following steps:

[0122] Construct a timeline;

[0123] According to the time ranges of the respective video segments in the video division information, mark corresponding video positions on the timeline for each of the video segments;

[0124] According to the time ranges of the respective semantic segments in the semantic division information, mark corresponding semantic positions on the timeline for each of the semantic segments;

[0125] Determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the timeline.

[0126] In one embodiment, when the determining unit 702 is used to determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the timeline, it is specifically used to perform the following steps:

[0127] Obtain the overlapping part between the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment;

[0128] Determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the overlapping part;

[0129] Wherein, if the overlapping part completely coincides with the time range of the first semantic segment, or the overlapping part completely coincides with the time range of the first video segment, then the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than the matching threshold;

[0130] If the overlapping part does not completely coincide with the time range of the first semantic segment, and the overlapping part does not completely coincide with the time range of the first video segment, then the matching degree between the time range of the first semantic segment and the time range of the first video segment is: the time difference value between the start time of the first semantic segment and the start time of the first video segment, or the time difference value between the end time of the first semantic segment and the end time of the first video segment.

[0131] In one embodiment, when the correction unit 703 is used to correct the video's video segmentation information based on the description object that matches the video segment to obtain multiple video segments included in the video, it is specifically used to perform the following steps:

[0132] Traverse multiple video segments included in the video segmentation information, and merge at least one video segment whose currently traversed matching description object is the same description object to obtain a video segment;

[0133] After the traversal is completed, multiple video segments included in the video are obtained.

[0134] In one embodiment, when the acquisition unit 701 is used to acquire the video's video segmentation information, it is specifically used to perform the following steps:

[0135] Obtain the video frame sequence of the video, where the video frame sequence includes multiple video frames;

[0136] Perform feature recognition processing on each video frame in the video frame sequence to obtain the video change label of each video frame, where the video change label indicates whether the video frame is a video change frame;

[0137] If the video change label of any video frame in the video frame sequence indicates that the any video frame is a video change frame, then perform segmentation processing on the video according to the time position of the any video frame in the video to obtain the video segmentation information.

[0138] In one embodiment, when the acquisition unit 701 is used to acquire the semantic segmentation information of the video, it is specifically used to perform the following steps:

[0139] Perform semantic analysis processing on the text content in the video to obtain the semantic text of the video and the timestamps of each character in the semantic text;

[0140] Perform object information extraction processing on the content description information of the video to obtain video content object information;

[0141] Based on the video content object information, perform semantic segmentation processing on the semantic text to obtain multiple semantic segments and the description objects corresponding to the semantic segments;

[0142] According to the characters included in each semantic segment and the timestamps of each character, determine the time range of each semantic segment.

[0143] Through the video segmentation device provided by the embodiments of the present application, the screen segmentation information and semantic segmentation information of the video can be obtained, and based on the semantic segmentation information and the screen segmentation information, a description object matching the screen segment can be determined for each screen segment, and based on the description object matching the screen segment, the screen segmentation information of the video can be corrected to obtain multiple video segments included in the video, so that the content of the video segments obtained by the correction processing is coherent, the semantics is complete, and the situation of audio truncation is not likely to occur; The present application provides two different methods to determine the matching degree between the time range of the semantic segment and the time range of the screen segment, which can accurately determine the matching degree, and then accurately determine whether a mapping relationship is established between the semantic segment and the screen segment, and the matching degree can also be selectively calculated, thereby saving computing resources and improving the processing efficiency of video segmentation; The corrected screen segments obtained by using the video segmentation method provided by the present application can be used as video materials, that is, the method provided by the present application can effectively improve the construction efficiency of video materials and save construction costs.

[0144] Based on the above method and device embodiments, the embodiments of the present application provide a computer device. Please refer to Figure 8 , Figure 8 which is a structural block diagram of a computer device provided by an embodiment of the present application. Figure 8 The computer device shown can be the above-mentioned Figure 1 video segmentation device 100. Figure 8 The computer device shown at least includes a processor 801, an input interface 802, an output interface 803, and a computer-readable storage medium 804. Among them, the processor 801, the input interface 802, the output interface 803, and the computer-readable storage medium 804 can be connected through a bus or other means.

[0145] The computer-readable storage medium 804 can be stored in the memory of the computer device. The computer-readable storage medium 804 is used to store a computer program, and the computer program includes program instructions (such as Figure 8 (the program instructions 1, program instructions 2, …, program instructions n therein), the processor 801 is used to execute the computer program stored in the computer-readable storage medium 804. The processor 801 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing the computer program, specifically suitable for loading and executing the computer program to implement the corresponding method flow or corresponding function.

[0146] The embodiment of the present application also provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the computer device is stored in this storage space. And, a computer program suitable for being loaded and executed by the processor is also stored in this storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0147] In specific implementation, the processor 801 can load and execute the computer program stored in the computer-readable storage medium 804 to implement the relevant Figure 2 or Figure 5 corresponding steps in the video segmentation method shown. In specific implementation, the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 as follows:

[0148] Obtain the frame segmentation information of the video. The frame segmentation information is obtained by segmenting the frame content in the video, and the frame segmentation information includes multiple frame segments;

[0149] Obtain the semantic segmentation information of the video. The semantic segmentation information is obtained by semantically segmenting the text content in the video, and the semantic segmentation information includes multiple semantic segments and the description objects corresponding to the semantic segments;

[0150] Based on each semantic segment in the semantic segmentation information and each frame segment in the frame segmentation information, determine a description object that matches the frame segment for each frame segment;

[0151] Based on the description objects matching the screen segments, the screen division information of the video is corrected to obtain multiple video segments included in the video.

[0152] In one embodiment, the screen division information further includes the time range of each screen segment, and the semantic division information further includes the time range of each semantic segment;

[0153] When the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to determine, for each screen segment, a description object matching the screen segment based on each semantic segment in the semantic division information and each screen segment in the screen division information, it is specifically used to execute the following steps:

[0154] Based on the matching degree between the time ranges of each semantic segment in the semantic division information and the time ranges of each screen segment in the screen division information, perform a mapping process on the semantic segments and the screen segments;

[0155] Use the description object corresponding to the semantic segment mapped by each screen segment as the description object matching the screen segment.

[0156] In one embodiment, any one semantic segment in the semantic division information is denoted as the first semantic segment, and any one screen segment in the screen division information is denoted as the first screen segment; when the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to perform a mapping process on the semantic segments and the screen segments based on the matching degree between the time ranges of each semantic segment in the semantic division information and the time ranges of each screen segment in the screen division information, it is specifically used to execute the following steps:

[0157] Obtain the matching degree between the time range of the first semantic segment and the time range of the first screen segment;

[0158] If the matching degree between the time range of the first semantic segment and the time range of the first screen segment is greater than or equal to the matching threshold, establish a mapping relationship between the first semantic segment and the first screen segment.

[0159] In one embodiment, when the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to obtain the matching degree between the time range of the first semantic segment and the time range of the first screen segment, it is specifically used to execute the following steps:

[0160] Obtain the time intersection between the time range of the first semantic segment and the time range of the first video segment, where the time intersection includes the overlapping time points between the time range of the first semantic segment and the time range of the first video segment;

[0161] Determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the number of overlapping time points included in the time intersection;

[0162] Wherein, the matching degree between the time range of the first semantic segment and the time range of the first video segment is: the number of overlapping time points in the time intersection, or the proportion of the number of overlapping time points in the time intersection; that the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than or equal to the matching threshold includes: the number of overlapping time points in the time intersection is greater than the quantity threshold, or the proportion of the number of overlapping time points in the time intersection is greater than the proportion threshold.

[0163] In one embodiment, when the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to obtain the matching degree between the time range of the first semantic segment and the time range of the first video segment, it is specifically used to execute the following steps:

[0164] Construct a timeline;

[0165] According to the time ranges of each of the video segments in the video division information, mark the corresponding video positions for each of the video segments on the timeline;

[0166] According to the time ranges of each of the semantic segments in the semantic division information, mark the corresponding semantic positions for each of the semantic segments on the timeline;

[0167] Determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the timeline.

[0168] In one embodiment, when the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to determine the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the timeline, it is specifically used to execute the following steps:

[0169] Obtain the overlapping part between the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment;

[0170] Determine the matching degree between the time range of the first semantic segment and the time range of the first picture segment according to the overlapping part;

[0171] Wherein, if the overlapping part completely coincides with the time range of the first semantic segment, or the overlapping part completely coincides with the time range of the first picture segment, the matching degree between the time range of the first semantic segment and the time range of the first picture segment is greater than the matching threshold;

[0172] If the overlapping part does not completely coincide with the time range of the first semantic segment, and the overlapping part does not completely coincide with the time range of the first picture segment, the matching degree between the time range of the first semantic segment and the time range of the first picture segment is: the time difference value between the start time of the first semantic segment and the start time of the first picture segment, or the time difference value between the end time of the first semantic segment and the end time of the first picture segment.

[0173] In one embodiment, when the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to correct the picture division information of the video based on the description object matching the picture segment and obtain multiple video segments included in the video, it is specifically used to execute the following steps:

[0174] Traverse multiple picture segments included in the picture division information, and merge at least one picture segment whose currently traversed matching description object is the same description object to obtain a video segment;

[0175] After the traversal is completed, multiple video segments included in the video are obtained.

[0176] In one embodiment, when the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to obtain the picture division information of the video, it is specifically used to execute the following steps:

[0177] Obtain the video frame sequence of the video, and the video frame sequence includes multiple video frames;

[0178] Perform feature recognition processing on each video frame in the video frame sequence to obtain the picture change label of each video frame, and the picture change label indicates whether the video frame is a picture change frame;

[0179] If the picture change label of any video frame in the video frame sequence indicates that the any video frame is a picture change frame, perform division processing on the video according to the time position of the any video frame in the video to obtain the picture division information.

[0180] In one embodiment, when the computer program in the computer-readable storage medium 804 is loaded and executed by the processor 801 to obtain the semantic segmentation information of the video, it is specifically used to perform the following steps:

[0181] Perform semantic analysis processing on the text content in the video to obtain the semantic text of the video and the timestamps of each character in the semantic text;

[0182] Perform object information extraction processing on the content description information of the video to obtain video content object information;

[0183] Based on the video content object information, perform semantic segmentation processing on the semantic text to obtain multiple semantic segments and the description objects corresponding to the semantic segments;

[0184] According to the characters included in each semantic segment and the timestamps of each character, determine the time range of each semantic segment.

[0185] Through the computer device provided by the embodiments of the present application, the frame segmentation information and semantic segmentation information of the video can be obtained, and based on the semantic segmentation information and frame segmentation information, a description object matching the frame segment is determined for each frame segment, and based on the description object matching the frame segment, the frame segmentation information of the video is corrected to obtain multiple video segments included in the video, so that the content of the video segments obtained by the correction processing is coherent, the semantics is complete, and the situation of audio truncation is not likely to occur; the present application provides two different methods to determine the matching degree between the time range of the semantic segment and the time range of the frame segment, which can accurately determine the matching degree, and then accurately determine whether a mapping relationship is established between the semantic segment and the frame segment, and can also selectively calculate the matching degree, thereby saving computing resources and improving the processing efficiency of video segmentation; the corrected frame segments obtained by using the video segmentation method provided by the present application can be used as video materials, that is, the method provided by the present application can effectively improve the construction efficiency of video materials and save the construction cost.

[0186] The embodiments of the present application also provide a computer-readable storage medium. Computer instructions are stored in the computer-readable storage medium. When it runs on a computer device, the computer device is enabled to execute the steps in the method embodiments of the present application to implement the video segmentation method provided by the embodiments of the present application. The specific implementation manner can refer to the foregoing description and will not be elaborated herein.

[0187] An embodiment of the present application further provides a computer program product, which includes a computer program or computer instructions, and the computer program or computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions, so that the computer device executes the steps in the method embodiments of the present application to implement the video segmentation method provided by the embodiments of the present application. The specific implementation manner can refer to the foregoing description and will not be elaborated here.

[0188] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled artisans can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.

[0189] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0190] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technical person familiar with the technical field of the present application can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims described.< / value> < / key> < / value> < / key>

Claims

1. A video partitioning method, characterized in that, The method includes: Obtaining the frame division information of the video, where the frame division information is obtained by dividing the frame content in the video, and the frame division information includes multiple frame segments; Obtaining the semantic division information of the video, where the semantic division information is obtained by semantically dividing the text content in the video, and the semantic division information includes multiple semantic segments and the description objects corresponding to the semantic segments; Based on each of the semantic segments in the semantic division information and each of the frame segments in the frame division information, determining a description object that matches the frame segment for each frame segment; Based on the description object that matches the frame segment, performing a correction process on the frame division information of the video to obtain multiple video segments included in the video.

2. The method according to claim 1, characterized in that, The frame division information further includes the time range of each frame segment, and the semantic division information further includes the time range of each semantic segment; The determining a description object that matches the frame segment for each frame segment based on each of the semantic segments in the semantic division information and each of the frame segments in the frame division information includes: Performing a mapping process on the semantic segment and the frame segment based on the matching degree between the time ranges of each of the semantic segments in the semantic division information and the time ranges of each of the frame segments in the frame division information; Taking the description object corresponding to the semantic segment mapped by each frame segment as the description object that matches the frame segment.

3. The method according to claim 2, characterized in that Any one of the semantic segments in the semantic division information is denoted as the first semantic segment, and any one of the frame segments in the frame division information is denoted as the first frame segment; The performing a mapping process on the semantic segment and the frame segment based on the matching degree between the time ranges of each of the semantic segments in the semantic division information and the time ranges of each of the frame segments in the frame division information includes: Obtaining the matching degree between the time range of the first semantic segment and the time range of the first frame segment; If the matching degree between the time range of the first semantic segment and the time range of the first frame segment is greater than or equal to the matching threshold, establishing a mapping relationship between the first semantic segment and the first frame segment.

4. The method according to claim 3, wherein The obtaining the matching degree between the time range of the first semantic segment and the time range of the first frame segment includes: Obtaining the time intersection between the time range of the first semantic segment and the time range of the first frame segment, where the time intersection includes the overlapping time points between the time range of the first semantic segment and the time range of the first frame segment; Determining the matching degree between the time range of the first semantic segment and the time range of the first frame segment according to the number of overlapping time points included in the time intersection. The matching degree between the time range of the first semantic segment and the time range of the first video segment is: the number of overlapping time points in the time intersection, or the proportion of the number of overlapping time points in the time intersection; that the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than or equal to a matching threshold includes: the number of overlapping time points in the time intersection is greater than a quantity threshold, or the proportion of the number of overlapping time points in the time intersection is greater than a ratio threshold.

5. The method according to claim 3, wherein Obtaining the matching degree between the time range of the first semantic segment and the time range of the first video segment includes: Constructing a time axis; Marking corresponding video positions on the time axis for each of the video segments according to the time ranges of the video segments in the video division information; Marking corresponding semantic positions on the time axis for each of the semantic segments according to the time ranges of the semantic segments in the semantic division information; Determining the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the time axis.

6. The method according to claim 5, characterized in that The determining the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment on the time axis includes: Obtaining the overlapping part between the first semantic position corresponding to the first semantic segment and the first video position corresponding to the first video segment; Determining the matching degree between the time range of the first semantic segment and the time range of the first video segment according to the overlapping part; Wherein, if the overlapping part completely coincides with the time range of the first semantic segment, or the overlapping part completely coincides with the time range of the first video segment, then the matching degree between the time range of the first semantic segment and the time range of the first video segment is greater than the matching threshold; If the overlapping part does not completely coincide with the time range of the first semantic segment and the overlapping part does not completely coincide with the time range of the first video segment, then the matching degree between the time range of the first semantic segment and the time range of the first video segment is: the time difference value between the start time of the first semantic segment and the start time of the first video segment, or the time difference value between the end time of the first semantic segment and the end time of the first video segment.

7. The method according to any one of claims 1 to 6, characterized in that, Based on the description object matching the video segment, performing a correction process on the video division information of the video to obtain multiple video segments included in the video, including: Traversing multiple video segments included in the video division information, and merging at least one video segment whose currently traversed matching description object is the same description object to obtain a video segment; After the traversal is completed, multiple video segments included in the video are obtained.

8. The method according to any one of claims 1-6, characterized in that Obtaining the screen division information of the video includes: Obtaining a video frame sequence of the video, where the video frame sequence includes multiple video frames; Performing feature recognition processing on each of the video frames in the video frame sequence to obtain a screen change label for each of the video frames, where the screen change label indicates whether the video frame is a screen change frame; If the screen change label of any video frame in the video frame sequence indicates that the any video frame is a screen change frame, then perform division processing on the video according to the time position of the any video frame in the video to obtain screen division information.

9. The method according to any one of claims 2-6, characterized in that, Obtaining the semantic division information of the video includes: Performing semantic analysis processing on the text content in the video to obtain the semantic text of the video and the timestamps of each word in the semantic text; Performing object information extraction processing on the content description information of the video to obtain video content object information; Performing semantic division processing on the semantic text based on the video content object information to obtain multiple semantic segments and the description objects corresponding to the semantic segments; Determining the time range of each semantic segment according to the words included in each semantic segment and the timestamps of each word.

10. A video partitioning device, characterized in that Including: An obtaining unit, configured to obtain the screen division information of the video, where the screen division information is obtained by dividing the screen content in the video, and the screen division information includes multiple screen segments; The obtaining unit is further configured to obtain the semantic division information of the video, where the semantic division information is obtained by semantically dividing the text content in the video, and the semantic division information includes multiple semantic segments and the description objects corresponding to the semantic segments; A determining unit, configured to determine, based on each of the semantic segments in the semantic division information and each of the screen segments in the screen division information, a description object that matches the screen segment for each screen segment; A correcting unit, configured to perform correction processing on the screen division information of the video based on the description object that matches the screen segment to obtain multiple video segments included in the video.

11. A computer device, characterized in that, The computer device includes: A processor, adapted to implement a computer program; A computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and implemented by the processor to implement the video division method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer program stored in the computer-readable storage medium is adapted to be loaded and implemented by the processor to implement the video division method according to any one of claims 1-9.

13. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the video division method according to any one of claims 1-9.

Citation Information

Cited By

  • AI-driven audio and video content semantic segmentation method and system

    CN121640341A