Audio and video resource processing methods, devices, media and electronic equipment

By adjusting the segmentation position in audio and video resources based on text information, the problem of users having difficulty concentrating is solved, and the appropriate length and content integrity of the segmented segments are achieved, thus improving the user experience.

CN116361492BActive Publication Date: 2026-04-07BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

When listening to or watching audio or video resources of a certain length, users may find it difficult to maintain full concentration, leading to missed content and reducing the overall listening or watching experience.

Method used

By acquiring the text information of audio and video resources, the resource segmentation reference position is determined based on the segmentation start position and the segment reference duration, and adjustments are made as necessary to avoid segmentation in the middle of natural sentences or semantically coherent positions, and to segment into segments of appropriate length.

Benefits of technology

Ensure that the segmented segments are of appropriate length and complete in content to improve user attention and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116361492B_ABST
    Figure CN116361492B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, medium, and electronic device for processing audio and video resources. The method includes: acquiring text information corresponding to the audio and video resources to be processed; determining a resource segmentation reference position based on the segmentation start position and segment reference duration corresponding to the audio and video resources; if the determined resource segmentation reference position is any of the adjustment positions, adjusting the resource segmentation reference position to obtain a resource segmentation target position, wherein the adjustment position includes the position within a natural sentence and the position between two semantically coherent natural sentences; and segmenting the audio and video resources according to the resource segmentation target position to obtain segmented segments. In this way, the sentences at the end of the obtained segmented segments can express complete semantics and have a suitable duration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio and video processing, and more specifically, to an audio and video resource processing method, apparatus, medium, and electronic device. Background Technology

[0002] Today, audio and video are important sources of information and knowledge for people. However, when listening to or watching long audio and video resources, due to attention limitations, it is difficult for people to maintain full concentration on the content. As a result, people are likely to miss some content when acquiring information or learning knowledge through audio and video, reducing the effectiveness of listening to or watching audio and video. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] Firstly, this disclosure provides a method for processing audio and video resources, including:

[0005] Obtain the text information corresponding to the audio and video resources to be processed;

[0006] The resource segmentation reference position is determined based on the segmentation start position and segment reference duration corresponding to the audio and video resources;

[0007] If the resource segmentation reference position is determined to be any of the adjustment positions, the resource segmentation reference position is adjusted to obtain the resource segmentation target position, wherein the adjustment position includes the position in the sentence of a natural sentence and the position between two semantically coherent natural sentences;

[0008] The audio and video resources are segmented according to the target location of the resource segmentation to obtain segmented segments.

[0009] Secondly, this disclosure provides an audio and video resource processing apparatus, comprising:

[0010] The acquisition module is used to acquire the text information corresponding to the audio and video resources to be processed;

[0011] The first determining module is used to determine the resource segmentation reference position based on the segmentation start position and segment reference duration corresponding to the audio and video resources;

[0012] An adjustment module is used to adjust the resource segmentation reference position if it is determined that the resource segmentation reference position is any of the adjustment positions, so as to obtain the resource segmentation target position, wherein the adjustment position includes the position in the sentence of a natural sentence and the position between two semantically coherent natural sentences.

[0013] The segmentation module is used to segment the audio and video resources according to the resource segmentation target position to obtain segmented segments.

[0014] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described audio and video resource processing method.

[0015] Fourthly, this disclosure provides an electronic device, comprising:

[0016] A storage device on which computer programs are stored;

[0017] A processing device is used to execute the computer program in the storage device to implement the steps of the above-described audio and video resource processing method.

[0018] In the above technical solution, the resource segmentation reference position is determined based on the segmentation start position and the reference duration of the corresponding audio and video resources. Then, when the resource segmentation reference position is determined as an adjustment position, it is adjusted to obtain the resource segmentation target position. This way, when segmenting audio and video resources according to the target position, the duration of the resulting segmented segments is close to the reference duration, avoiding excessively long or short segments and ensuring that the duration of the segmented segments is within a certain range. This allows users to better focus their attention on the content of each segment of the audio and video resource, thereby improving overall attention to the audio and video resource. Furthermore, the adjustment position during segmentation includes positions within natural sentences and positions between two semantically coherent natural sentences. This ensures that the segmented segments do not end at positions within natural sentences or between two semantically coherent natural sentences. The ending sentences of the resulting segmented segments can express complete semantics, guaranteeing the integrity and independence of each segmented segment's content, further enhancing the user experience.

[0019] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0021] Figure 1 This is a flowchart of an audio and video resource processing method provided according to one embodiment of the present disclosure.

[0022] Figure 2 This is a flowchart of step S103 of an audio / video resource processing method provided according to an embodiment of the present disclosure.

[0023] Figure 3 This is a flowchart of an audio and video resource processing method provided according to one embodiment of the present disclosure.

[0024] Figures 4a-4d This is a schematic diagram of a directory used when displaying audio and video resources in an interactive interface, according to an embodiment of the present disclosure.

[0025] Figure 5 This is a block diagram of an audio / video resource processing apparatus provided according to one embodiment of the present disclosure.

[0026] Figure 6 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0036] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0037] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0038] Figure 1 This is a flowchart of an audio / video resource processing method according to one embodiment of this disclosure. Figure 1 As shown, the audio and video resource processing method includes steps S101 to S104.

[0039] In step S101, the text information corresponding to the audio and video resources to be processed is obtained.

[0040] The audio and video resource processing method disclosed herein can be used to process audio and video resources. The audio and video resources to be processed can be video resources or audio resources (e.g., audiobook resources).

[0041] The text information corresponding to the audio / video resources to be processed can be the text content contained in the audio / video (i.e., text content in natural language that humans can read, such as English text, Chinese text, etc.) and the timeline corresponding to the text content contained in the audio / video. For example, the audio / video resource to be processed can be a documentary, where the narration is its text content, and the timeline corresponding to the narration records the time when each sentence in the narration appears in the video (i.e., the time difference between each sentence and the opening sequence during playback). Another example is a course video, where the text content is its text content, and the timeline corresponding to the text content records the time when each sentence in the text appears in the video. Yet another example is an audiobook, where the text played by the audiobook is its text content, and the timeline corresponding to the text played by the audiobook records the time when each sentence in the audiobook appears in the audio.

[0042] When acquiring text information corresponding to audio and video resources to be processed, for audio resources (such as audiobooks), speech recognition can be performed on the audio resources to obtain the text content in the audio resources, and the timeline corresponding to the recognized text content can be obtained based on the time points corresponding to the speech in the audio resources; for video resources, the audio in the video resources can be extracted, speech recognition can be performed on the extracted audio to identify the text content in the extracted audio, and the timeline corresponding to the identified text content can be obtained.

[0043] In step S102, the resource segmentation reference position is determined based on the segmentation start position and segment reference duration corresponding to the audio and video resources.

[0044] When segmenting audio and video resources, we can obtain the segmented segments (i.e., the desired segments) and the remaining parts after segmentation. The segmentation start position can be used to indicate the position of the beginning of the segment to be segmented in the audio and video resources before segmentation (that is, the beginning position of the remaining parts obtained after the last segmentation of the audio and video resources).

[0045] In one implementation, the initial start of the audio / video resource (e.g., at 0 minutes and 0 seconds) can be used as the segmentation starting point. The segment reference duration is the desired length of the resulting segments. The segment reference duration can be preset. For example, when users learn using video courses, they are more likely to remember the first 3 minutes of the video content, resulting in better learning outcomes. If the video content is too short, such as 1 minute, the course video is too fragmented; conversely, if the video content is too long, such as 8 minutes, the learning effect may be reduced due to the excessive length. Therefore, the segment reference duration can be preset to 3 minutes, allowing the length of the segmented data to be adjusted within a 3-minute range. Similarly, for audio resources, such as audiobooks, users typically maintain focus for about 8 minutes; therefore, the segment reference duration for such audio resources can be preset to 8 minutes.

[0046] The resource segmentation reference position refers to the initial location determined when segmenting audio and video resources. In other words, initially, it is expected that the audio and video resources will be segmented at the resource segmentation reference position. However, in the actual segmentation process, due to limitations such as resource content and other rules, the resource segmentation reference position may be adjusted, and the final segmentation may not necessarily take place at the resource segmentation reference position.

[0047] The resource segmentation reference position can be determined based on the segmentation start position and the segment reference duration corresponding to the audio and video resources. For example, if the segmentation start position is 0 minutes and 0 seconds of the audio and video resources, and the segment reference duration is 3 minutes, the resource segmentation reference position can be determined to be 3 minutes and 0 seconds of the audio and video resources.

[0048] In step S103, if the resource segmentation reference position is determined to be any of the adjustment positions, the resource segmentation reference position is adjusted to obtain the resource segmentation target position. The adjustment position includes the position in the sentence of a natural sentence and the position between two semantically coherent natural sentences.

[0049] An adjustment position refers to a location within an audio or video resource that meets certain conditions and cannot be used for segmentation. In other words, if the resource segmentation reference position is an adjustment position, it needs to be adjusted to a location that is no longer an adjustment position.

[0050] A natural sentence is a complete sentence within the text content, such as a complete English sentence or a complete Chinese sentence. Punctuation marks in the text can be used to determine if a natural sentence is coherent. For example, in English, exclamation marks and periods can be used as markers for the end of a natural sentence. Two semantically coherent natural sentences are those with strong semantic relevance. Specifically, when determining whether two natural sentences are semantically coherent, a semantic relevance score can be calculated between them. If the semantic relevance score is greater than a preset threshold (a higher score indicates stronger semantic relevance), then the two natural sentences can be considered semantically coherent.

[0051] In one implementation, the adjustment position includes both the position within a sentence of a natural language statement and the position between two semantically coherent natural language statements. That is, when segmenting audio and video resources, if a position is either the position within a sentence of a natural language statement in the text content corresponding to the audio or video resource, or the position between two semantically coherent natural language statements in the text content corresponding to the audio or video resource, then the audio or video resource cannot be segmented at that position. Accordingly, in this step, the resource segmentation reference position can be adjusted to a position that is not an adjustment position, thus obtaining the target resource segmentation position.

[0052] As another example, if the resource splitting reference position is not the adjustment position, the resource splitting reference position can be directly determined as the resource splitting target position.

[0053] In step S104, the audio and video resources are segmented according to the resource segmentation target position to obtain segmented segments. That is, the audio and video resources can be segmented at the resource segmentation target position determined in step S103 to obtain segmented segments.

[0054] In the above technical solution, the resource segmentation reference position is determined based on the segmentation start position and the reference duration of the corresponding audio and video resources. Then, when the resource segmentation reference position is determined as an adjustment position, it is adjusted to obtain the resource segmentation target position. This way, when segmenting audio and video resources according to the target position, the duration of the resulting segmented segments is close to the reference duration, avoiding excessively long or short segments and ensuring that the duration of the segmented segments is within a certain range. This allows users to better focus their attention on the content of each segment of the audio and video resource, thereby improving overall attention to the audio and video resource. Furthermore, the adjustment position during segmentation includes positions within natural sentences and positions between two semantically coherent natural sentences. This ensures that the segmented segments do not end at positions within natural sentences or between two semantically coherent natural sentences. The ending sentences of the resulting segmented segments can express complete semantics, guaranteeing the integrity and independence of each segmented segment's content, further enhancing the user experience.

[0055] After executing step S104, the audio and video resources are divided into segmented segments and the remaining audio and video resources. The remaining audio and video resources can then be further segmented. Optionally, the method further includes:

[0056] If the duration of the remaining audio and video resources obtained after segmenting the segment is greater than the reference duration of the segment, then the starting position of the remaining audio and video resources is taken as the new segmentation starting position, and the process returns to step S102, which determines the resource segmentation reference position based on the segmentation starting position and the reference duration of the segment, until the duration of the remaining audio and video resources is less than or equal to the reference duration of the segment.

[0057] For example, during step S104, the audio and video resources are divided into segmented segments and remaining audio and video resources at the resource segmentation target location. If the duration of the remaining audio and video resources is greater than the segment reference duration, the starting position of the remaining audio and video resources can be used as the new segmentation starting position to continue segmenting the remaining audio and video resources based on this new segmentation starting position. The segmentation method can be as described in steps S102-S104 above, and will not be repeated here. The above steps can be executed repeatedly to segment the audio and video resources into multiple segmented segments until, after segmenting the audio and video resources, the remaining audio and video resources are less than or equal to the segment reference duration, at which point the segmentation can be stopped (i.e., after this segmentation, the remaining audio and video resources after this segmentation will not be segmented again).

[0058] In this embodiment, after each segmentation of the audio / video resource, if the remaining audio / video resource is longer than the reference segment duration, the remaining audio / video resource is segmented again to obtain another segment, until the duration of the remaining audio / video resource is less than or equal to the reference segment duration. In this way, an audio / video resource can be segmented into multiple segments of moderate length with relatively complete semantic meaning. Users can obtain a better user experience when using the segmented audio / video resource (multiple segments of moderate length with relatively complete semantic meaning).

[0059] Figure 2 This is a flowchart of step S103 of an audio / video resource processing method provided according to an embodiment of the present disclosure.

[0060] like Figure 2 As shown, step S103 adjusts the resource partitioning reference position to obtain the resource partitioning target position. An exemplary implementation method is as follows, which may include steps S1031 and S1032.

[0061] In step S1031, the semantic end position that is adjacent to the resource splitting reference position and does not belong to the adjustment position is determined.

[0062] In step S1032, the resource splitting target position is determined from the adjacent semantic end positions.

[0063] A semantic ending position can be the end of a natural statement where there is no semantic coherence between that natural statement and its next natural statement. In one implementation, a position that is not a transition position can be used as a semantic ending position.

[0064] In one implementation, during the execution of step S1031, the semantic end position that is closest to the resource segmentation reference position and is located 10 seconds before the resource segmentation reference position can be determined from the resource segmentation reference position. Alternatively, the semantic end position that is closest to the resource segmentation reference position and is located 12 seconds after the resource segmentation reference position can be determined from the resource segmentation reference position and is located 10 seconds after the resource segmentation reference position.

[0065] As described above, during step S1031, two semantic end positions adjacent to the resource segmentation reference position are determined forward and backward. In step S1032, the semantic end position closest in time to the resource segmentation reference position can be used as the resource segmentation target position. Continuing the example above, the adjacent semantic end position before the resource segmentation reference position can be determined as the resource segmentation target position. If adjacent semantic end positions are equidistant from the resource segmentation reference position in time, one of them can be randomly selected as the resource segmentation target position.

[0066] In this embodiment, the target position for resource segmentation is determined from the adjacent semantic end positions corresponding to the resource segmentation reference position, so that the determined target position for resource segmentation is a position close to the resource segmentation reference position. The segmented segments obtained after segmentation end with the semantic end position, which makes it easier for users to understand the text content contained in the audio and video resources. In addition, the duration of the obtained segmented segments is similar to the duration of the segment reference, which brings a better user experience.

[0067] Optionally, the method further includes:

[0068] If the duration of the last segment of the audio / video resource is less than the merging threshold, then the last segment is merged with the preceding segment. The last segment of the audio / video resource is the segment corresponding to the remaining audio / video resource after the segmentation process.

[0069] The merging threshold can be preset according to the actual application scenario, such as being set to 1 minute. If, after multiple segmentations, the last segment is less than 1 minute, it can be considered that the duration of the segment is too short and the segment is too fragmented. In this case, the last segment can be merged with the previous segment. For example, if the audio and video resources are segmented into 5 segments in the order A1, A2, A3, A4, and A5, and the duration of A5 is 52 seconds, then A5 can be merged into A4, that is, the audio and video resources are finally segmented into 4 segments, namely A1, A2, A3, A4' (A4+A5).

[0070] In this embodiment, if the duration of the last segment of the audio and video resource is less than the merging threshold, the last segment is merged with the previous segment to avoid the last segment being too short. This ensures that users can obtain rich content while listening to or watching each segment, thus improving the user experience.

[0071] Optionally, the reference duration of a segment is determined in the following way:

[0072] If the total duration of the audio and video resources is less than or equal to the product of the preset maximum number of segments and the first duration, the first duration will be determined as the segment reference duration.

[0073] If the total duration of the audio and video resources is greater than the product of the maximum number of segments and the first duration, the second duration is determined as the reference duration of the segments. The second duration is the quotient of the total duration of the audio and video resources to be processed and the maximum number of segments.

[0074] If a single audio / video resource is divided into too many segments, it becomes inconvenient for users to watch and learn. Therefore, the maximum number of segments a given audio / video resource can be divided into can be preset according to the actual application scenario. The maximum number of segments is the maximum number of segments obtained after dividing the audio / video resource. For example, for resources under a video course, the maximum number of segments can be preset to 8. Thus, audio / video resources under a video course can be divided into a maximum of 8 segments. Similarly, for audio / video resources under an audiobook course, the maximum number of segments can also be preset to 8.

[0075] The first duration can be a preset duration intended as a reference duration for a segment for the audio / video resource. For example, the first duration of an audio / video resource corresponding to a learning course in a video genre can be set to 3 minutes. If the duration of the audio / video resource is less than or equal to 24 minutes, such as 20 minutes, meaning the total duration of the audio / video resource is less than or equal to the product of the preset maximum number of segments (8) and the first duration of 3 minutes, then the segment reference duration can be determined as 3 minutes. If the duration of the audio / video resource is greater than 24 minutes, such as 40 minutes, meaning the total duration of the audio / video resource is greater than the product of the maximum number of segments (8) and the first duration of 3 minutes, then the segment reference duration can be determined based on the total duration of the audio / video resource and the maximum number of segments. In this example, the segment reference duration can be determined as 5 (40 / 8) minutes.

[0076] In this embodiment, the reference duration of a segment is determined based on the total duration of the audio and video resources, the maximum number of segments, and the first duration. This ensures that when segmenting the audio and video resources, both the number of segments and the duration of the segments are considered, avoiding the segmentation of the audio and video resources into too many segments that would affect the user's learning efficiency and improving the user experience.

[0077] Figure 3 This is a flowchart of an audio / video resource processing method according to one embodiment of this disclosure. Figure 3 As shown, compared to Figure 1The method further includes steps S105, S106 and S107, which are executed before step S103 (if the resource partitioning reference position is determined to be an adjustment position, the resource partitioning reference position is adjusted to obtain the resource partitioning target position).

[0078] In step S105, a third duration is determined between the beginning of the natural segment where the resource segmentation reference position is located and the segmentation start position corresponding to the audio and video resources, and a fourth duration is determined between the end of the natural segment where the resource segmentation reference position is located and the segmentation start position.

[0079] After step S102 is executed, the resource segmentation reference position can be determined, and this resource segmentation reference position may fall within a paragraph of the text content. Generally, the semantic coherence within a paragraph is relatively high. Therefore, in this embodiment, the time distance between the resource segmentation reference position and the beginning of the paragraph, and the time distance between the resource segmentation reference position and the end of the paragraph can be determined respectively.

[0080] In step S106, if both the third and fourth durations are within the segment duration interval, the beginning or end of the natural segment where the resource segmentation reference position is located is determined as the new resource segmentation reference position.

[0081] The segment duration range can be determined based on the segment reference duration. For example, the segment duration range can be set to fluctuate by 20 seconds above and below the segment reference duration. For instance, if the segment reference duration is 4 minutes, the segment duration range can be set to 3 minutes and 40 seconds to 4 minutes and 20 seconds.

[0082] For example, if the third duration is 3 minutes and 42 seconds and the fourth duration is 4 minutes and 10 seconds, both durations fall within the segment duration range. This means that the beginning or end of the segment containing the resource segmentation reference position meets the segmentation duration requirement, and either one can be selected as the new resource segmentation reference position. The specific choice between the beginning and end of the segment containing the resource segmentation reference position is not limited here. For example, the beginning of the segment containing the resource segmentation reference position can be prioritized, as can the end of the segment containing the resource segmentation reference position. Alternatively, the position corresponding to the duration of the third or fourth duration that is closer to the segment reference duration (the third duration corresponds to the beginning of the segment, and the fourth duration corresponds to the end of the segment) can be selected as the new resource segmentation reference position.

[0083] In step S107, if only one of the third and fourth durations falls within the segment duration interval, the position corresponding to the duration within the segment duration interval is determined as the new resource segmentation reference position. For example, if the third duration is 3 minutes and 29 seconds and the fourth duration is 4 minutes and 10 seconds, and the fourth duration falls within the segment duration interval, then the end of the natural segment containing the resource segmentation reference position can be determined as the new resource segmentation reference position.

[0084] In this embodiment, after determining the resource segmentation reference position, the duration corresponding to the beginning and end of the natural segment where the resource segmentation reference position is located is further determined, so as to adjust the resource segmentation reference position to the beginning or end of the natural segment where it is located, thereby further improving the rationality and effectiveness of segmenting the fragments.

[0085] Optionally, the audio and video resources are in the genre of audiobooks, and the position adjustment also includes the adjacent positions after the preset structure. The preset structure includes one or more of the following: the position of preset punctuation, the position at the end of the title, and the position between the image and the image description.

[0086] Preset punctuation marks can be predefined; for example, in English, the predefined punctuation marks could be colons and question marks. Since the text following a colon usually explains the text preceding it, and the text following a question mark usually answers the question, audio and video resources can be left unsegmented after the colon and question mark to ensure the integrity of the segmented content.

[0087] For audiobooks, since books typically contain chapters, audiobooks also usually include chapter titles, and the end of the chapter title is the title end position. Generally, in the text content corresponding to audio-visual resources of the audio-visual genre, the title is clearly identified. For example, the title can be "Chapter 1, Section 1 XXX", "Chapter 2, Section 1 XXX", or "section1 XXX". The phrases or sentences following these titles can be clearly identified by such identifiers. In one implementation, the preset structure may include the title end position.

[0088] For audiobook-style audio and video resources, it can include not only audio, but also text content and illustrations corresponding to the audio when the application plays the audiobook resource. For audiobook resources with illustrations, the text content corresponding to the audio may include explanations of the illustrations, which are called picture descriptions. Since the pictures are inserted into the text content of the audio, the timeline corresponding to the text content of the audio is known, and the position where the pictures are inserted into the text content when the application displays the pictures is also known, so the position of the pictures can be known. Picture descriptions may contain obvious identifiers, such as sentences in the text content with the structure "picture 1", "picture 2", etc., which are "picture" followed by numbers, which may be picture descriptions. In one implementation, the preset structure may include the position between the pictures and the picture descriptions.

[0089] In this embodiment, the adjustment of position includes the adjacent position after the preset structure. Audio and video resources are not cut at the adjacent position after the preset structure, so that closely related content will not be cut into two segments, which is convenient for users and improves the user experience.

[0090] Optionally, the method further includes:

[0091] The audio and video resources are displayed according to their respective segments.

[0092] For example, when displaying audio and video resources, they can be displayed in a hierarchical manner. Directory identifiers can be used to distinguish resources within the same directory level.

[0093] A parent directory for audio and video resources may contain multiple audio and video resources. For example, the parent directory for audio and video resources could be "albums," and each album may contain multiple audio and video resources. Different "albums" can be distinguished by directory identifiers such as "Album 1," "Album 2," etc. Audio and video resources under an "album" can be distinguished by their names or titles as directory identifiers. After audio and video resources are segmented, each audio and video resource may contain multiple segments, which can be distinguished by directory identifiers such as "Segment 1," "Segment 2," etc.

[0094] Figures 4a-4d This is a schematic diagram of a directory used when displaying audio and video resources in an interactive interface, according to an embodiment of the present disclosure. Figures 4a-4d In the middle, the directory display page is used to show the directories corresponding to audio and video resources. Figures 4a-4d Four ways to display the catalog corresponding to audio and video resources are shown. When displaying audio and video resources, at least one of the following (1)-(4) can be included.

[0095] (1) If there is only one audio / video resource in the parent directory of the audio / video resource, and the audio / video resource contains a segment, then the parent directory and the directory identifier of the audio / video resource will be displayed in the directory display page.

[0096] For example, if an "album" contains only one audio / video resource, and that resource contains only one segment, the directory used to display the audio / video resources under that album can be presented in the structure "'Album'-'Audio / Video Resource Name'". For instance, if album 1 contains only one audio / video resource, and that resource contains only one segment, the directory used to display the audio / video resources under album 1 can be presented as follows: Figure 4a The table shows the directories corresponding to the audio and video resources.

[0097] (2) If there is only one audio / video resource in the parent directory, and the audio / video resource contains at least two segmented segments, then the directory identifiers of the parent directory and the segmented segments will be displayed in the directory display page.

[0098] For example, if an "album" contains only one audio / video resource, and that resource includes at least two segments, then the directory used to display the audio / video resources under that album can be presented in the structure "'Album' - 'Directory identifier of segment'". That is, the directory omits the directory identifier of the audio / video resource and simplifies the hierarchical structure corresponding to the "audio / video resource". For instance, if album 2 contains only one audio / video resource, and that resource includes four segments (segments 1-4), then when displaying the audio / video resources under album 2, it can be presented as follows: Figure 4b The table shows the directories corresponding to the audio and video resources.

[0099] (3) If the parent directory contains at least two audio and video resources, and each audio and video resource in the parent directory is a segment, then the directory identifier of the parent directory and each audio and video resource in the parent directory will be displayed in the directory display page.

[0100] For example, an "album" contains at least two audio / video resources, and each audio / video resource in the "album" contains only one segment. The directory used to display the audio / video resources under this album can be presented in the structure "'Album'-'Audio / Video Resource Name'". For instance, album 3 contains three audio / video resources, 2-4, and each audio / video resource 2-4 contains only one segment. When displaying the audio / video resources under album 3, it can be presented as follows: Figure 4c The table shows the directories corresponding to the audio and video resources.

[0101] (4) If the parent directory contains at least two audio and video resources, and the number of segments of at least one audio and video resource in the parent directory is more than one, then the parent directory, the directory identifier of each audio and video resource under the parent directory, and the directory identifier of each segment under each audio and video resource will be displayed in the directory display page.

[0102] For example, if an "album" contains at least two audio / video resources, and at least one of the audio / video resources in the "album" has more than one segment, then the directory used to display the audio / video resources under that album can be presented in the structure "'Album'-'Audio / Video Resource Name'-Directory Identifier of Segment'". For instance, album 4 contains two audio / video resources, audio / video resources 5 and 6, and audio / video resource 5 contains four segments (segments 5-8), while audio / video resource 6 contains one segment. When displaying the audio / video resources under album 4, it can be presented as follows: Figure 4d The table shows the directories corresponding to the audio and video resources.

[0103] In this embodiment, the audio and video resources are displayed according to the number of audio and video resources contained in the parent directory and the number of segments contained in each audio and video resource in the parent directory. The number of directory levels used during display is minimized and the directories used meet the requirements for distinguishing each audio and video resource and each segment under the directory, thereby improving the user experience.

[0104] Figure 5 This is a block diagram of an audio / video resource processing apparatus provided according to one embodiment of the present disclosure. Figure 5 As shown, the audio and video resource processing device 400 includes an acquisition module 401, a first determination module 402, an adjustment module 403, and a segmentation module 404.

[0105] The acquisition module 401 is used to acquire the text information corresponding to the audio and video resources to be processed;

[0106] The first determining module 402 is used to determine the resource segmentation reference position based on the segmentation start position and segment reference duration corresponding to the audio and video resources;

[0107] The adjustment module 403 is used to adjust the resource segmentation reference position if the resource segmentation reference position is determined to be any of the adjustment positions, so as to obtain the resource segmentation target position. The adjustment position includes the position in the sentence of a natural sentence and the position between two semantically coherent natural sentences.

[0108] The segmentation module 404 is used to segment audio and video resources according to the resource segmentation target position to obtain segmented segments.

[0109] Optionally, the audio and video resource processing device 400 also includes a control module.

[0110] The control module is used to, if the duration of the remaining audio and video resources obtained after segmentation is greater than the reference duration of the segment, take the starting position of the remaining audio and video resources as the new segmentation starting position, and return to execute the step of determining the resource segmentation reference position according to the segmentation starting position and the reference duration of the segment, until the duration of the remaining audio and video resources is less than or equal to the reference duration of the segment.

[0111] Optionally, the adjustment module 403 includes a first determining submodule and a second determining submodule.

[0112] The first determining submodule is used to determine the semantic end position that is adjacent to the adjustment position but does not belong to the resource splitting reference position.

[0113] The second determination submodule is used to determine the target location for resource splitting from adjacent semantic end positions.

[0114] Optionally, the audio and video resource processing device 400 also includes a merging module.

[0115] The merging module is used to merge the last segment of an audio or video resource with the preceding segment if the duration of the last segment is less than the merging threshold.

[0116] Optionally, the reference duration of a segment is determined in the following way:

[0117] If the total duration of the audio and video resources is less than or equal to the product of the preset maximum number of segments and the first duration, the first duration will be determined as the segment reference duration.

[0118] If the total duration of the audio and video resources is greater than the product of the maximum number of segments and the first duration, the second duration is determined as the reference duration of the segments. The second duration is the quotient of the total duration of the audio and video resources to be processed and the maximum number of segments.

[0119] Optionally, the audio and video resource processing device 400 further includes a second determining module, a third determining module, and a fourth determining module.

[0120] The second determining module is used to determine the third duration between the beginning of the natural segment where the resource segmentation reference position is located and the segmentation start position corresponding to the audio and video resources, and to determine the fourth duration between the end of the natural segment where the resource segmentation reference position is located and the segmentation start position.

[0121] The third determination module is used to determine the beginning or end of the natural segment where the resource segmentation reference position is located as the new resource segmentation reference position if both the third duration and the fourth duration are within the segment duration interval.

[0122] The fourth determination module is used to determine the position corresponding to the duration within the segment duration interval as the new resource segmentation reference position if only one of the third duration and the fourth duration belongs to the segment duration interval.

[0123] Optionally, the audio and video resources are in the genre of audiobooks, and the position adjustment also includes the adjacent positions after the preset structure. The preset structure includes one or more of the following: the position of preset punctuation, the position at the end of the title, and the position between the image and the image description.

[0124] Optionally, the audio and video resource processing device 400 also includes a display module.

[0125] The display module is used to display the audio and video resources according to the various segments corresponding to the audio and video resources;

[0126] Displaying audio and video resources includes at least one of the following:

[0127] If there is only one audio / video resource in the parent directory of the audio / video resource, and the audio / video resource contains a segment, then the directory identifiers of the parent directory and the audio / video resource will be displayed in the directory display page.

[0128] If there is only one audio / video resource in the parent directory, and the audio / video resource contains at least two segmented segments, then the directory identifiers of the parent directory and the segmented segments will be displayed in the directory display page.

[0129] If the parent directory contains at least two audio and video resources, and each audio and video resource in the parent directory is a segment, then the directory identifier of the parent directory and each audio and video resource in the parent directory will be displayed in the directory display page.

[0130] If the parent directory contains at least two audio / video resources, and the number of segments of at least one audio / video resource in the parent directory is more than one, then the directory display page will show the parent directory, the directory identifier of each audio / video resource under the parent directory, and the directory identifier of each segment under each audio / video resource.

[0131] The following is for reference. Figure 6 This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0132] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0133] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0134] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0135] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0136] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0137] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0138] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following: to acquire text information corresponding to the audio and video resources to be processed; to determine a resource segmentation reference position based on the segmentation start position and segment reference duration corresponding to the audio and video resources; if the determined resource segmentation reference position is any of the adjustment positions, to adjust the resource segmentation reference position to obtain a resource segmentation target position, wherein the adjustment position includes the position within a natural sentence and the position between two semantically coherent natural sentences; and to segment the audio and video resources according to the resource segmentation target position to obtain segmented segments.

[0139] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0141] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, an acquisition module can also be described as "a module for acquiring text information corresponding to audio and video resources to be processed".

[0142] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0143] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0144] According to one or more embodiments of this disclosure, Example 1 provides an audio and video resource processing method, including: acquiring text information corresponding to the audio and video resource to be processed; determining a resource segmentation reference position based on the segmentation start position and segment reference duration corresponding to the audio and video resource; if the determined resource segmentation reference position is any of the adjustment positions, adjusting the resource segmentation reference position to obtain a resource segmentation target position, wherein the adjustment position includes the position in the sentence of a natural sentence and the position between two semantically coherent natural sentences; segmenting the audio and video resource according to the resource segmentation target position to obtain segmented segments.

[0145] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the method further includes: if the duration of the remaining audio and video resources obtained after segmenting the segment is greater than the reference duration of the segment, then the starting position of the remaining audio and video resources is taken as the new segmentation starting position, and the step of determining the resource segmentation reference position according to the segmentation starting position and the reference duration of the segment corresponding to the audio and video resources is returned to be executed until the duration of the remaining audio and video resources is less than or equal to the reference duration of the segment.

[0146] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 1, wherein adjusting the resource segmentation reference position to obtain the resource segmentation target position includes: determining adjacent semantic end positions that do not belong to the adjusted position corresponding to the resource segmentation reference position; and determining the resource segmentation target position from the adjacent semantic end positions.

[0147] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 1, wherein the method further includes: if the duration of the last segment of the audio / video resource is less than a merging threshold, then merging the last segment with the preceding segment.

[0148] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 1, wherein the reference duration of the segment is determined in the following manner: if the total duration of the audio and video resources is less than or equal to the product of a preset maximum number of segments and a first duration, the first duration is determined as the reference duration of the segment; if the total duration of the audio and video resources is greater than the product of the maximum number of segments and the first duration, a second duration is determined as the reference duration of the segment, wherein the second duration is the quotient of the total duration of the audio and video resources to be processed and the maximum number of segments.

[0149] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 1, wherein, before the step of adjusting the resource segmentation reference position to obtain the resource segmentation target position if the resource segmentation reference position is determined to be an adjustment position, the method further includes: determining a third duration between the beginning of the natural segment where the resource segmentation reference position is located and the segmentation start position corresponding to the audio / video resource, and determining a fourth duration between the end of the natural segment where the resource segmentation reference position is located and the segmentation start position; if both the third duration and the fourth duration belong to the segment duration interval, the beginning or end of the natural segment where the resource segmentation reference position is located is determined as the new resource segmentation reference position; if only one of the third duration and the fourth duration belongs to the segment duration interval, the position corresponding to the duration within the segment duration interval is determined as the new resource segmentation reference position.

[0150] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein the genre of the audio and video resource is an audiobook genre, and the adjustment position further includes the adjacent position after a preset structure, wherein the preset structure includes one or more of the following: the position of a preset punctuation mark, the position at the end of the title, and the position between the image and the image description.

[0151] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 1, wherein the method further includes: displaying the audio and video resources according to each segment corresponding to the audio and video resources; wherein displaying the audio and video resources includes at least one of the following: if there is only one audio and video resource in the parent directory of the audio and video resource, and the audio and video resource contains one segment, then the parent directory and the directory identifier of the audio and video resource are displayed in the directory of the directory display page; if there is only one audio and video resource in the parent directory, and the audio and video resource contains at least two segment, then the directory identifier of the parent directory and the audio and video resource are displayed in the directory of the directory display page. The directory identifiers of the parent directory and the segmented segments are displayed. If the parent directory contains at least two audio and video resources, and each audio and video resource in the parent directory has one segment, then the directory identifiers of the parent directory and each audio and video resource under the parent directory are displayed in the directory display page. If the parent directory contains at least two audio and video resources, and at least one audio and video resource in the parent directory has more than one segment, then the directory identifiers of the parent directory, each audio and video resource under the parent directory, and the directory identifiers of the segmented segments under each audio and video resource are displayed in the directory display page.

[0152] According to one or more embodiments of this disclosure, Example 9 provides an audio / video resource processing apparatus, the apparatus comprising: an acquisition module, configured to acquire text information corresponding to the audio / video resource to be processed; a first determination module, configured to determine a resource segmentation reference position based on the segmentation start position and segment reference duration corresponding to the audio / video resource; an adjustment module, configured to adjust the resource segmentation reference position to obtain a resource segmentation target position if the determined resource segmentation reference position is any of the adjustment positions, wherein the adjustment position includes the position within a sentence of a natural language and the position between two semantically coherent natural language sentences; and a segmentation module, configured to segment the audio / video resource according to the resource segmentation target position to obtain segmented segments.

[0153] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-8.

[0154] According to one or more embodiments of this disclosure, Example 11 provides an electronic device, including:

[0155] A storage device on which computer programs are stored;

[0156] A processing device for executing the computer program in the storage device to implement the steps of any one of the methods in Examples 1-8.

[0157] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0158] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0159] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A method for processing audio and video resources, characterized in that, include: Obtain the text information corresponding to the audio and video resources to be processed; The resource segmentation reference position is determined based on the segmentation start position and segment reference duration corresponding to the audio and video resources; If the resource segmentation reference position is determined to be any of the adjustment positions, the resource segmentation reference position is adjusted to obtain the resource segmentation target position, wherein the adjustment position includes the position in the sentence of a natural sentence and the position between two semantically coherent natural sentences; The audio and video resources are segmented according to the target location of the resource segmentation to obtain segmented segments; The method further includes: If the duration of the remaining audio and video resources obtained after segmenting the segment is greater than the reference duration of the segment, then the starting position of the remaining audio and video resources is taken as the new segmentation starting position, and the step of determining the resource segmentation reference position based on the segmentation starting position and the reference duration of the segment is returned to be executed until the duration of the remaining audio and video resources is less than or equal to the reference duration of the segment. The reference duration of the segment is determined in the following way: If the total duration of the audio and video resources is less than or equal to the product of the preset maximum number of segments and the first duration, the first duration is determined as the reference duration of the segment; If the total duration of the audio and video resources is greater than the product of the maximum number of segments and the first duration, the second duration is determined as the reference duration of the segments, and the second duration is the quotient of the total duration of the audio and video resources to be processed and the maximum number of segments.

2. The method according to claim 1, characterized in that, The step of adjusting the resource partitioning reference position to obtain the resource partitioning target position includes: Identify the semantic end positions that are adjacent to the adjusted position but do not belong to the resource segmentation reference position; The resource segmentation target position is determined from the adjacent semantic end positions.

3. The method according to claim 1, characterized in that, The method further includes: If the duration of the last segment of the audio / video resource is less than the merging threshold, then the last segment will be merged with the preceding segment.

4. The method according to claim 1, characterized in that, Before the step of adjusting the resource partitioning reference position to obtain the resource partitioning target position if the resource partitioning reference position is determined to be the adjustment position, the method further includes: A third duration is determined between the beginning of the natural segment where the resource segmentation reference position is located and the segmentation start position corresponding to the audio and video resource; and a fourth duration is determined between the end of the natural segment where the resource segmentation reference position is located and the segmentation start position. If both the third duration and the fourth duration are within the segment duration interval, the beginning or end of the natural segment where the resource segmentation reference position is located will be determined as the new resource segmentation reference position. If only one of the third duration and the fourth duration falls within the segment duration interval, then the position corresponding to the duration within the segment duration interval is determined as the new resource segmentation reference position.

5. The method according to claim 1, characterized in that, The audio and video resources are in the genre of audiobooks. The adjustment position also includes the adjacent position after the preset structure. The preset structure includes one or more of the following: the position of preset punctuation, the position at the end of the title, and the position between the image and the image description.

6. The method according to claim 1, characterized in that, The method further includes: The audio and video resources are displayed according to their respective segments; Displaying the audio and video resources includes at least one of the following: If there is only one audio / video resource in the parent directory of the audio / video resource, and the audio / video resource contains a segment, then the parent directory and the directory identifier of the audio / video resource will be displayed in the directory display page. If there is only one audio / video resource in the parent directory, and the audio / video resource contains at least two segmented segments, then the directory identifiers of the parent directory and the segmented segments will be displayed in the directory display page. If the parent directory contains at least two audio and video resources, and each audio and video resource in the parent directory is a segment, then the directory identifier of the parent directory and each audio and video resource under the parent directory will be displayed in the directory display page. If the parent directory contains at least two audio and video resources, and the number of segments of at least one audio and video resource in the parent directory is more than one, then the parent directory, the directory identifier of each audio and video resource under the parent directory, and the directory identifier of each segment under the audio and video resource are displayed in the directory display page.

7. An audio and video resource processing device, characterized in that, include: The acquisition module is used to acquire the text information corresponding to the audio and video resources to be processed; The first determining module is used to determine the resource segmentation reference position based on the segmentation start position and segment reference duration corresponding to the audio and video resources; An adjustment module is used to adjust the resource segmentation reference position if it is determined that the resource segmentation reference position is any of the adjustment positions, so as to obtain the resource segmentation target position, wherein the adjustment position includes the position in the sentence of a natural sentence and the position between two semantically coherent natural sentences. The segmentation module is used to segment the audio and video resources according to the resource segmentation target position to obtain segmented segments; The control module is configured to, if the duration of the remaining audio and video resources obtained after segmenting the segment is greater than the reference duration of the segment, take the starting position of the remaining audio and video resources as the new segmentation starting position, and return to execute the step of determining the resource segmentation reference position based on the segmentation starting position and the reference duration of the segment corresponding to the audio and video resources, until the duration of the remaining audio and video resources is less than or equal to the reference duration of the segment; The reference duration of the segment is determined in the following way: If the total duration of the audio and video resources is less than or equal to the product of the preset maximum number of segments and the first duration, the first duration is determined as the reference duration of the segment; If the total duration of the audio and video resources is greater than the product of the maximum number of segments and the first duration, the second duration is determined as the reference duration of the segments, and the second duration is the quotient of the total duration of the audio and video resources to be processed and the maximum number of segments.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-6.

9. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Audio editing system and audio editing method

    CN102543080A

  • Video slicing method based on slicing file duration threshold

    CN109525893A