Cross-screen continuous playing method and device, electronic equipment and storage medium

By analyzing user behavior data and copyright information, dynamically adjusting the slice length, and combining multimodal feature extraction, the problem of inconsistent playback of copyrighted videos in cross-screen continuation playback was solved, achieving a highly accurate cross-screen continuation playback experience.

CN121567904APending Publication Date: 2026-02-24MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511479609.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively solve the problem of inconsistent playback times caused by different video versions provided by different copyright holders when resuming playback across screens, resulting in users encountering situations such as plot skipping or repeated playback.

Method used

By analyzing user behavior data, the event density and copyright information of the video to be resumed are determined, the slice length is dynamically adjusted, and the resumed playback position is accurately determined through multimodal feature extraction and similarity calculation.

Benefits of technology

It enables precise matching of video content across different copyrighted video libraries, ensuring seamless cross-screen playback and improving the smoothness and user experience of cross-screen playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567904A_ABST
    Figure CN121567904A_ABST
Patent Text Reader

Abstract

The invention provides a cross-screen continuous playing method and device, electronic equipment and a storage medium, and the method comprises the steps: responding to a received video playing request which is transmitted by first equipment and comprises user behavior data of a first video, and determining a second video to be continuously played according to the user behavior data; determining the event density of the second video, and determining the slice length according to the event density and the copyright information of the second video; wherein the event density refers to the number of events occurring in unit time; determining at least one candidate slice from the second video according to the slice length and the playing stop time point of the first video; and according to the at least one candidate slice, determining a playing start time point of the second video, so as to play the second video on the first device from the playing start time point. Therefore, the slice length can be dynamically determined, the accuracy of the continuous playing position is improved, the fluency of cross-screen continuous playing is improved, and the smooth cross-screen continuous playing experience is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video technology, and in particular to a method, apparatus, electronic device and storage medium for cross-screen playback continuation. Background Technology

[0002] Cross-screen resume playback is an intelligent media service feature that allows users to automatically resume playback on another device after pausing audio or video content on one device, without having to search for the content or reposition the playback progress.

[0003] In related technologies, continuation playback mainly relies on simple playback progress record matching. However, when faced with video versions provided by different copyright holders, this method often fails to effectively resolve the issue of inconsistent playback times, leading to situations such as plot jumps or repeated playback when users resume playback across different screens. Summary of the Invention

[0004] This application aims to at least partially address one of the technical problems in the related art.

[0005] Therefore, the first objective of this application is to propose a method for cross-screen playback continuation.

[0006] The second objective of this application is to propose a cross-screen playback continuation device.

[0007] The third objective of this application is to propose an electronic device.

[0008] The fourth objective of this application is to provide a computer-readable storage medium.

[0009] The fifth objective of this application is to provide a computer program product.

[0010] To achieve the above objectives, the first aspect of this application proposes a cross-screen playback continuation method, comprising: In response to receiving a video playback request sent by a first device, wherein the video playback request includes user behavior data of a first video, a second video to be resumed is determined based on the user behavior data; wherein the first video is a video that has stopped playing on the second device, and the second video is a different version from the first video; The event density of the second video is determined, and the slice length is determined based on the event density and the copyright information of the second video; wherein, the event density refers to the number of events occurring per unit time. Based on the slice length and the stop playback time of the first video, at least one candidate slice is determined from the second video; Based on the at least one candidate slice, the start time point of the second video is determined so that the second video is played on the first device from the start time point.

[0011] To achieve the above objectives, a second aspect of this application provides a cross-screen playback continuation device, comprising: A first determining module is configured to respond to receiving a video playback request sent by a first device, wherein the video playback request includes user behavior data of a first video, and determine a second video to be resumed playback based on the user behavior data; wherein the first video is a video that has stopped playing on the second device, and the second video is a different version from the first video; The second determining module is used to determine the event density of the second video and determine the slice length based on the event density and the copyright information of the second video; wherein, the event density refers to the number of events occurring per unit time. The third determining module is used to determine at least one candidate slice from the second video based on the slice length and the stop playback time of the first video. The fourth determining module is configured to determine the start playback time point of the second video based on the at least one candidate slice, so as to play the second video on the first device from the start playback time point.

[0012] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method described in the first aspect embodiment above.

[0013] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method described in the first aspect of the present application.

[0014] To achieve the above objectives, a fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect of the application.

[0015] The cross-screen playback continuation method, apparatus, electronic device, and storage medium provided in this application, if receiving a video playback request sent by a first device, and the video playback request includes user behavior data of the first video, determines the second video to be resumed based on the user behavior data, and dynamically determines the slice length based on the event density and copyright information of the second video, can ensure that the slice length matches the content characteristics. Then, based on the slice length and the stop playback time of the first video, the candidate slices are determined from the second video, which can improve the accuracy of the candidate slices. Thus, based on the candidate slices, the start playback time of the second video is determined, which can improve the accuracy of the continuation position. This can ensure that the user can seamlessly and accurately continue watching the content on the new device, improve the smoothness of cross-screen playback continuation, and achieve a smooth cross-screen playback continuation experience.

[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0017] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating a cross-screen playback continuation method provided in an embodiment of this application; Figure 2 A flowchart illustrating another cross-screen playback method provided in this application embodiment; Figure 3 A flowchart illustrating another cross-screen playback method provided in this application embodiment; Figure 4 This is a schematic diagram of a cross-screen playback device provided in an embodiment of this application. Detailed Implementation

[0018] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0019] It should be noted that the acquisition, storage, use, and processing of data in this application comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.

[0020] The following description, with reference to the accompanying drawings, describes a cross-screen playback method, apparatus, electronic device, and storage medium according to embodiments of this application.

[0021] Figure 1 This is a flowchart illustrating a cross-screen playback method provided in an embodiment of this application.

[0022] like Figure 1 As shown, this cross-screen playback continuation method includes the following steps: Step 101: In response to receiving a video playback request sent by the first device, and the video playback request includes user behavior data of the first video, determine the second video to be played based on the user behavior data.

[0023] In this application, in the cross-screen playback scenario, when a user switches from a second device to a first device to continue watching video content, the first device can send a video playback request to the server, which may include user behavior data of the first video.

[0024] The first video is the video that has stopped playing on the second device, which can be a different device from the first. For example, the first device could be a mobile phone and the second device could be a television, or the first device could be a television and the second device could be a mobile phone, or the first device could be a tablet and the second device could be a television, etc. There are no restrictions on this.

[0025] User behavior data may include, but is not limited to, user identifiers, content identifiers, operation types, context sources, device information of the first device, and session identifiers. The content identifier may refer to the video identifier of the first video, the operation type may be a click or trigger, and the context source may be a history playback page or screen mirroring, etc.

[0026] As an application scenario, a user watches a first video in a video application on a second device. When the user pauses playback at a certain point, the user later opens the playback history of the same video application on the first device. The user can click on or trigger the playback history associated with the first video. At this time, the first device can send a video playback request to the server. This video playback request can include the user's behavior data for the first video.

[0027] As another application scenario, users can trigger the screen mirroring function to cast the first video playing on the second device to the first device for continued playback. At this time, the first device can send a video playback request to the server, which may include user behavior data of the first video.

[0028] When a user switches from one device to another to continue watching content, different devices may involve different copyrighted videos. For example, the same movie or TV series may have different copyrights, such as large-screen copyrights and small-screen copyrights. In other words, the videos played on devices with different screen sizes may have different copyrights.

[0029] Based on this, in this application, if the video playback request sent by the first device includes user behavior data of the first video, it can be assumed that the user wants to continue playing the first video on the first device. Therefore, based on the user behavior data, a second video to be continued can be searched in a video library with different copyrights. The second video can be a version different from the first video.

[0030] For example, based on the video identifier of the first video in the user behavior data, a video matching the video identifier of the first video can be found in the video library, and then a video matching the device information of the first device can be found from these videos, which can be used as the second video to be played.

[0031] For minor differences in different versions of videos due to copyright requirements (such as adjustments to screen colors, additions or subtractions of audio effects, modifications to subtitles, etc.), the relevant technologies cannot detect and effectively compare them by matching basic information such as file names, thus failing to ensure the consistency and continuity of content when playing back across screens.

[0032] Based on this, for example, in order to improve the accuracy of the second video, a large model can be used to filter out the second video that is highly similar to the content of the first video from the video library based on user behavior data, thereby reducing misjudgments and omissions and improving the success rate of cross-screen playback.

[0033] Step 102: Determine the event density of the second video, and determine the slice length based on the event density and the copyright information of the second video.

[0034] Event density refers to the number of events that occur in the second video per unit of time. The higher the event density, the more drastic the content changes within that time period.

[0035] In this application, the content of a video segment in the second video can be analyzed to determine the atomic events occurring within the video segment, and the ratio between the number of atomic events and the duration of the video segment can be determined as the event density. Atomic events can be considered as indivisible smallest units, representing key points in the content that have significant changes or characteristics.

[0036] For example, atomic events in a video clip can be determined based on visual signals (such as large shot transitions, sudden changes in motion intensity, etc.), audio signals (such as sudden changes in volume, silent segments, etc.), and text signals (such as changes in subtitle segments, keyword triggering, etc.).

[0037] Because of different copyrights, the extent to which copyright holders modify the content may vary. Therefore, in order to improve the accuracy of replay, this application determines the slice length based on the event density and the copyright information of the second video after determining the event density.

[0038] The slice length is used to extract video segments of length equal to the slice length from the second video.

[0039] Step 103: Based on the slice length and the stop playback time of the first video, determine at least one candidate slice from the second video.

[0040] In this application, the historical playback information of the first video can be found based on the video identifier of the first video in the user behavior data. The time point when the first video stopped playing can be found from the historical playback information. Then, based on the slice length and the time point when the first video stopped playing, at least one candidate slice can be determined from the second video. Each candidate slice is a video segment in the second video.

[0041] For example, the segmentation range in the second video can be determined based on the stop playback time of the first video, and at least one candidate slice with a length equal to the segmentation length can be segmented within the segmentation range.

[0042] For example, the second video can be used as a slice range by pre-setting the length of the time point that is the same as the stop playback time point of the first video.

[0043] Step 104: Determine the start time point of the second video based on at least one candidate slice, so as to play the second video on the first device from the start time point.

[0044] In this application, a target slice can be determined from at least one candidate slice, and the playback time point of the second video can be determined based on the target slice, so that the second video can be played on the first device from the start playback time point.

[0045] For example, the candidate slice with the highest similarity to the slice near the stop playback time of the first video can be determined from at least one candidate slice and used as the target slice.

[0046] In this embodiment, if a video playback request is received from a first device, and the video playback request includes user behavior data of the first video, a second video to be resumed is determined based on the user behavior data. The slice length is dynamically determined based on the event density and copyright information of the second video, ensuring that the slice length matches the content characteristics. Based on the slice length and the stop playback time of the first video, candidate slices are determined from the second video, improving the accuracy of the candidate slices. Thus, based on the candidate slices, the start playback time of the second video is determined, improving the accuracy of the resume playback position. This ensures that users can seamlessly and accurately continue watching content on new devices, improving the smoothness of cross-screen resume playback and achieving a silky smooth cross-screen resume playback experience.

[0047] Figure 2 This is a flowchart illustrating another cross-screen playback method provided in an embodiment of this application.

[0048] like Figure 2 As shown, the cross-screen playback continuation method may include the following steps: Step 201: In response to receiving a video playback request sent by the first device, and the video playback request includes user behavior data of the first video, determine the second video to be played based on the user behavior data.

[0049] Step 202: Determine the event density of the second video.

[0050] In this application, steps 201-202 can be implemented in any of the embodiments of this application, so they will not be described in detail here.

[0051] Step 203: Determine the copyright coefficient based on the copyright information.

[0052] The copyright coefficient can be used to reflect the extent to which the copyright holder has modified the content. For example, the higher the copyright coefficient, the higher the possibility of content differences, and more attention needs to be paid to it when slicing.

[0053] In this application, different copyrights correspond to different copyright coefficients. A mapping relationship between copyright information and copyright coefficients can be established in advance. Based on the copyright information of the second video, the copyright coefficient corresponding to the second video can be determined by querying this preset relationship.

[0054] For example, online version, copyright coefficient r =1; Foreign version, copyright coefficient r =1.5; TV version or movie version, copyright coefficient r =2.

[0055] Step 204: Adjust the basic time unit according to the event density and copyright coefficient to obtain the slice length.

[0056] In this application, the slice adjustment coefficient can be determined based on the event density and copyright coefficient, and then the basic time unit can be adjusted according to the slice adjustment coefficient to obtain the slice length.

[0057] In some embodiments, the sum of the target value and the event density can be determined, and the slice adjustment coefficient can be determined based on the ratio of the target value to the sum and the copyright coefficient.

[0058] The target value can be a number greater than or equal to 1.

[0059] For example, the product of the ratio of the target value to the sum and the copyright coefficient can be determined as the slice adjustment coefficient.

[0060] As an example, with a target value of 1, the slice adjustment factor can be calculated using the following formula (1): (1) in, Indicates the slice adjustment factor. d Indicates event density, This represents the copyright coefficient.

[0061] According to formula (1), the event density d The higher the value, the more drastic the changes in the video content, requiring more precise slicing processing. (Slicing adjustment coefficient) The smaller the value; the lower the copyright coefficient The larger the value, the larger the slice adjustment factor. The larger the event density. For example, when the event density is higher... d Larger and copyright coefficient When smaller, It can be 0.8; when the event density... d Smaller and copyright coefficient When it is large, It can be 1.2.

[0062] Therefore, the slice adjustment coefficient can be dynamically determined based on the event density and copyright coefficient, and the slice length can be dynamically adjusted according to the slice adjustment coefficient to adapt to the characteristics and copyright requirements of different content.

[0063] For example, the slice adjustment factor can be determined by multiplying the slice adjustment factor by the base time unit. For instance, the slice adjustment factor... ,in Represents the basic time unit. This represents the slice adjustment factor.

[0064] For example, the basic time unit can be a pre-set fixed value or it can be dynamically determined; there is no limitation on this.

[0065] For example, the initial value can be adjusted based on one or more of the rhythm information, plot information, event density, shot switching frequency, and user viewing habits of the second video to obtain the basic time unit.

[0066] Among them, the rhythm information of the second video can refer to the speed of the second video's rhythm, and the plot information of the second video can refer to the speed of the plot development of the second video.

[0067] The initial values ​​for different types of videos can be the same or different; there is no restriction on this.

[0068] For example, the basic time units for different content types can be implemented using any of the following dynamically determined methods: a) Utilize artificial intelligence algorithms to analyze content types (such as movies, TV series, etc.) and dynamically calculate basic time units based on their rhythmic characteristics and plot development.

[0069] b) For different types of content, set an initial value T0 and adjust it according to the dynamic characteristics of the content (such as event density, camera switching frequency, etc.).

[0070] c) By learning from user behavior data feedback, we can study users' viewing habits for different content types, adjust the initial values, and obtain the basic time unit.

[0071] For example, for film content, the initial value T0 is initially set at 5 minutes, and dynamically adjusted according to the actual pace of the film (e.g., action films are fast-paced, so the initial value T0 can be shortened; drama films are slow-paced, so the initial value T0 can be extended). For TV series content, the initial value T0 is initially set at 10 minutes, and similarly, the initial value T0 can be adjusted according to the specific plot development of the series and the viewing habits of the users.

[0072] Therefore, by dynamically adjusting the initial values ​​based on the rhythm information, plot information, dynamic characteristics of the content, and user viewing habits of the second video, the accuracy of the basic time unit can be improved.

[0073] Step 205: Based on the slice length and the stop playback time of the first video, determine at least one candidate slice from the second video.

[0074] Step 206: Determine the start time point of the second video based on at least one candidate slice, so as to play the second video on the first device from the start time point.

[0075] In this application, steps 205-206 can be implemented in any of the embodiments of this application, so they will not be described in detail here.

[0076] In this embodiment, by determining the copyright coefficient based on the copyright information, the copyright information is converted into a numerical value. Then, the basic time unit is dynamically adjusted according to the event density and the copyright coefficient to obtain the slice length. This ensures that the slice length matches the content characteristics, thereby improving the accuracy of the slice length and, consequently, the accuracy of the resume playback position and the precision of the resume playback.

[0077] Figure 3 This is a flowchart illustrating another cross-screen playback method provided in an embodiment of this application.

[0078] like Figure 3 As shown, the cross-screen playback continuation method may include the following steps: Step 301: In response to receiving a video playback request sent by the first device, and the video playback request includes user behavior data of the first video, determine the second video to be played based on the user behavior data.

[0079] Step 302: Determine the event density of the second video, and determine the slice length based on the event density and the copyright information of the second video.

[0080] In this application, steps 301-302 can be implemented in any of the embodiments of this application, so they will not be described in detail here.

[0081] Step 303: Based on the slice length and the stop playback time of the first video, determine at least one candidate slice from the second video.

[0082] In some embodiments, the estimated time position can be determined based on the stop playback time of the first video and the metadata of the second video, and multiple candidate slices with a length equal to the slice length can be extended forward and backward in the second video with the estimated time position as the center point.

[0083] The metadata for the second video may include its playback duration, chapter markers, etc.

[0084] For example, if the first video is episode 20 of a TV series and stops playing at 4 minutes, and based on the metadata of the second video, it is determined that there is a 15-second version information introduction at the beginning of the second video, then the estimated time position can be determined to be 4 minutes and 15 seconds of episode 20.

[0085] For example, a video segment in the second video with the estimated time position as the center point and a length equal to the slice length can be used as a candidate slice. This candidate slice can be called an estimated slice. Within the range of the estimated slice, the center point can be moved with a set step size to obtain multiple candidate slices, so as to determine the start playback time of the second video based on all the obtained candidate slices.

[0086] For example, the estimated start time of the slice = Estimated end time of slice = ,in Indicates the estimated time and location. Indicates the slice length.

[0087] As an example, a sliding window is set up with the same length as the slice length. By moving the sliding window in the second video with a set step size, the center point of the sliding window moves within the range of the estimated slice, resulting in multiple candidate slices.

[0088] As an example, you can also start from the estimated time position, set a step size to move the window forward and backward, so that the center point of the sliding window moves within the estimated slice range, and obtain multiple candidate slices.

[0089] It should be noted that the center point of the sliding window can also move within a certain range that can contain the estimated slices to obtain multiple candidate slices, and there is no limitation on this.

[0090] Step 304: Perform multi-dimensional feature extraction on the candidate slice and the playback slice in the first video respectively to obtain the first multi-dimensional feature and the second multi-dimensional feature.

[0091] Here, a playback slice can refer to a video segment in the first video between the target time point and the stop playback time point, where the target time point can be earlier than the stop playback time point.

[0092] To improve the accuracy of the start time of the second video, this application performs multi-dimensional feature extraction on both the candidate slice and the playback slice, such as image feature extraction, audio feature extraction, and text feature extraction, to obtain the first multi-dimensional feature and the second multi-dimensional feature. The multi-dimensional features may include, but are not limited to, image features, audio features, and text features.

[0093] For example, image features may include, but are not limited to, color histograms, textures, and changes in the position of key objects. Among these, color histograms can reflect the color distribution of the image, texture features can reflect the details and texture of the image, and changes in the position of key objects can capture the movement trajectory of the main objects in the image.

[0094] For example, audio feature extraction may include, but is not limited to, spectrum analysis, rhythm pattern extraction, and key sound effect recognition. Correspondingly, audio features may include, but are not limited to, spectrum, rhythm pattern, and key sound effects. Among them, spectrum analysis can reflect the frequency components and energy distribution of audio, rhythm pattern can capture the rhythmic regularity of audio, and key sound effect recognition can locate specific sound effect elements.

[0095] For example, text features may include, but are not limited to, semantics, keywords, and subtitle layout features. Among them, subtitle layout features may include the layout and format of the subtitles, such as the font, size, and position of the subtitles.

[0096] Step 305: Determine the comprehensive similarity between the candidate slice and the playback slice based on the first multi-dimensional feature and the second multi-dimensional feature.

[0097] In this application, the similarity between the same type of features in the first multi-dimensional feature and the second multi-dimensional feature can be determined, and the comprehensive similarity between the candidate slice and the playback slice can be determined based on the similarity between the same type of features.

[0098] In some embodiments, the image similarity between the candidate slice and the playback slice can be determined based on the image features in the first multi-dimensional features and the image features in the second multi-dimensional features; the audio similarity between the candidate slice and the playback slice can be determined based on the audio features in the first multi-dimensional features and the audio features in the second multi-dimensional features; the text similarity between the candidate slice and the playback slice can be determined based on the text features in the first multi-dimensional features and the text features in the second multi-dimensional features; and finally, a comprehensive similarity can be determined based on the image similarity, audio similarity, and text similarity.

[0099] For example, image features may include color histograms, textures, and key object position changes. The similarity between the color histograms of candidate slices and playback slices can be calculated to obtain color histogram similarity. The texture similarity can be obtained based on the Euclidean distance between the texture feature vectors of candidate slices and playback slices. The key object position change similarity can be obtained based on the difference in position coordinates between the key object in the candidate slice and the playback slice. Finally, the color histogram similarity, texture similarity, and key object position change similarity are weighted to obtain the image similarity.

[0100] Among them, the color histogram similarity is used to represent the degree of similarity between the candidate slice and the playback slice in terms of color distribution.

[0101] Texture similarity is used to measure the degree of similarity between candidate slices and playback slices in terms of texture features.

[0102] Among them, the key object position change similarity is used to represent the degree of similarity between the positions of key objects in the candidate slice and the playback slice.

[0103] As an example, the color histogram similarity can be calculated using the following formula (2): (2) in, Indicates the similarity of color histograms. This represents the color histogram of the playback slice. A color histogram representing candidate slices.

[0104] Therefore, the color similarity between two slices can be measured by calculating the ratio of the shared color portion between the candidate slice and the playback slice to the color distribution of the playback slice. The closer to 1, the more similar the color distribution.

[0105] As an example, texture similarity can be calculated using the following formula (3): (3) in, Indicates texture similarity, The distance between the texture feature vectors of the candidate slice and the playback slice represents the degree of difference in texture features.

[0106] According to formula (3), the smaller the Euclidean distance, the more similar the texture features, and the closer the texture similarity is to 1.

[0107] As an example, the similarity of key object position changes can be calculated using the following formula (4): (4) in, Indicates the similarity of changes in the positions of key objects. This indicates the difference in position coordinates between the key object and the playback slice.

[0108] According to formula (4), the smaller the difference in position coordinates between the key object in the candidate slice and the playback slice, the closer the similarity of the key object's position change is to 1, indicating that the position of the key object is more similar in the two images.

[0109] For example, audio features may include, but are not limited to, spectrum, rhythm pattern, number of key sound effects, etc. The similarity of the audio in the candidate slice and the audio in the playback slice in the spectrum can be calculated to obtain the spectrum similarity. The rhythm pattern similarity can be obtained based on the rhythm feature sequence of the audio in the candidate slice and the rhythm feature sequence of the audio in the playback slice. The key sound effect similarity can be obtained based on the number of key sound effects in the candidate slice and the number of key sound effects in the playback slice. Then, the spectrum similarity, rhythm pattern similarity and key sound effect similarity are weighted to obtain the audio similarity.

[0110] Among them, spectral similarity is used to measure the degree of similarity between two audio segments in terms of their spectrum.

[0111] Among them, rhythm pattern similarity is used to represent the degree of similarity in rhythmic patterns between the audio of the candidate slice and the audio of the playback slice.

[0112] Among them, key sound effect similarity is used to measure the degree of similarity between the key sound effects in the audio of the candidate slice and the key sound effects in the audio of the playback slice.

[0113] As an example, spectral similarity can be calculated using the following formula (5): (5) in, Indicates spectral similarity. Let A represent the angle between the spectral vector of the candidate slice's audio and the spectral vector of the playback slice's audio. Let B represent the spectral vector of the candidate slice's audio and B represent the spectral vector of the playback slice's audio.

[0114] In other words, the cosine similarity is obtained by representing the spectra of the candidate audio slice and the playback audio slice as vectors, and calculating the ratio of their dot product to their magnitude. The closer this value is to 1, the more similar the spectral distributions of the two audio segments are.

[0115] As an example, rhythmic pattern similarity can be calculated using the following formula (6): (6) in, Indicates rhythmic pattern similarity. The correlation coefficient represents the rhythmic features of the audio from the candidate slice and the audio from the played slice, measuring the consistency of rhythmic changes. X represents the rhythmic feature sequence of the audio from the candidate slice, and Y represents the rhythmic feature sequence of the audio from the played slice. This represents the covariance of X and Y. The standard deviation of X is represented by X. This represents the standard deviation of Y.

[0116] The correlation coefficient measures the degree of linear correlation between rhythmic features; the closer the absolute value is to 1, the more similar the rhythmic patterns.

[0117] As an example, the key sound effect similarity can be calculated using the following formula (7): (7) in, Indicates the similarity of key sound effects. This indicates the number of key audio effects that match the candidate slice and the playback slice. This indicates the total number of key sound effects.

[0118] Therefore, the similarity of the audio of the candidate slice and the audio of the playback slice in terms of key sound effects can be evaluated by statistically analyzing the ratio of the key sound effects matched between the candidate slice and the playback slice to the total key sound effects.

[0119] For example, text features may include, but are not limited to, semantic features, keywords, and subtitle layout features. The semantic similarity between the text of the candidate slice and the text of the playback slice can be calculated to obtain semantic similarity. Keyword similarity can be obtained based on the keywords in the text of the candidate slice and the keywords in the text of the playback slice. Subtitle layout similarity can be obtained based on the number of layout features of the subtitles in the candidate slice and the number of layout features of the subtitles in the playback slice. Then, the semantic similarity, keyword similarity, and subtitle layout similarity are weighted to obtain audio similarity.

[0120] The text of the candidate slice can be obtained by speech recognition of the candidate slice, or it can be the subtitle text of the candidate slice.

[0121] The text in the playback segment is obtained through speech recognition of the playback segment, or it is the subtitle text of the playback segment.

[0122] Semantic similarity can be used to represent the degree of semantic similarity between the text of the candidate slice and the text of the playing slice.

[0123] Keyword similarity can be used to represent the degree of similarity between the text of the candidate slice and the text of the playing slice in terms of keywords.

[0124] Among them, the subtitle layout similarity can be used to indicate the degree of similarity in layout between the subtitles of the candidate slice and the subtitles of the playback slice.

[0125] As an example, semantic similarity can be calculated using the following formula (8): (8) in, Indicates semantic similarity. This describes a method for calculating text similarity using a pre-trained language model. The text indicating the playback slice, The text representing the candidate slice.

[0126] As an example, keyword similarity can be calculated using the following formula (9): (9) in, Indicates keyword similarity. This represents the set of keywords shared by the text of the candidate slice and the text of the played slice. This represents the union of all keywords in the text of the candidate slice and the text of the played slice. Wherein, The larger the value, the more keywords the two texts share, and the higher the keyword similarity.

[0127] As an example, the following formula (10) can be used to calculate the similarity of subtitle layouts: (10) in, This indicates that the subtitle layouts are similar. This indicates the number of typographic features (such as font, size, position, etc.) that match the subtitles of the candidate slice with the subtitles of the playback slice. This indicates the total number of typesetting features.

[0128] For example, the aforementioned image similarity, audio similarity, and text similarity can be weighted to obtain a comprehensive similarity score. The weighting values ​​can be determined according to actual needs and are not limited thereto.

[0129] Step 306: Determine the target slice from at least one candidate slice based on the comprehensive similarity.

[0130] In this application, the candidate slice with the highest overall similarity (greater than the similarity threshold) can be identified as the target slice.

[0131] Step 307: Determine the start playback time point based on the timestamp corresponding to the target slice, so as to play the second video on the first device from the start playback time point.

[0132] For example, if the length of the target slice is less than a preset threshold, the timestamp corresponding to the target slice can be directly used as the start playback time point; if the length of the target slice is greater than the preset threshold, any time point in the target slice can be determined as the start playback time point.

[0133] In this embodiment, multi-dimensional features are extracted from both the candidate slice and the playback slice in the first video. Based on the extracted multi-dimensional features, the comprehensive similarity between the candidate slice and the playback slice is calculated. Based on the comprehensive similarity, the target slice is determined, and the continuation position of the second video is determined based on the target slice. Therefore, by capturing subtle differences between the candidate slice and the playback slice in multiple dimensions, the similarity between the two slices can be more comprehensively evaluated, improving the accuracy of the target slice and accurately locating the playback node for cross-screen continuation. This ensures the consistency and integrity of content during cross-screen continuation, meeting users' needs for a high-quality viewing experience. Furthermore, determining the continuation position based on multi-modal comparison ensures a high degree of correspondence with the original device's playback time, enabling seamless transitions when users resume playback on a new device, avoiding plot jumps or repeated playback, and enhancing the continuity of the viewing experience.

[0134] To achieve the above embodiments, this application also proposes a cross-screen playback device. Figure 4 This is a schematic diagram of a cross-screen playback device provided in an embodiment of this application.

[0135] like Figure 4 As shown, the cross-screen playback continuation device 400 includes: The first determining module 410 is configured to respond to receiving a video playback request sent by a first device, wherein the video playback request includes user behavior data of a first video, and determine a second video to be resumed based on the user behavior data; wherein the first video is a video that has stopped playing on the second device, and the second video is a different version from the first video; The second determining module 420 is used to determine the event density of the second video and determine the slice length based on the event density and the copyright information of the second video; wherein, the event density refers to the number of events occurring per unit time. The third determining module 430 is used to determine at least one candidate slice from the second video based on the slice length and the stop playback time of the first video. The fourth determining module 440 is configured to determine the start playback time point of the second video based on the at least one candidate slice, so as to play the second video on the first device from the start playback time point.

[0136] Furthermore, in one possible implementation of this application embodiment, the second determining module 420 is used for: Based on the copyright information, a copyright coefficient is determined; wherein the copyright coefficient is used to reflect the degree of modification of the content by the copyright holder of the second video; The basic time unit is adjusted based on the event density and the copyright coefficient to obtain the slice length.

[0137] Furthermore, in one possible implementation of this application embodiment, the second determining module 420 is used for: The slice adjustment coefficient is determined based on the event density and the copyright coefficient. The base time unit is adjusted according to the slice adjustment coefficient to obtain the slice length.

[0138] Furthermore, in one possible implementation of this application embodiment, the basic time unit is determined in the following manner; The initial values ​​are adjusted based on one or more of the following factors: rhythm information, plot information, event density, shot switching frequency, and user viewing habits, to obtain the basic time unit.

[0139] Furthermore, in one possible implementation of this application embodiment, the fourth determining module 440 is used for: Multi-dimensional features are extracted from the candidate slices and the playback slices in the first video respectively to obtain the first multi-dimensional features and the second multi-dimensional features; wherein, the playback slice refers to the video segment in the first video between the target time point and the stop playback time point, and the target time point is less than the stop playback time point; Based on the first multi-dimensional feature and the second multi-dimensional feature, the comprehensive similarity between the candidate slice and the playback slice is determined; Based on the comprehensive similarity, the target slice is determined from the at least one candidate slice; The start playback time is determined based on the timestamp corresponding to the target slice.

[0140] Furthermore, in one possible implementation of this application embodiment, the fourth determining module 440 is used for: Based on the image features in the first multi-dimensional features and the image features in the second multi-dimensional features, the image similarity between the candidate slice and the playback slice is determined. Based on the audio features in the first multi-dimensional features and the audio features in the second multi-dimensional features, the audio similarity between the candidate slice and the playback slice is determined. Based on the text features in the first multi-dimensional features and the text features in the second multi-dimensional features, the text similarity between the candidate slice and the playback slice is determined. The overall similarity is determined based on the image similarity, the audio similarity, and the text similarity.

[0141] Furthermore, in one possible implementation of this application embodiment, the third determining module 430 is used for: Based on the time point at which playback stopped and the metadata of the second video, the estimated time position is determined; In the second video, multiple candidate slices with a length equal to the slice length are extended forward and backward from the estimated time position as the center point.

[0142] It should be noted that the foregoing explanation of the cross-screen playback method embodiment also applies to the cross-screen playback device of this embodiment, and will not be repeated here.

[0143] In this embodiment, if a video playback request is received from a first device, and the video playback request includes user behavior data of the first video, a second video to be resumed is determined based on the user behavior data. The slice length is dynamically determined based on the event density and copyright information of the second video, ensuring that the slice length matches the content characteristics. Based on the slice length and the stop playback time of the first video, candidate slices are determined from the second video, improving the accuracy of the candidate slices. Thus, based on the candidate slices, the start playback time of the second video is determined, improving the accuracy of the resume playback position. This ensures that users can seamlessly and accurately continue watching content on new devices, improving the smoothness of cross-screen resume playback and achieving a silky smooth cross-screen resume playback experience.

[0144] The cross-screen insurance renewal method of this application embodiment has the following advantages: (1) Improve the accuracy of media asset matching: Through preliminary screening by large models and comparison of multimodal slice detection, it is possible to find videos that are highly similar to the original viewing content more accurately in video libraries with different copyrights, reduce misjudgments and omissions, and improve the success rate of cross-screen playback. (2) Ensure the accuracy of playback nodes: The playback nodes determined based on multimodal comparison can correspond closely to the playback time of the original device, enabling users to seamlessly continue playback on the new device, avoiding plot jumps or repeated playback, and enhancing the continuity of the viewing experience. (3) Enhanced sensitivity to content differences: The multimodal detection function can capture subtle differences in the picture, audio and text of different versions of video, thereby more comprehensively assessing the similarity of media assets, ensuring the consistency and integrity of content when playing back across screens, and meeting users' needs for a high-quality viewing experience.

[0145] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments. To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0146] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0147] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0148] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0149] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0150] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0151] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0152] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0153] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0154] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0155] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.

[0156] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0157] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for cross-screen playback continuation, characterized in that, include: In response to receiving a video playback request sent by a first device, wherein the video playback request includes user behavior data of a first video, a second video to be resumed is determined based on the user behavior data; wherein the first video is a video that has stopped playing on the second device, and the second video is a different version from the first video; The event density of the second video is determined, and the slice length is determined based on the event density and the copyright information of the second video; wherein, the event density refers to the number of events occurring per unit time. Based on the slice length and the stop playback time of the first video, at least one candidate slice is determined from the second video; Based on the at least one candidate slice, the start time point of the second video is determined so that the second video is played on the first device from the start time point.

2. The method as described in claim 1, characterized in that, Determining the slice length based on the event density and the copyright information of the second video includes: Based on the copyright information, a copyright coefficient is determined; wherein the copyright coefficient is used to reflect the degree of modification of the content by the copyright holder of the second video; The basic time unit is adjusted based on the event density and the copyright coefficient to obtain the slice length.

3. The method as described in claim 2, characterized in that, The step of adjusting the base time unit based on the event density and the copyright coefficient to obtain the slice length includes: The slice adjustment coefficient is determined based on the event density and the copyright coefficient. The base time unit is adjusted according to the slice adjustment coefficient to obtain the slice length.

4. The method as described in claim 2, characterized in that, The basic time unit is determined in the following manner; The initial values ​​are adjusted based on one or more of the following factors: rhythm information, plot information, event density, shot switching frequency, and user viewing habits, to obtain the basic time unit.

5. The method as described in claim 1, characterized in that, Determining the start playback time of the second video based on the at least one candidate slice includes: Multi-dimensional features are extracted from the candidate slices and the playback slices in the first video respectively to obtain the first multi-dimensional features and the second multi-dimensional features; wherein, the playback slice refers to the video segment in the first video between the target time point and the stop playback time point, and the target time point is less than the stop playback time point; Based on the first multi-dimensional feature and the second multi-dimensional feature, the comprehensive similarity between the candidate slice and the playback slice is determined; Based on the comprehensive similarity, the target slice is determined from the at least one candidate slice; The start playback time is determined based on the timestamp corresponding to the target slice.

6. The method as described in claim 5, characterized in that, The step of determining the comprehensive similarity between the candidate slice and the playback slice based on the first multi-dimensional feature and the second multi-dimensional feature includes: Based on the image features in the first multi-dimensional features and the image features in the second multi-dimensional features, the image similarity between the candidate slice and the playback slice is determined. Based on the audio features in the first multi-dimensional features and the audio features in the second multi-dimensional features, the audio similarity between the candidate slice and the playback slice is determined. Based on the text features in the first multi-dimensional features and the text features in the second multi-dimensional features, the text similarity between the candidate slice and the playback slice is determined. The overall similarity is determined based on the image similarity, the audio similarity, and the text similarity.

7. The method as described in claim 1, characterized in that, The step of determining at least one candidate slice from the second video based on the slice length and the stop playback time of the first video includes: Based on the time point at which playback stopped and the metadata of the second video, the estimated time position is determined; In the second video, multiple candidate slices with a length equal to the slice length are extended forward and backward from the estimated time position as the center point.

8. A cross-screen playback continuation device, characterized in that, include: A first determining module is configured to respond to receiving a video playback request sent by a first device, wherein the video playback request includes user behavior data of a first video, and determine a second video to be resumed playback based on the user behavior data; wherein the first video is a video that has stopped playing on the second device, and the second video is a different version from the first video; The second determining module is used to determine the event density of the second video and determine the slice length based on the event density and the copyright information of the second video; wherein, the event density refers to the number of events occurring per unit time. The third determining module is used to determine at least one candidate slice from the second video based on the slice length and the stop playback time of the first video. The fourth determining module is configured to determine the start playback time point of the second video based on the at least one candidate slice, so as to play the second video on the first device from the start playback time point.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.

11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-7.