Resource playing method and device, electronic equipment, medium and product

By automatically recognizing video plot content through electronic devices and adding advertising resources to candidate frames, and playing them in a picture-in-picture format, the high cost of embedding advertisements in videos is solved, achieving a natural and smooth integration and improved viewing experience.

CN121940385APending Publication Date: 2026-04-28HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-10-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

When embedding advertising resources into video resources, existing technologies require video creators to have a deep understanding of the plot, resulting in high manpower and time costs and affecting the video viewing experience.

Method used

The electronic device automatically identifies the plot content of the first media resource, matches candidate frames to add a second media resource, and plays it in a picture-in-picture format to reduce obstruction of the original subject and improve the viewing experience.

Benefits of technology

It reduces manpower and time costs, achieves a natural and smooth integration of advertising resources into video resources, and enhances the video viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940385A_ABST
    Figure CN121940385A_ABST
Patent Text Reader

Abstract

The invention discloses a resource playing method and apparatus, an electronic device, a medium and a product, relates to the technical field of Internet, and can reduce labor cost and time cost on the premise of ensuring natural fluency of media resource fusion. The method comprises the steps of displaying a first interface; receiving a first media resource and a second media resource input by a user on the first interface; and in response to the input first media resource and the second media resource, playing a target media resource, the target media resource being a media resource obtained after the second media resource is added in the first media resource based on the candidate frame, and the content semantics of the candidate frame being matched with the content semantics of the second media resource.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to a resource playback method, apparatus, electronic device, medium and product. Background Technology

[0002] With the rapid development and widespread adoption of the internet, the application of media resources such as video and image resources is becoming increasingly common. For example, in the advertising field, advertising videos or images can be embedded into video resources to monetize advertising based on video resources.

[0003] However, in order to seamlessly integrate advertising resources into video resources, video creators need to have a deep understanding of the video's plot and content in order to reasonably arrange the timing and frequency of advertising resources, which requires a high amount of manpower and time. Summary of the Invention

[0004] This application provides a resource playback method, apparatus, electronic device, medium, and product that can reduce labor and time costs while ensuring the natural and smooth integration of media resources.

[0005] To achieve the above objectives, the embodiments of this application provide the following technical solutions:

[0006] Firstly, a resource playback method is provided. This method can be executed by an electronic device, or by a component of the electronic device, such as the processor, chip, or chip system of the electronic device. It can also be implemented by a logic module or software that can realize all or part of the functions of the electronic device, such as the client or server of an application in the electronic device.

[0007] Using an electronic device as the executing entity, the method includes:

[0008] The electronic device displays a first interface. For example, the first interface is an interface for triggering the fusion of different media resources, such as an interface for triggering the embedding of advertising resources into video resources.

[0009] The electronic device receives a first media resource and a second media resource input by a user on a first interface. The first and second media resources input by the user may include first and second media resources uploaded by the user, or first and second media resources selected by the user. In this embodiment, the user refers to a video developer such as an advertiser or video creator.

[0010] The electronic device plays a target media resource in response to input first and second media resources. The target media resource is a media resource obtained by adding the second media resource to the first media resource based on candidate frames. The content semantics of the candidate frames match the content semantics of the second media resource. For example, if the target media resource is an image resource, it is a media resource obtained by adding the second media resource to the candidate frames of the first media resource. As another example, if the target media resource is a video resource, it is a media resource obtained by adding the second media resource to the first media resource with a candidate frame as the starting frame.

[0011] In the above technical solution, candidate frames that match the plot content of the second media resource can be automatically determined from the first media resource. This eliminates the need for video creators to manually understand the plot content of the first media resource to determine candidate frames, significantly reducing labor and time costs. Furthermore, by adding the second media resource to the first media resource based on candidate frames to obtain the target media resource, the second media resource can be added naturally and smoothly to the first media resource in a more efficient and economical way. This results in better integration between the first and second media resources, ensuring a natural and smooth fusion of media resources.

[0012] In conjunction with the first aspect above, in one possible implementation, the electronic device plays the target media resource, including:

[0013] The second media resource is played in a picture-in-picture format within the first media resource.

[0014] In the above implementation method, by playing the second media resource in a picture-in-picture format, it is possible to avoid obscuring the original subject in the first media resource, and reduce the impact on video viewers (such as consumer users) watching the first media resource, thereby improving the viewing experience of video viewers.

[0015] In conjunction with the first aspect above, in one possible implementation, before the electronic device plays the target media resource, the method further includes:

[0016] The content semantics of the first media resource and the content semantics of the second media resource are obtained. Then, based on the content semantics of the first media resource and the content semantics of the second media resource, candidate frames are obtained from the first media resource.

[0017] In the first implementation method, the candidate frames and the content semantics of the second media resource meet a similarity condition. That is, based on the content semantics of the first media resource and the content semantics of the second media resource, video frames from the first media resource that meet the content semantics similarity condition with the second media resource are selected as candidate frames.

[0018] In the second implementation method, the content semantics of the candidate frames include one or more of the following: entity type, scene type, expression type, or behavior type. That is, based on the content semantics of the first media resource and the content semantics of the second media resource, video frames from the first media resource whose content semantics include one or more of the following: entity type, scene type, expression type, or behavior type are selected as candidate frames.

[0019] The above implementation provides two methods for determining candidate frames. Method one determines candidate frames based on video plot content, such as points determined by the similarity between the plot content of the first media resource and the advertising content of the second media resource. Method two determines candidate frames based on preset templates, such as points determined based on the plot content of the first media resource, including one or more content types such as entity type, scene type, expression type, or behavior type.

[0020] In conjunction with the first aspect described above, in one possible implementation, the first media resource is a video resource, and the content semantics of the first media resource includes both audio and video content semantics. The second media resource is either a video resource or an image resource. When the second media resource is a video resource, its content semantics includes both audio and video content semantics; when the second media resource is an image resource, its content semantics includes only the image content semantics.

[0021] In conjunction with the first aspect described above, in one possible implementation, the process of acquiring the semantic content of audio includes:

[0022] Speaker recognition and speech recognition are performed on the audio to obtain the recognition results, which serve as the semantic content of the audio; or,

[0023] Speaker recognition and speech recognition are performed on the audio to obtain the recognition results. Keywords are extracted from the recognition results as the semantic content of the audio; or,

[0024] Speaker recognition is performed on the audio, and subtitle recognition is performed on the corresponding video frames to obtain the audio recognition result, which serves as the semantic content of the audio; or,

[0025] Speaker recognition is performed on the audio, and subtitle recognition is performed on the corresponding video frames to obtain the audio recognition results. Keywords are extracted from the audio recognition results as the semantic content of the audio.

[0026] The above implementation methods provide four ways to obtain the semantic content of audio. Among them, speaker recognition combined with speech recognition or speaker recognition combined with subtitle recognition can quickly determine the audio recognition result. Furthermore, the audio recognition result can be directly determined as the audio's semantic content, or keywords can be extracted from the audio recognition result to obtain the audio's semantic content. By extracting keywords, not only can effective semantic content be determined, but the processing load on electronic devices can also be reduced, helping to improve the processing efficiency of media resources.

[0027] In conjunction with the first aspect described above, in one possible implementation, the process of acquiring the semantic content of a video includes:

[0028] Perform at least one of entity recognition, scene recognition, face detection, and behavior detection on the video frames included in the video to obtain video frame recognition results, which serve as the content semantics of the video; the video frame recognition results include at least one of entity recognition results, scene recognition results, face detection results, and behavior detection results; or,

[0029] Extract keyframes from the video frames included in the video, and perform at least one of entity recognition, scene recognition, face detection, and behavior detection on the keyframes to obtain keyframe recognition results, which are used as the content semantics of the video; the keyframe recognition results include at least one of entity recognition results, scene recognition results, face detection results, and behavior detection results;

[0030] Among them, entity recognition results are used to indicate entity type, scene recognition results are used to indicate scene type, face detection results are used to indicate the expression type of a person, and behavior detection results are used to indicate the behavior type of an entity.

[0031] The above implementation provides two methods for acquiring the semantic content of videos. One method utilizes entity recognition, scene recognition, face detection, and behavior detection technologies to quickly determine video frame recognition results or keyframe recognition results. Extracting keyframes not only identifies valid semantic content but also reduces the processing load on electronic devices, thus improving the efficiency of media resource processing.

[0032] In conjunction with the first aspect described above, in one possible implementation, the process of acquiring the content semantics of an image includes:

[0033] Perform at least one of entity recognition, scene recognition, face detection, and behavior detection on an image to obtain an image recognition result, which serves as the semantic content of the image. The image recognition result includes at least one of entity recognition result, scene recognition result, face detection result, and behavior detection result. The entity recognition result is used to indicate the entity type, the scene recognition result is used to indicate the scene type, the face detection result is used to indicate the expression type of a person, and the behavior detection result is used to indicate the behavior type of the entity.

[0034] In conjunction with the first aspect above, in one possible implementation, the content semantics of the first media resource includes the content semantics of multiple shot segments. The process of obtaining the content semantics of the multiple shot segments includes: performing shot analysis on the first media resource to obtain multiple shot segments of the first media resource, and then performing content understanding on each of the multiple shot segments to obtain the content semantics of the multiple shot segments.

[0035] In the above implementation method, the first media resource can be divided into multiple shot fragments by shot analysis, and then the content semantics of multiple shot fragments can be obtained by content understanding.

[0036] In conjunction with the first aspect above, in one possible implementation, the content semantics of the first media resource includes the content semantics of the target shot segment, and the process of acquiring the content semantics of the target shot segment includes:

[0037] Keywords are extracted from the textual information of secondary media resources. Target shot segments containing keywords or synonyms are selected from multiple shot clips, provided the similarity between the synonyms and keywords reaches a similarity threshold. Subsequently, content understanding is performed on the target shot segments to obtain their semantic content.

[0038] In the above implementation method, by using the keywords of the second media resource to determine the target shot segments including the keywords or synonyms, it is possible to identify the shot segments in the first media resource that are related to the second media resource. This not only identifies effective shot segments but also reduces the processing content of electronic devices, which helps to improve the processing efficiency of media resources.

[0039] In conjunction with the first aspect described above, in one possible implementation, the electronic device obtains candidate frames from the first media resource, including:

[0040] Based on the content semantics of the shot fragments in the first media resource and the content semantics of the second media resource, candidate shot fragments are determined from the first media resource. The content similarity between the candidate shot fragments and the second media resource satisfies a similarity condition. That is, by determining the content similarity between each shot fragment and the second media resource, candidate shot fragments whose content similarity satisfies the similarity condition can be identified from the first media resource. Content similarity measures the degree of similarity between two content semantics (feature vectors), and can be cosine similarity, Jaccard similarity, Manhattan distance, Hamming distance, Mahalanobis distance, or other similarity metrics.

[0041] Then, candidate frames are obtained from the candidate shot segments. A candidate frame is either the last frame of a candidate shot segment or the first frame of the next shot segment. In other words, the last frame of a candidate shot segment or the first frame of the next shot segment is determined as the candidate frame.

[0042] In the above implementation, a method for determining candidate frames based on the content semantics of shot fragments is provided. This method first determines candidate shot fragments that meet the similarity conditions, and then determines candidate frames from these candidate shot fragments, so that the determined candidate frames are candidate frames that match the plot content of the second media resource.

[0043] In conjunction with the first aspect described above, in one possible implementation, the electronic device obtains candidate frames from the first media resource, including:

[0044] Clustering is performed on consecutive shots in the primary media resource to obtain at least one scene segment.

[0045] Based on the content semantics of at least one scene fragment in the first media resource and the content semantics of the second media resource, candidate scene fragments are determined from the at least one scene fragment. The content similarity between the candidate scene fragments and the second media resource satisfies a similarity condition. That is, by determining the content similarity between at least one scene fragment and the second media resource, candidate scene fragments whose content similarity satisfies the similarity condition are determined from the at least one scene fragment.

[0046] Based on the content semantics of the shot segments included in the candidate scene segment and the content semantics of the second media resource, candidate shot segments are determined from the shot segments included in the candidate scene segment. The content similarity between the candidate shot segments and the second media resource satisfies a similarity condition. In other words, by determining the content similarity between the shot segments included in the candidate scene segment and the second media resource, candidate shot segments whose content similarity satisfies the similarity condition can be determined from the shot segments included in the candidate scene segment.

[0047] Then, candidate frames are obtained from the candidate shot segments. A candidate frame is either the last frame of a candidate shot segment or the first frame of the next shot segment. In other words, the last frame of a candidate shot segment or the first frame of the next shot segment is determined as the candidate frame.

[0048] In the above implementation, a method is provided to determine candidate frames based on the content semantics of scene fragments and shot fragments. This method first determines candidate scene fragments that meet the similarity conditions, then determines candidate shot fragments from the candidate scene fragments, and then determines candidate frames from the candidate shot fragments, so that the determined candidate frames are candidate frames that match the plot content of the second media resource.

[0049] In conjunction with the first aspect above, in one possible implementation, the electronic device determines the content similarity between at least one scene segment and a second media resource, including:

[0050] Based on the content semantics of the shot segments included in at least one scene segment, the content semantics of at least one scene segment are determined. Based on the content semantics of at least one scene segment and the content semantics of the second media resource, a first content similarity is determined between each of the at least one scene segment and the second media resource.

[0051] Based on the content semantics of the shot segments included in at least one scene segment, a text summary of at least one scene segment is generated; based on the text summary of at least one scene segment and the text information of the second media resource, a second content similarity between at least one scene segment and the second media resource is determined.

[0052] Based on the first content similarity and the second content similarity, determine the content similarity between at least one scene segment and the second media resource.

[0053] The above implementation provides a way to determine the content similarity between a scene segment and a second media resource. This method not only considers the content semantics of the scene segment and the content semantics of the second media resource, but also the text summary of the scene segment and the text information of the second media resource. This increases the amount of information considered in determining the content similarity and improves the accuracy of determining the content similarity.

[0054] In conjunction with the first aspect above, in one possible implementation, before the electronic device plays the target media resource, the method further includes:

[0055] Candidate frames of the first media resource are displayed. At this point, the user can confirm whether the candidate frames of the first media resource are suitable and perform a feedback operation. The electronic device then receives the user's feedback operation on the candidate frames. This feedback operation instructs the user to confirm the candidate frames or the video frames corrected by the user.

[0056] Accordingly, the target media resource is the media resource obtained by adding a second media resource to the first media resource based on the target frame. Specifically, when the feedback operation is used to instruct the user to confirm a candidate frame, the target frame is the candidate frame; when the feedback operation is used to instruct the user to correct the video frame, the target frame is the video frame corrected by the user.

[0057] In the above implementation, by displaying candidate frames of the first media resource to the user, so that the user can choose whether to optimize or correct the candidate frames, the flexibility of media resource processing is improved, and more personalized target media resources can be generated based on user feedback.

[0058] In conjunction with the first aspect above, in one possible implementation, the number of candidate frames is multiple, and the electronic device displays candidate frames of the first media resource, including: displaying multiple candidate frames in descending order of the degree of matching between each candidate frame and the second media resource.

[0059] The above implementation provides a method for displaying multiple candidate frames based on their matching degree. Displaying candidate frames in descending order of their matching degree with the second media resource allows users to quickly identify the most suitable candidate frames, improving human-computer interaction efficiency.

[0060] In conjunction with the first aspect described above, in one possible implementation, the method further includes:

[0061] The text of the connecting words is displayed in the first media resource. The text of the connecting words is matched with the semantic content of the first media resource and the semantic content of the second media resource.

[0062] The above implementation provides a method for displaying transition text. This transition text matches the semantic content of both the first and second media resources. In other words, the transition text connects not only to the plot content of the first media resource but also to the plot content of the second media resource, thus improving the playback effect of the target media resource.

[0063] In conjunction with the first aspect described above, in one possible implementation, the method further includes:

[0064] The electronic device generates the connecting text of the second media resource based on the content semantics of the candidate frame, the shot segments before the candidate frame, the shot segments after the candidate frame, and the content semantics of the second media resource.

[0065] In conjunction with the first aspect above, in one possible implementation, the electronic device generates transition text for the second media resource based on the content semantics of the candidate frame, the shot segments preceding the candidate frame, the shot segments following the candidate frame, and the content semantics of the second media resource, including:

[0066] Based on the semantic content of the shot segments before and after the candidate frames, plot summary texts before and after the candidate frames are generated.

[0067] Furthermore, based on the plot summary text before the candidate frame, the plot summary text after the candidate frame, the content semantics of the candidate frame, and the content semantics of the second media resource, the connecting word text of the second media resource is generated.

[0068] In the above implementation, a method for generating transition text for a second media resource is provided. Specifically, by generating plot summary text before and after the candidate frame, and using the determined plot summary text before and after the candidate frame, transition text for the second media resource is generated. This method can generate transition text that is relevant to both the plot content of the first and second media resources.

[0069] In conjunction with the first aspect above, in one possible implementation, before the electronic device plays the target media resource, the method further includes:

[0070] The connecting text of the second media resource is displayed. At this time, the user can confirm whether the connecting text of the second media resource is appropriate and perform a feedback operation. Subsequently, the electronic device receives the user's feedback operation on the connecting text. The feedback operation is used by the user to confirm the connecting text or the text corrected by the user.

[0071] Accordingly, the electronic device displays the linking word text in the first media resource, including: displaying the target text in the first media resource based on a feedback operation. Wherein, if the feedback operation is used to instruct the user to confirm the linking word text, the target text is the linking word text; if the feedback operation is used to instruct the user to correct the text, the target text is the text corrected by the user.

[0072] In the above implementation, by displaying the linking text of the second media resource to the user, so that the user can choose whether to optimize or correct the linking text, the flexibility of media resource processing is improved, and more personalized target media resources can be generated based on user feedback.

[0073] In conjunction with the first aspect above, in one possible implementation, the number of connector texts is multiple, and the electronic device displays the connector texts of the second media resource, including: displaying multiple connector texts in descending order of the degree of matching between each connector text and the content semantics of the first media resource and the content semantics of the second media resource.

[0074] Among the above implementation methods, a way to display multiple linking word texts based on matching degree is provided. Specifically, displaying the linking word texts in descending order of matching degree with the content semantics of the first media resource and the second media resource allows users to quickly identify the more matching linking word texts, improving human-computer interaction efficiency.

[0075] In conjunction with the first aspect above, in one possible implementation, the electronic device plays the target media resource, including:

[0076] Within the candidate resource area of ​​the first media resource, the second media resource is played according to its aspect ratio. The candidate resource area is the area within the first media resource used to display the second media resource.

[0077] In conjunction with the first aspect described above, in one possible implementation, the method further includes:

[0078] Foreground detection is performed on the candidate video frames to obtain the foreground and background regions of the video frames.

[0079] Furthermore, based on the foreground and background regions of the video frame, the frame size of the video frame, the edge regions of the video frame, and the frame size of the second media resource, the frame scaling ratio and candidate resource regions are determined.

[0080] The above implementation provides a way to determine the image scaling ratio and candidate resource area, which can quickly and efficiently determine the image scaling ratio and candidate resource area.

[0081] In conjunction with the first aspect described above, in one possible implementation, determining the aspect ratio and candidate resource region based on the foreground and background regions of the video frame, the aspect ratio of the video frame, the edge regions of the video frame, and the aspect ratio of the second media resource includes:

[0082] The optimization objective is to maximize the aspect ratio of the second media resource, while the constraint is that the second media resource does not touch the edge area or the foreground area of ​​the video frame. Based on the foreground and background areas of the video frame, the aspect ratio of the video frame, the edge area of ​​the video frame, and the aspect ratio of the second media resource, the aspect ratio and candidate resource areas are determined.

[0083] In the above implementation method, the scaling ratio and candidate resource area that maximize the image size of the second media resource can be determined to ensure the playback effect of the subsequent second media resource.

[0084] In conjunction with the first aspect above, in one possible implementation, before the electronic device plays the target media resource, the method further includes:

[0085] The candidate resource area for the second media resource is displayed. At this point, the user can confirm whether the candidate resource area is suitable and perform a feedback operation. The electronic device then receives the user's feedback operation on the candidate resource area. This feedback operation is used by the user to confirm the candidate resource area or the area corrected by the user.

[0086] Accordingly, the electronic device displays the second media resource in the candidate resource area of ​​the first media resource according to the aspect ratio of the second media resource, including: based on a feedback operation, displaying the second media resource in the target area of ​​the first media resource according to the aspect ratio of the second media resource. Wherein, when the feedback operation is used to instruct the user to confirm the candidate resource area, the target area is the candidate resource area; when the feedback operation is used to instruct the user to correct the area, the target area is the area corrected by the user.

[0087] In the above implementation, by displaying candidate resource areas of the second media resource to the user, so that the user can choose whether to optimize or modify the candidate resource area, the flexibility of media resource processing is improved, and more personalized target media resources can be generated based on user feedback.

[0088] In conjunction with the first aspect above, in one possible implementation, the number of candidate resource regions is multiple, and the electronic device displays candidate resource regions of the second media resource, including: displaying multiple candidate resource regions in descending order of their frame size.

[0089] Among the above implementation methods, a way to display multiple candidate resource areas based on the image size is provided. Displaying the candidate resource areas in descending order of image size allows users to quickly identify the larger candidate resource areas, improving human-computer interaction efficiency.

[0090] In conjunction with the first aspect described above, in one possible implementation, the method further includes:

[0091] Play the target audio from the second media resource within the first media resource.

[0092] In the above implementation, in addition to playing the second media resource, the target audio of the second media resource is also played, which improves the playback effect of the second media resource.

[0093] In conjunction with the first aspect described above, in one possible implementation, the method further includes:

[0094] Method 1: Extract audio features from an audio segment of a preset duration preceding the candidate frame; input the audio content of the second media resource and the audio features of the audio segment into a speech synthesizer to obtain the target audio of the second media resource. The audio features of the target audio are identical to those of the audio segment. Thus, by simulating the original audio of the first media resource and generating target audio with the same audio features as the original, it ensures consistency in video style between the first and second media resources, and enhances the auditory integration of the target media resource.

[0095] Method 2: Extract the audio features of the preset audio. Input the audio content of the second media resource and the audio features of the preset audio into a speech synthesizer to obtain the target audio of the second media resource. The audio features of the target audio are identical to those of the preset audio. Thus, by simulating the preset audio, a target audio with the same audio features as the preset audio is generated, enabling the generation of audio in the desired style and achieving personalized generation of the target audio.

[0096] In conjunction with the first aspect above, in one possible implementation, before the electronic device plays the target media resource, the method further includes:

[0097] The first and second media resources are validated for legality. If both the first and second media resources pass the validation, the process of generating the target media resource is executed.

[0098] In the above implementation method, the security of media resource processing can be guaranteed through legality verification.

[0099] Secondly, a resource playback method is provided. This method can be executed by an electronic device, or by a component of the electronic device, such as the processor, chip, or chip system of the electronic device. It can also be implemented by a logic module or software that can realize all or part of the functions of the electronic device, such as the client or server of an application in the electronic device.

[0100] Using an electronic device as the executing entity, the method includes:

[0101] The electronic device displays an identifier for the target video resource. For example, the electronic device displays the identifier for the target video resource on a second interface. For instance, the second interface could be a video playback interface, etc.

[0102] The target video resource includes a first segment and a second segment, where the second segment is either the segment preceding or following the first segment. This means that the first and second segments are two adjacent segments. It is understood that the segments can be video clips. The first segment embeds video footage used to showcase product features, which match the plot of either the first or second segment. For example, the video can be a moving video or a still image carrying audio.

[0103] The electronic device receives a user's playback operation on the identifier and, in response, plays the target video resource, wherein the first segment is played synchronously with the embedded video. For example, the playback operation can be a trigger operation on the identifier, such as a single click or double click. Alternatively, the identifier may include a playback control, and the playback operation can be a trigger operation on the playback control. In this embodiment, the user refers to a video viewer, such as a consumer.

[0104] In the above technical solution, while playing the first segment of the target film and television resource, the video embedded in the first segment can also be played simultaneously. On the one hand, since the product characteristics displayed in the video match the plot of the first or second segment (i.e., the segment adjacent to the first segment), the smoothness of the playback of the target film and television resource can be ensured, avoiding a disjointed feeling for the user when watching the film and television resource. On the other hand, by embedding the video in the first segment and playing the first segment and the video simultaneously, it can be ensured that the first segment and the embedded video play synchronously, so that the user can watch the embedded video without affecting the viewing of the first segment, thereby improving the viewing experience of the video viewer.

[0105] In conjunction with the second aspect described above, in one possible implementation, a video is embedded in the background area of ​​the first segment. For example, when playing the first segment, the embedded video is played synchronously in a picture-in-picture format. Here, the background area refers to the portion of the first segment's frame excluding the foreground area (i.e., the region of interest). Thus, by embedding the video in the background area of ​​the first segment, obstruction of the foreground area of ​​the first segment is avoided, allowing users to watch the embedded video without affecting their viewing of the first segment, thereby improving the viewing experience.

[0106] In conjunction with the second aspect described above, in one possible implementation, the method further includes: displaying target text, which is used to introduce product features in conjunction with the plot of the first or second segment. That is, the content of the target text not only matches the plot of the first or second segment but also matches the product features, enabling the introduction of product features while connecting the first or second segment, effectively enhancing the viewing experience for video viewers.

[0107] In conjunction with the second aspect described above, in one possible implementation, the method further includes: playing audio corresponding to the target text. The audio features of this audio are the same as those of the first or second segment. Thus, by playing audio with the same audio features as the first or second segment, the smoothness of the target video resource in terms of listening experience can be ensured. Alternatively, the audio features of this audio are the same as those of preset audio. This allows for the generation of audio with a desired style, achieving personalized audio generation.

[0108] Thirdly, a resource playback device is provided for implementing any of the methods provided in the first aspect. This resource playback device includes modules, units, or means corresponding to the aforementioned methods. The actions performed by these modules, units, or means can be implemented in hardware, software, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the aforementioned functions.

[0109] In one possible implementation, the device may include a display module, an input module, and a playback module; wherein:

[0110] The display module is used to display the first interface.

[0111] The input module is used to receive the first media resource and the second media resource input by the user on the first interface;

[0112] The playback module is used to play a target media resource in response to the input first media resource and second media resource. The target media resource is a media resource obtained by adding the second media resource to the first media resource based on candidate frames. The content semantics of the candidate frames match the content semantics of the second media resource.

[0113] Fourthly, a resource playback device is provided for implementing any of the methods provided in the second aspect above. The resource playback device includes modules, units, or means that implement the methods described above. The actions performed by these modules, units, or means can be implemented in hardware, software, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the functions described above.

[0114] In one possible implementation, the device may include a display module, a receiving module, and a playback module; wherein:

[0115] The display module is used to display the identifier of the target film and television resource. The target film and television resource includes a first segment and a second segment. The first segment contains embedded video, which is used to showcase product features. The product features match the plot of the first or second segment. The second segment is the segment before or after the first segment.

[0116] The receiving module is used to receive the user's playback operation for the identifier;

[0117] The playback module is used to respond to playback operations and play the target video resource, wherein the first segment is played synchronously with the embedded video.

[0118] Fifthly, an electronic device is provided, comprising: a memory and a processor, the memory and the processor being connected; the memory being used to store computer-executed instructions; and the processor being used to invoke the computer-executed instructions to implement the methods of the first aspect, the second aspect, or any implementation thereof described above.

[0119] The electronic device in the fifth aspect can be: the electronic device described in the first aspect, the second aspect, or any implementation thereof, or a device containing the electronic device, or a device included in the electronic device, such as a chip. For example, the electronic device can be a terminal or a server.

[0120] In a sixth aspect, a chip is provided, comprising: a processor and an interface circuit; the interface circuit being configured to receive computer execution instructions and transmit them to the processor; and the processor being configured to execute the computer execution instructions to perform the methods of the first aspect, the second aspect, or any implementation thereof described above.

[0121] In a seventh aspect, a computer-readable storage medium is provided, comprising computer-executable instructions that, when executed on an electronic device, cause the electronic device to perform the methods described in the first aspect, the second aspect, or any implementation thereof.

[0122] Eighthly, a computer program product is provided, including computer execution instructions that, when executed on an electronic device, cause the electronic device to perform the first aspect, the second aspect, or any implementation thereof described above.

[0123] The technical effects of any of the implementation methods in aspects two through eight can be found in the technical effects of the corresponding implementation methods in aspect one, and will not be repeated here.

[0124] All possible implementations of any of the above aspects can be combined, provided that the solutions do not contradict each other. Attached Figure Description

[0125] Figure 1 A schematic diagram of the system architecture for a resource playback method provided in an embodiment of this application;

[0126] Figure 2 A schematic diagram of the module division of a system architecture provided for an embodiment of this application;

[0127] Figure 3 A schematic diagram illustrating a media resource processing flow provided in an embodiment of this application;

[0128] Figure 4 This is a schematic diagram illustrating an example of media resource processing provided in an embodiment of this application.

[0129] Figure 5 A schematic diagram of the hardware structure of a terminal provided in an embodiment of this application;

[0130] Figure 6 A schematic diagram of the hardware structure of a server provided in an embodiment of this application;

[0131] Figure 7 A flowchart illustrating a resource playback method provided in an embodiment of this application;

[0132] Figure 8 A flowchart illustrating another resource playback method provided in an embodiment of this application;

[0133] Figure 9 A schematic diagram of an audio / video parsing process provided in an embodiment of this application;

[0134] Figure 10 A schematic diagram of a keyframe in a shot segment provided in an embodiment of this application;

[0135] Figure 11 A schematic diagram of a candidate resource area for a second media resource provided in an embodiment of this application;

[0136] Figure 12 A flowchart illustrating a resource playback method in an advertiser decision-making scenario provided in this application embodiment;

[0137] Figure 13 This is a schematic diagram illustrating an example of media resource embedding provided in an embodiment of this application.

[0138] Figure 14 A flowchart illustrating the module interaction in an advertiser decision-making scenario is provided in this application embodiment.

[0139] Figure 15 A flowchart illustrating a resource playback method in a video creator scenario provided in this application embodiment;

[0140] Figure 16 This application provides a schematic diagram of an interface display in a video creator scenario.

[0141] Figure 17 A flowchart illustrating the module interaction in a video creator scenario provided in this application embodiment;

[0142] Figure 18 A flowchart illustrating another resource playback method provided in an embodiment of this application;

[0143] Figure 19 This is a schematic diagram of the structure of a resource playback device provided in an embodiment of this application;

[0144] Figure 20 This is a schematic diagram of another resource playback device provided in an embodiment of this application. Detailed Implementation

[0145] In the description of this application, unless otherwise stated, "multiple" means two or more. At least one of the following or similar expressions refer to any combination of these terms, including any combination of single or plural terms. For example, at least one of a, b, and / or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0146] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.

[0147] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.

[0148] It is understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, throughout the specification, various embodiments do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0149] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Correspondingly, the apparatus given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.

[0150] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following embodiments of this application do not constitute a limitation on the scope of protection of this application.

[0151] The following provides an exemplary description of the application scenarios of the embodiments of this application.

[0152] With the rapid development and widespread adoption of the internet, the application of media resources such as video and image resources is becoming increasingly widespread. For example, in the advertising field, advertising videos or images can be embedded in video resources to achieve advertising monetization based on video resources. Video resources include long-form videos (such as TV dramas) and derivative videos (such as narration).

[0153] Currently, the common methods for embedding advertisements in video resources are: pre-roll, mid-roll, or overlay ads. However, pre-roll, mid-roll, or overlay ads are relatively direct and may negatively impact the viewing experience, potentially causing viewer aversion.

[0154] To seamlessly integrate advertising into video content without disrupting the viewing experience or causing viewer aversion, video creators need a deep understanding of the video's storyline to strategically time and frequency ad placement. However, while integrating ads into the narrative may seem natural, the creative design and content integration require meticulous planning, consuming significant human and time resources, and potentially involving substantial economic costs such as production difficulties.

[0155] Therefore, in the field of advertising, how to seamlessly and naturally integrate advertising resources into video resources in a more efficient and economical way is a challenging and practical issue.

[0156] This application provides two methods for embedding advertisements in related technologies, which are described below.

[0157] Related Technology 1: The process involves obtaining the initial ad placement selection instruction, marking the corresponding video frame of the initial ad placement location within the video segment as a standard frame, and using an image matching algorithm to calculate the spatial mapping relationship between all video frames and the standard frame in the video segment to obtain a first spatial transformation mapping matrix. Based on this matrix, the ad placement position of the advertisement to be inserted in each frame of the video segment is determined. A second spatial transformation mapping matrix of the advertisement to be inserted in the video segment is then calculated based on the ad placement position in each frame. Based on this matrix, the advertisement to be inserted is placed into the corresponding ad placement position in the video segment. Specifically, image matching technology is used to track and locate the ad placement position in the video segment, and the spatial transformation mapping matrix is ​​used to achieve the placement of the advertisement to be inserted into the located ad placement position in the video segment.

[0158] However, in the first related technology, it is still necessary to provide external instructions for selecting the initial ad placement location. In other words, users need to understand the video content and ad content themselves to determine the appropriate placement location, which requires high manpower and time costs.

[0159] Related Technology 2: Acquire the video to be processed, segment the video to obtain multiple video segments, determine the image frames to be identified from each video segment, and classify the pixels in the image frames to be identified to obtain the pixel types. Determine whether a target object exists in the image frames to be identified based on the pixel types. When a target object exists in the image frames to be identified, obtain the frame sequence segments containing the target object according to preset rules, and use the location of the target object in the frame sequence segments as the information implantation area.

[0160] However, in related technology two, recognition based solely on image frames in video segments can only ensure that the embedded advertising resources do not obscure the main content of the video, such as people, but cannot achieve the natural and smooth embedding of advertising resources in video resources.

[0161] In view of this, embodiments of this application provide a resource playback method that can be applied to scenarios involving advertising insertion, such as inserting advertisements into long videos or derivative videos. In some embodiments, embodiments of this application can be applied to advertiser decision-making scenarios, that is, assisting advertisers in deciding which video resources to target for advertising. For example, advertisers can use the resource playback method provided by embodiments of this application to preview the effects of advertising insertion on multiple different types of video resources, thereby assisting advertisers in determining their advertising placement choices. In other embodiments, embodiments of this application can be applied to video creator scenarios, that is, assisting video creators in inserting advertising resources into video resources. For example, video creators (such as video UP masters) can use the resource playback method provided by embodiments of this application to insert advertisements into a video resource, completing the process of producing new video content. The embodiments of this application will subsequently use these two application scenarios as examples to illustrate the solution. In the embodiments of this application, a first media resource is used to refer to the video resource to which the advertising resource is to be inserted, and a second media resource is used to refer to the advertising resource.

[0162] In this embodiment, candidate frames matching the plot content of the second media resource can be automatically determined from the first media resource, eliminating the need for video creators to manually understand the plot content of the first media resource to determine candidate frames, thus significantly reducing labor and time costs. Furthermore, by adding the second media resource to the first media resource based on candidate frames to obtain the target media resource, the second media resource can be added naturally and smoothly to the first media resource in a more efficient and economical manner, resulting in better integration between the first and second media resources and ensuring a natural and smooth fusion of media resources.

[0163] To facilitate understanding of the embodiments of this application, the following points will be explained before introducing the embodiments of this application.

[0164] 1. In the embodiments of this application, "instruction" can include direct instruction and indirect instruction, as well as explicit instruction and implicit instruction. The information indicated by a certain piece of information is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is a relationship between the other information and the information to be instructed. It can also indicate only a part of the information to be instructed, while the other parts of the information to be indicated are known or pre-agreed.

[0165] 2. "Pre-setting" can be achieved by pre-saving the corresponding code, table or other means that can be used to indicate relevant information in the device (e.g. electronic device). The embodiments of this application do not limit the specific implementation method.

[0166] 3. In the embodiments of this application, the descriptions such as "in the case of", "if" and "if" all refer to the fact that the device (e.g., electronic device) will make corresponding processing under certain objective circumstances. They are not time limits, nor do they require the device (e.g. electronic device) to have a judgment action when it is implemented, nor do they mean that there are other limitations.

[0167] Furthermore, the system architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0168] Furthermore, the actions, terms, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages used for interaction between devices in the embodiments of this application are merely examples, and other names may be used in specific implementations without limitation.

[0169] The system architecture of the embodiments of this application will be described below as an example.

[0170] In some embodiments, the resource playback method provided in this application can be applied to, for example... Figure 1 In the system architecture shown, for example, Figure 1 This is a schematic diagram of the system architecture for a resource playback method provided in an embodiment of this application. See also... Figure 1 The system architecture may include: Terminal 101.

[0171] The terminal 101 can be at least one of the following devices: smartphone, smartwatch, printer, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. In some embodiments, the terminal 101 has communication capabilities and can access wired or wireless networks.

[0172] In some embodiments, terminal 101 includes a development terminal and an application terminal.

[0173] The development terminal is used to execute the media resource processing procedure to obtain the target media resource. It is understood that the development terminal is also the operating terminal of the video developer, such as the advertiser or video creator. In this embodiment, for the above-mentioned scheme of executing the media resource processing procedure to obtain the target media resource, the development terminal can utilize the resource playback method provided in this embodiment to display the target media resource in response to the user inputting a first media resource and a second media resource on a first interface. The target media resource is the media resource obtained by adding a second media resource based on candidate frames to the first media resource. Here, the user refers to the video developer, such as the advertiser or video creator.

[0174] An application terminal is used to display and trigger the playback of target video resources. It is understood that the application terminal is also the operating terminal of the video viewer, such as a consumer user. Regarding the above-mentioned scheme for displaying and triggering the playback of target video resources, the application terminal can utilize the resource playback method provided in this application embodiment to display an identifier of the target video resource, receive a user's playback operation on the identifier, and, in response to the playback operation, play the target video resource. Here, "user" refers to a video viewer, such as a consumer user.

[0175] The system architecture may also include a server 102. In some embodiments, the terminal 101 may communicate with the server 102 via a wired or wireless network.

[0176] Server 102 can be a standalone physical server, a server cluster consisting of multiple physical servers, a distributed file system, or at least one of the following cloud servers providing basic cloud computing services: cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data or artificial intelligence platforms. This disclosure does not limit the scope of the embodiments. Of course, server 102 can also include other functions to provide more comprehensive and diversified services.

[0177] The above-described method for processing media resources to obtain the target media resource can also be accomplished by the development terminal and the server 102 working together. For example, in some embodiments, the development terminal sends a request to the server 102 to add a second media resource to the first media resource. In response to the request, the server 102 adds the second media resource to the first media resource based on candidate frames to obtain the target media resource, and returns the target media resource to the development terminal for display and playback.

[0178] Figure 2 This is a schematic diagram illustrating the modular division of a system architecture provided in an embodiment of this application. See also... Figure 2As shown in (2-1), when the development terminal performs the media resource processing to obtain the target media resource, the development terminal may include an interaction module, an audio / video parsing module, a content understanding module, an audio generation module, and an audio / video embedding module. See also Figure 2 As shown in (2-2), when the development terminal and the server cooperate to complete the media resource processing process to obtain the target media resource, the development terminal may include an interaction module, and the server may include an audio and video parsing module, a content understanding module, an audio generation module, and an audio and video embedding module.

[0179] The following example illustrates the modular division of the system architecture, using the example of a development terminal and server working together to process media resources and obtain the target media resources.

[0180] The interaction module provides an interface for interaction with users (i.e., video creators). In some embodiments, users interact with the interaction module to trigger the server to process media resources. It can be understood that the interaction module serves as the entry point for the entire system.

[0181] In some embodiments, the interaction module can be divided into a user interaction unit, a resource input / output unit, and a validity verification unit.

[0182] In some possible implementations, the user operates the user interaction unit to obtain the first media resource and the second media resource, and uploads the first media resource and the second media resource to the server to trigger the server to execute the media resource processing process.

[0183] In some possible implementations, after the server returns candidate frames (such as ad insertion points) in the first media resource, linking text (such as ad linking words) in the second media resource, and candidate resource areas (such as ad embedding areas) in the second media resource, the user can verify and process the results returned by the server by operating the resource input and output units, and further trigger the subsequent media resource processing flow.

[0184] In some possible implementations, a legality verification unit is used to verify the legality of the first and second media resources, such as format and length.

[0185] The following section introduces the audio and video parsing module, content understanding module, audio generation module, and audio and video embedding module included in the server.

[0186] The audio and video parsing module is used to perform preliminary parsing and processing of the first and second media resources uploaded by the development terminal for use by the subsequent content understanding module. In some embodiments, the audio and video parsing module can be divided into an audio and video separation unit, a speech recognition unit, and a video segmentation unit. The video segmentation unit is used to perform shot parsing on the first media resource to obtain multiple shot segments of the first media resource. The audio and video separation unit separates the audio track of each shot segment to obtain the audio of that shot segment. The speech recognition unit is used to obtain the semantic content of the audio of the multiple shot segments.

[0187] The content understanding module performs processing related to video content understanding. In some embodiments, the content understanding module can be divided into an ad placement identification unit, a linking word generation unit, and an implantation region identification unit. The ad placement identification unit identifies candidate frames in the first media resource. The linking word generation unit generates corresponding linking word text for each candidate frame. The implantation region identification unit identifies candidate resource regions in the first media resource.

[0188] The audio generation module is used to convert the linking text of the second media resource and the text information (such as advertising copy) of the second media resource into speech. In some embodiments, the audio generation module can be divided into a speech cloning unit and a speech synthesis unit. The speech cloning unit is used to clone the audio features of an audio segment or preset audio from the first media resource. The speech synthesis unit is used to convert the audio content and linking text of the second media resource into target audio with the same audio features as the audio segment or preset audio.

[0189] The audio and video embedding module is used to add a second media resource to a first media resource to obtain a target media resource. In some embodiments, the audio and video embedding module can be divided into an audio track embedding unit and a video embedding unit. The audio track embedding unit is used to embed the audio track of the second media resource into the audio track of the first media resource. The video embedding unit is used to embed the video frame of the second media resource into the video frame of the first media resource.

[0190] For example, Figure 3 This is a schematic diagram illustrating a media resource processing flow provided in an embodiment of this application. See also... Figure 3 The media resource processing workflow can include the following four steps.

[0191] Step ①: Uploading Media Resources. In one implementation, the user uploads a first media resource and a second media resource. For example, the user uploads the first and second media resources by operating on a development terminal. Alternatively, in another implementation, the database retrieves the first and second media resources. For example, the user selects the first and second media resources from the media resources stored in the database by operating on a development terminal.

[0192] Step ②, Video Content Understanding. Based on relevant services for video content understanding, output candidate frames from the first media resource (such as...). Figure 3 The advertising placement points shown), and the connecting text of the second media resources (such as...) Figure 3 The advertised link keywords shown) and the candidate resource areas of the second media resources (such as...) Figure 3 (The area where the advertisement is embedded is shown).

[0193] Step 3: Audio Generation. Based on speech cloning technology, the audio features of audio segments before and after candidate frames in the first media resource are cloned. Furthermore, based on speech synthesis technology, the connecting words and text information (such as advertising copy) of the second media resource are converted into target audio with the same audio features as the audio segments.

[0194] Step 4: Audio and Video Embedding. Audio Track Embedding: Embed the audio track of the second media resource into the audio track of the first media resource. For example, embed the converted target audio into the first media resource. Video Embedding: Embed the video frame of the second media resource into the video frame of the first media resource. For example, embed the video frame of the second media resource into the background area of ​​the first media resource in a picture-in-picture format. In this way, the second media resource can be added to the first media resource to obtain the target media resource.

[0195] For example, Figure 4 This is a schematic diagram illustrating an example of media resource processing provided in an embodiment of this application. See also... Figure 4 Based on the content understanding of the first and second media resources, it is possible to identify candidate frames in the first media resource that are strongly correlated with the plot content of the second media resource (such as...). Figure 4 The candidate frames of the first media resource shown can be represented by points on the timeline, and the text of connecting words in the second media resource that are strongly related to the plot content of the first media resource (such as...) is generated. Figure 4 The transition text shown can be represented by the duration of the transition text played on the timeline, to determine the candidate resource area in the first media resource for embedding the second media resource (e.g., Figure 4The candidate resource region is shown. Furthermore, target audio for the second media resource is generated using speech cloning and speech synthesis techniques. This target audio may include connective text and the audio content of the second media resource (such as...). Figure 4 The audio content of the second media resource shown can be represented by the duration of playback on the time track. Furthermore, by using the aforementioned candidate frames, connecting word text, candidate resource regions, and target audio, adding the second media resource to the first media resource allows the second media resource to be relatively naturally embedded into the first media resource, resulting in the target media resource and ensuring the fusion effect of the media resources.

[0196] In one example of this application, a schematic diagram of the terminal's hardware structure is shown below. Figure 5 As shown. Figure 5 This is a schematic diagram of the hardware structure of a terminal provided in an embodiment of this application.

[0197] See Figure 5 The terminal 500 may include a processor 510, an external memory interface 520, an internal memory 521, a universal serial bus (USB) interface 530, a charging management module 540, a power management module 541, a battery 542, antenna 1, antenna 2, a mobile communication module 550, a wireless communication module 560, an audio module 570, a sensor module 580, a camera 590, and a display screen 591. The sensor module 580 may include a pressure sensor 580A, a gyroscope sensor 580B, an accelerometer sensor 580C, a proximity sensor 580D, a touch sensor 580E, etc.

[0198] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the terminal 500. In other embodiments of this application, the terminal 500 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0199] Processor 510 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. The different processing units may be independent devices or integrated into one or more processors.

[0200] The controller can serve as the central nervous system and command center of the terminal 500. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.

[0201] The processor 510 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory can store instructions or data that the processor 510 has just used or that are used repeatedly. If the processor 510 needs to use the instruction or data again, it can retrieve it directly from this memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.

[0202] In some embodiments, the processor 510 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, etc.

[0203] USB port 530 is a USB standard compliant interface, which can be a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 530 can be used to connect a charger to charge terminal 500, and can also be used for data transfer between terminal 500 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0204] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal 500. In other embodiments of this application, the terminal 500 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.

[0205] The charging management module 540 receives charging input from a charger, which can be either a wireless or wired charger. The power management module 541 connects to the battery 542, the charging management module 540, and the processor 510. The power management module 541 receives input from the battery 542 and / or the charging management module 540, providing power to the processor 510, internal memory 521, external memory, display 591, camera 590, and wireless communication module 560, etc.

[0206] The wireless communication function of terminal 500 can be implemented through antenna 1, antenna 2, mobile communication module 550, wireless communication module 560, modem processor and baseband processor.

[0207] The mobile communication module 550 can provide wireless communication solutions including 2G / 3G / 4G / 5G for use on the terminal 500. The wireless communication module 560 can provide wireless communication solutions for use on the terminal 500 including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.

[0208] In some embodiments, the antenna 1 of the terminal 500 is coupled to the mobile communication module 550, and the antenna 2 is coupled to the wireless communication module 560, so that the terminal 500 can communicate with the network and other devices through wireless communication technology.

[0209] Terminal 500 implements display functions through a GPU, display screen 591, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 591 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 510 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0210] Display screen 591 is used to display images, videos, etc. Display screen 591 includes a display panel. In some embodiments, terminal 500 may include one or N displays screens 591, where N is a positive integer greater than 1.

[0211] Terminal 500 can perform shooting functions through an ISP, camera 590, video codec, GPU, display 591, and application processor. Camera 590 is used to capture still images or videos.

[0212] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in terminals, such as image recognition, facial recognition, speech recognition, and text understanding.

[0213] The external storage interface 520 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal 500. The external storage card communicates with the processor 510 through the external storage interface 520 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0214] Internal memory 521 can be used to store computer executable program code, which includes instructions. Processor 510 executes various functional applications and data processing of terminal 500 by running the instructions stored in internal memory 521.

[0215] Terminal 500 can implement audio functions, such as music playback and recording, through audio module 570 and application processor.

[0216] The pressure sensor 580A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 580A may be disposed on the display screen 591.

[0217] The gyroscope sensor 580B can be used to determine the motion attitude of the terminal 500. In some embodiments, the angular velocity of the terminal 500 about three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 580B.

[0218] The accelerometer 580C can detect the magnitude of acceleration of the terminal 500 in various directions (generally three axes). When the terminal 500 is stationary, it can detect the magnitude and direction of gravity.

[0219] A distance sensor 580D is used to measure distance. The terminal 500 can measure distance via infrared or laser. In some embodiments, during a shooting scene, the terminal 500 can utilize the distance sensor 580D to measure distance for rapid focusing.

[0220] Touch sensor 580E, also known as a "touch panel," can be located on display screen 591. The touch sensor 580E and display screen 591 together form a touchscreen, also known as a "touch screen." Touch sensor 580E detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 591. In other embodiments, touch sensor 580E may also be located on the surface of terminal 500, in a different position than display screen 591.

[0221] It should be pointed out that, Figure 5 The structure shown does not constitute a limitation on this terminal, except... Figure 5 In addition to the components shown, the terminal may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0222] In one example of this application, a schematic diagram of the server's hardware structure is shown below. Figure 6 As shown. Figure 6 This is a schematic diagram of the hardware structure of a server provided in an embodiment of this application.

[0223] See Figure 6 , Figure 6 The server shown may include a processor 601, a memory 602, a communication module 603, and a bus 604. The processor 601, the memory 602, and the communication module 603 can be connected via the bus 604.

[0224] The processor 601 is the control center of the server and can be a general-purpose central processing unit such as a CPU, or other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor. In this embodiment, the processor 601 in the server can execute the media resource processing procedure in the resource playback method.

[0225] As an example, processor 601 may include one or more CPUs, for example Figure 6 CPU 0 and CPU 1 are shown in the diagram.

[0226] The memory 602 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. In this embodiment, training data, such as first training data or second training data, may be stored in the memory 602 of the server.

[0227] In one possible implementation, the memory 602 may exist independently of the processor 601. The memory 602 can be connected to the processor 601 via a bus 604 and is used to store data, instructions, or program code. When the processor 601 calls and executes the instructions or program code stored in the memory 602, it can implement the media resource processing procedure in the resource playback method provided in this application embodiment.

[0228] In another possible implementation, the memory 602 can also be integrated with the processor 601.

[0229] The communication module 603 is used for connecting the server to other devices via a communication network, which can be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. The communication module 603 may include a receiving unit for receiving data and a transmitting unit for transmitting data.

[0230] Bus 604 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0231] It should be pointed out that, Figure 6 The structure shown does not constitute a limitation on the server, except... Figure 6 In addition to the components shown, the server may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0232] For ease of understanding, the resource playback method provided in the embodiments of this application will be described below with reference to the above system architecture and accompanying drawings.

[0233] It is understood that in the embodiments of this application, the terminal (development terminal or application terminal) or the server can execute some or all of the steps in the embodiments of this application. These steps or operations are only examples. The embodiments of this application can also perform other operations or variations of various operations.

[0234] Figure 7 This is a flowchart illustrating a resource playback method provided in an embodiment of this application. In some possible implementations, this resource playback method can be executed by the development terminal or server in the above system architecture, see [link to relevant documentation]. Figure 7 Using an electronic device as the execution subject, the method includes the following steps S701-S703.

[0235] S701, the electronic device displays the first interface.

[0236] In this embodiment, the first interface is used to trigger the fusion of different media resources, such as triggering the addition of a second media resource to a first media resource. For example, the first interface can be used to trigger the embedding of an advertising resource into a video resource. In this embodiment, the term "first media resource" will be used to refer to the video resource, and "second media resource" will be used to refer to the advertising resource.

[0237] S702, The electronic device receives the first media resource and the second media resource input by the user on the first interface.

[0238] The first media resource can be a video resource, such as a long video or a derivative video. The second media resource can be an advertising resource, such as the video resource of an advertising video or the image resource of an advertising image. In some embodiments, the embodiments of this application can be applied to scenarios where advertising videos are embedded in video resources, or in other embodiments, the embodiments of this application can be applied to scenarios where advertising images are embedded in video resources. The embodiments of this application are not limited in this respect.

[0239] The first and second media resources input by the user can include first and second media resources uploaded by the user, or first and second media resources selected by the user.

[0240] In some embodiments, the first interface includes a resource upload control, which allows users to upload a first media resource and a second media resource based on their operations on the resource upload control.

[0241] In other embodiments, the first interface includes multiple candidate resources and a resource selection control, and the user selects a first media resource and a second media resource from the multiple candidate resources based on the trigger operation of the resource selection control.

[0242] S703, the electronic device responds to the input first media resource and second media resource and plays the target media resource.

[0243] The target media resource is the media resource obtained by adding a second media resource based on the candidate frame to the first media resource.

[0244] In this embodiment, the content semantics of the candidate frame matches the content semantics of the second media resource. In some embodiments, the candidate frame can be a candidate frame based on the video plot content, such as a point determined based on the similarity between the plot content of the first media resource and the advertising content of the second media resource. In other embodiments, the candidate frame can be a candidate frame based on a preset template, such as a point determined based on a preset content type determined based on the plot content of the first media resource. For the process of determining the candidate frame, please refer to the following sections. Figure 8 The detailed process is shown below.

[0245] In some embodiments, the process of an electronic device playing a target media resource can be as follows: playing a second media resource in a picture-in-picture format within a first media resource. By playing the second media resource in a picture-in-picture format, obscuring the original subject in the first media resource can be avoided, and the impact on the video viewer's viewing of the first media resource is reduced, thus improving the video viewer's viewing experience.

[0246] The technical solution provided in this application can automatically determine candidate frames from the first media resource that match the plot content of the second media resource, eliminating the need for video creators to manually understand the plot content of the first media resource to determine candidate frames, thus greatly reducing labor and time costs. Furthermore, by adding the second media resource to the first media resource based on candidate frames to obtain the target media resource, the second media resource can be added naturally and smoothly to the first media resource in a more efficient and economical way, resulting in better integration of the first and second media resources and ensuring a natural and smooth fusion of media resources.

[0247] In the above Figure 7Based on this, at least one of the following can be generated: the transition text of the second media resource, the candidate resource region of the second media resource, and the target audio of the second media resource. Then, based on at least one of the transition text of the second media resource, the candidate resource region of the second media resource, and the target audio of the second media resource, the target media resource is generated and played.

[0248] The following is based on Figure 8 This section provides a detailed introduction to the resource playback method. Figure 8 This is a flowchart illustrating another resource playback method provided in an embodiment of this application. See also... Figure 8 Taking an electronic device as the execution subject as an example, the method includes the following steps S801-S810.

[0249] S801, the electronic device displays the first interface.

[0250] S802, The electronic device receives the first media resource and the second media resource input by the user on the first interface.

[0251] S803: The electronic device responds to the input first media resource and second media resource by performing a validity check on the first media resource and the second media resource. If both the first media resource and the second media resource pass the check, then S804 is executed.

[0252] In some embodiments, the electronic device verifies the legality of the first media resource and the second media resource by including at least one of the following verification items (1) to (4).

[0253] Verification item (1) Video format check.

[0254] In some embodiments, the electronic device acquires the video formats of the first media resource and the second media resource, and determines whether the video formats of the first media resource and the second media resource are video formats supported by the platform. If the video formats of both the first media resource and the second media resource are video formats supported by the platform, then the video formats of the first media resource and the second media resource are determined to have passed the verification. If the video format of either the first media resource or the second media resource is not a video format supported by the platform, then the video formats of the first media resource and the second media resource are determined to have failed the verification.

[0255] Optionally, the platform supports video formats including Moving Picture Experts Group Vudio Layer 4 (MP4), WebM, and Audio Video Interleaved (AVI). It is understood that if the second media resource is an image resource, there is no need to check the video format of the second media resource.

[0256] Verification item (2), encoding format check.

[0257] In some embodiments, the electronic device acquires the encoding formats of the first media resource and the second media resource, and determines whether the encoding formats of the first media resource and the second media resource are encoding formats supported by the platform. If the encoding formats of both the first media resource and the second media resource are encoding formats supported by the platform, then the encoding formats of the first media resource and the second media resource are determined to have passed the verification. If the encoding format of either the first media resource or the second media resource is not an encoding format supported by the platform, then the encoding formats of the first media resource and the second media resource are determined to have failed the verification.

[0258] The encoding format can include both video and audio encoding formats. For example, video encoding formats include Advanced Video Coding (AVC, also known as H.264), High Efficiency Video Coding (HEVC, also known as H.265), and Video Processor 9 (VP9). Audio encoding formats include Advanced Audio Coding (AAC) and Moving Picture Experts Group Audio Layer 3 (MP3). It is understood that when the second media resource is an image resource, there is no need to check the encoding format of the second media resource.

[0259] Verification item (3), file size check.

[0260] In some embodiments, the electronic device obtains the file sizes of the first media resource and the second media resource, and determines whether the file sizes of the first media resource and the second media resource exceed the file size required by the platform. If the file sizes of both the first media resource and the second media resource do not exceed the file size required by the platform, then the file sizes of the first media resource and the second media resource are determined to have passed the verification. If the file size of either the first media resource or the second media resource exceeds the file size required by the platform, then the file sizes of the first media resource and the second media resource are determined to have failed the verification.

[0261] This avoids uploading excessively large media resource files, thus saving storage space and transmission bandwidth. It's understood that the first and second media resources can be in file format, such as video files.

[0262] Verification item (4), security check.

[0263] In some embodiments, the electronic device performs a security check on the first media resource and the second media resource to determine whether malicious code or viruses exist in them. If no malicious code or viruses are found in the first media resource or the second media resource, the first media resource and the second media resource are deemed to have passed the verification. If malicious code or viruses are found in either the first media resource or the second media resource, the first media resource and the second media resource are deemed to have failed the verification.

[0264] Optionally, the electronic device may employ antivirus software to scan the first and second media resources to complete the aforementioned security checks. This prevents malicious code or viruses from spreading through the media resources, thereby ensuring the security of data transmission.

[0265] It is worth noting that during the verification process of the above-mentioned verification items (1) to (4), the electronic device can use an open-source audio and video codec tool (Fast Forward Moving Picture Experts Group, ffmpeg) to parse the first media resource and the second media resource, obtain the corresponding verification item information, and compare the verification item information of the first media resource and the second media resource with the legality threshold required by the platform to complete the verification. Of course, in some other embodiments, the electronic device may also include other verification items, which are not limited in this application embodiment.

[0266] S804. Electronic devices acquire the content semantics of the first media resource and the content semantics of the second media resource.

[0267] In this embodiment, the first media resource is a video resource, and the content semantics of the first media resource includes the content semantics of audio and video. The second media resource is either a video resource or an image resource. When the second media resource is a video resource, its content semantics includes the content semantics of audio and video; when the second media resource is an image resource, its content semantics includes the content semantics of the image.

[0268] In some embodiments, the process by which an electronic device acquires the semantic content of audio may include any one of the following four implementation methods.

[0269] Method 1: Electronic devices perform speaker recognition and speech recognition on audio, and obtain the recognition result of the audio as the semantic content of the audio.

[0270] Speaker recognition is a technology that identifies a speaker by recognizing personal characteristics in audio, that is, identifying "who is speaking" through sound signals. Speech recognition is a technology that converts audio into text to obtain the audio's content text, that is, converting human speech into a sequence of characters, such as text. The audio recognition result can include speaker recognition results and speech recognition results, that is, obtaining human voice dialogue, including the speaker's identity information (such as gender) and content text (or plot text).

[0271] In some possible implementations, the process of speaker recognition in audio by an electronic device can be as follows: based on speaker diarization technology, different voices included in the audio are distinguished to obtain speaker recognition results. Optionally, before implementing this scheme, the electronic device iteratively trains an initial deep learning model based on speaker diarization technology to obtain a speaker recognition model (such as a voiceprint recognition model). Then, the speaker recognition model is used to distinguish different voices included in the audio to obtain speaker recognition results.

[0272] In some possible implementations, the process of electronic devices performing speech recognition on audio can be as follows: based on automatic speech recognition (ASR) technology, the audio is converted into text to obtain the speech recognition result. Optionally, before implementing this solution, the electronic device iteratively trains an initial deep learning model based on ASR technology, or fine-tunes a pre-trained speech recognition model to obtain a speech recognition model. Then, using the speech recognition model, the audio is converted into text to obtain the speech recognition result. Optionally, the electronic device completes the above speech recognition process by calling a speech recognition tool from a third-party library, such as calling the application programming interface (API) or software development kit (SDK) service of the speech recognition tool.

[0273] Method 2: Electronic devices perform speaker recognition and speech recognition on audio, obtain the audio recognition results, and extract keywords from the audio recognition results as the semantic content of the audio.

[0274] The process of extracting keywords from the audio recognition results by the electronic device is as follows: based on the natural language semantic understanding model, keywords are extracted from the audio recognition results to obtain multiple keywords. Optionally, the natural language semantic understanding model is a bidirectional encoder representations from transformers (BERT) model or a generative pre-training transformer (GPT) model.

[0275] Method 3: The electronic device performs speaker recognition on the audio and subtitle recognition on the corresponding video frames to obtain the audio recognition result, which serves as the semantic content of the audio.

[0276] The audio recognition results can include speaker recognition results and subtitle recognition results, which means that the human voice dialogue is obtained, including the speaker's identity information (such as gender) and the content text.

[0277] In some possible implementations, the process of subtitle recognition for the video frame corresponding to the audio by the electronic device can be as follows: based on a character recognition model, subtitle detection and recognition are performed on the video frame to obtain the subtitle recognition result. Optionally, before implementing this scheme, the electronic device iteratively trains an initial deep learning model to obtain a subtitle recognition model. Then, the subtitle recognition model is used to perform subtitle detection and recognition to obtain the subtitle recognition result.

[0278] Implementation Method 4: The electronic device performs speaker recognition on the audio and subtitle recognition on the corresponding video frames to obtain the audio recognition result. Keywords are extracted from the audio recognition result as the semantic content of the audio.

[0279] The above embodiments provide four methods for obtaining the content semantics of audio. Among them, speaker recognition combined with speech recognition or speaker recognition combined with subtitle recognition can quickly determine the audio recognition result. Furthermore, the audio recognition result can be directly determined as the audio content semantics, or keywords can be extracted from the audio recognition result to obtain the audio content semantics. By extracting keywords, not only can effective content semantics be determined, but the processing content of electronic devices can also be reduced, which helps to improve the processing efficiency of media resources.

[0280] In some embodiments, the process by which an electronic device acquires the semantic content of a video may include either of the following two implementation methods.

[0281] Implementation Method 1: The electronic device performs at least one of entity recognition, scene recognition, face detection, and behavior detection on the video frames included in the video to obtain the video frame recognition result, which serves as the content semantics of the video.

[0282] The video frame recognition results include at least one of the following: entity recognition results, scene recognition results, face detection results, and behavior detection results.

[0283] Entity recognition refers to the process of identifying and classifying entities with specific meanings, such as people, vehicles, and roads, within video frames. The entity recognition results indicate the type of entity in the video frame. In some possible implementations, the entity recognition results can be in the form of entity labels.

[0284] In some possible implementations, the electronic device performs entity recognition on the video frame based on an entity recognition model to obtain the entity recognition result. For example, the electronic device inputs the video frame into an entity recognition model, and the model performs entity recognition on the video frame to obtain the entity recognition result. Alternatively, if the electronic device already has the semantic content of the audio, it can input the semantic content of the video frame and its corresponding audio into an entity recognition model, and the model performs entity recognition on the semantic content of the video frame and its corresponding audio to obtain the entity recognition result.

[0285] Optionally, before implementing this solution, the electronic device iteratively trains an initial deep learning model based on an image dataset labeled with entities, or fine-tunes a pre-trained multi-label classification model to obtain an entity recognition model. For example, the entity recognition model may include a Swin Transformer sub-model and a bidirectional encoder representations from transformers (BERT) sub-model, wherein the Swin Transformer sub-model is used to extract entity features from video frames, and the BERT sub-model is used to output entity labels based on the entity features in the video frames, the entity labels indicating the entity type.

[0286] Scene recognition refers to the analysis of video frames to intelligently understand information such as objects, positions, and appearances in different scenes, thereby achieving scene identification. The scene recognition result indicates the scene type of the video frame. In some possible implementations, the scene recognition result can be in the form of scene labels.

[0287] In some possible implementations, the electronic device performs scene recognition on the video frame based on a scene recognition model to obtain the scene recognition result. For example, the electronic device inputs the video frame into a scene recognition model, and the scene recognition model performs scene recognition on the video frame to obtain the scene recognition result. Alternatively, if the electronic device already has the semantic content of the audio, it can input the semantic content of the video frame and its corresponding audio into a scene recognition model, and the scene recognition model performs scene recognition on the semantic content of the video frame and its corresponding audio to obtain the scene recognition result.

[0288] Optionally, before implementing this solution, the electronic device iteratively trains an initial deep learning model based on a scene-annotated image dataset, or fine-tunes a pre-trained multi-label classification model to obtain a scene recognition model. For example, the scene recognition model may include a Swin Transformer sub-model and a BERT sub-model, wherein the Swin Transformer sub-model is used to extract scene features from video frames, and the BERT sub-model is used to output scene labels based on the scene features in the video frames, the scene labels indicating the scene type.

[0289] Face detection refers to the detection and localization of the position, size, and expression of faces in a video frame. The face detection results are used to indicate the type of expression of a person in the video frame. In some possible implementations, the face detection results can be in the form of expression labels.

[0290] In some possible implementations, the electronic device performs face detection on the video frame based on a face detection algorithm to obtain the face detection result. For example, the electronic device inputs the video frame into a face detection algorithm, which then performs face detection on the video frame to obtain the face detection result. Alternatively, if the electronic device already possesses the semantic content of the audio, it can input the semantic content of the video frame and its corresponding audio into a face detection algorithm, which then performs face detection on the semantic content of the video frame and its corresponding audio to obtain the face detection result.

[0291] For example, the face detection algorithm can be a target detection algorithm (YOLO, you only look once), a face detection algorithm (retina face), or other types of detection algorithms. Optionally, before implementing this solution, the electronic device iteratively trains an initial deep learning model (such as a closed-set deep learning expression recognition model) based on the face detection algorithm to obtain a face detection model. Then, the face detection model is used to perform face detection on video frames to obtain the face detection result.

[0292] After identifying face regions based on face detection algorithms, face alignment and normalization are performed on these regions. Then, expression features are extracted from the processed face regions to identify face detection results that indicate expression types. Face alignment involves detecting key points in the face image, automatically locating key facial feature points, and then correcting and standardizing the image based on these feature points, such as through affine transformations, to facilitate subsequent expression analysis. Affine transformations convert face images of different poses and sizes into a unified coordinate system, reducing variations caused by these factors and eliminating the impact of pose and size changes.

[0293] Action detection, also known as behavior detection, refers to the process of extracting spatial, temporal, and fused features of entities in a video frame to capture dynamic patterns and temporal changes, thereby identifying the actions and behaviors of entities within the video frame, such as human posture, behavior, and activities. The action detection results indicate the type of behavior of entities in the video frame. In some possible implementations, the action detection results can be in the form of action labels.

[0294] In some possible implementations, the electronic device performs behavior detection on the video frame based on an action recognition model, obtaining the behavior detection result. For example, the electronic device inputs the video frame into the action recognition model, which then performs behavior detection on the video frame to obtain the behavior detection result. Alternatively, if the electronic device already possesses the semantic content of the audio, it can input the semantic content of the video frame and its corresponding audio into the action recognition model, which then performs behavior detection on the semantic content of the video frame and its corresponding audio to obtain the behavior detection result.

[0295] Optionally, before implementing this scheme, the electronic device iteratively trains an initial deep learning model based on an image dataset labeled with actions of specific entities to obtain an action detection model. Then, the action detection model is used to detect actions in video frames to obtain the action detection result. The deep learning model can be a two-stream network or a 3D convolutional neural network (CNN).

[0296] In the above embodiments, the process of determining the semantic content of a video is illustrated by taking entity recognition, scene recognition, face detection, and behavior detection of video frames as examples. It can be seen that the above embodiments illustrate the process of determining the semantic content of a video by taking the separate execution of entity recognition, scene recognition, face detection, and behavior detection as examples. In other embodiments, the processing of video frames (i.e., entity recognition, scene recognition, face detection, and behavior detection) can also directly call a multimodal large model to process the video frames and obtain a content description of the video frames. This content description includes at least one of the following: entity recognition result, scene recognition result, face detection result, and behavior detection result.

[0297] For example, when a video frame is input into a multimodal large model, the built-in prompt is "Please try to understand this image and summarize the content description of the image from the following aspects: (1) the entities contained in the image (2) the scene in the image (3) the people in the image, please ignore if they do not exist (4) the expressions of the people, please ignore if they do not exist (5) the interactive behavior, please ignore if they do not exist".

[0298] Method 2: The electronic device extracts keyframes from the video frames included in the video, and performs at least one of entity recognition, scene recognition, face detection and behavior detection on the keyframes to obtain the keyframe recognition result, which serves as the content semantics of the video.

[0299] The keyframe recognition results include at least one of entity recognition results, scene recognition results, face detection results, and behavior detection results.

[0300] In some possible implementations, the process of keyframe extraction by an electronic device is as follows: From the intermediate frames of the video (excluding the first and last frames), based on an inter-frame difference algorithm, video frames with an inter-frame difference intensity greater than a threshold are selected as candidate keyframes. Based on a sharpness detection algorithm, motion-blurred video frames are filtered from the candidate keyframes, and the filtered video frames are deduplicated to obtain the keyframes.

[0301] Among them, the inter-frame difference method is a method to obtain the contour of a moving target by performing a difference operation on two consecutive frames of images in a video. In this way, by extracting keyframes, not only can representative keyframes in the video be obtained, but it is also unnecessary to perform entity recognition, scene recognition, face detection, and behavior detection on all video frames, thereby reducing the computational load of electronic devices and improving their computational efficiency.

[0302] The above embodiments provide two methods for acquiring the semantic content of videos. Specifically, by employing technologies such as entity recognition, scene recognition, face detection, and behavior detection, the video frame recognition results or keyframe recognition results can be quickly determined. Furthermore, extracting keyframes not only identifies valid semantic content but also reduces the processing load on electronic devices, thus improving the processing efficiency of media resources.

[0303] In some embodiments, the process by which an electronic device acquires the content semantics of an image may include: the electronic device performing at least one of entity recognition, scene recognition, face detection, and behavior detection on the image to obtain an image recognition result, which serves as the content semantics of the image. The image recognition result includes at least one of entity recognition result, scene recognition result, face detection result, and behavior detection result.

[0304] In the above embodiments, the process of acquiring the content semantics of audio, video, and image is described respectively.

[0305] The following section introduces the process of acquiring the semantic meaning of content from primary media resources.

[0306] In some embodiments, the content semantics of the first media resource includes the content semantics of multiple shot segments. The process of obtaining the content semantics of the multiple shot segments may include the following steps (1-1) to (1-2).

[0307] Step (1-1): The electronic device performs shot analysis on the first media resource to obtain multiple shot fragments of the first media resource.

[0308] Shot analysis can include shot detection and shot segmentation. Specifically, shot boundaries are detected by detecting feature changes between video frames in the first media resource. These feature changes can include variations in color, texture, shape, and motion. The detected shot boundaries are then used to segment the first media resource into multiple shot segments. Each shot segment consists of multiple consecutive video frames.

[0309] In some embodiments, the electronic device performs lens detection and lens segmentation on the first media resource to obtain multiple lens segments that have undergone lens detection and lens segmentation.

[0310] In some possible implementations, the electronic device can calculate the pixel histogram difference index between two adjacent video frames in the first media resource based on a histogram analysis algorithm. By determining whether the pixel histogram difference index between two adjacent video frames exceeds a difference threshold, shot boundaries in the first media resource can be detected. It is understood that if the pixel histogram difference index between two adjacent video frames exceeds the difference threshold, it indicates a shot change, and these two adjacent video frames are determined to be shot boundaries. Furthermore, segmenting these two adjacent video frames can divide the first media resource into multiple shot segments.

[0311] For example, electronic devices can store histogram analysis algorithms in open-source third-party video detection and segmentation tools (such as PySceneDetect) or libraries optimized based on PySceneDetect. PySceneDetect helps users automatically identify scene transition points in videos, thereby enabling video segmentation, editing, or analysis. Optionally, the electronic device can perform the aforementioned scene detection and segmentation process by calling PySceneDetect or libraries optimized based on PySceneDetect.

[0312] In other possible implementations, the electronic device can detect shot boundaries in the first media resource based on a shot boundary detection (SBD) model. Then, based on the detected shot boundaries, segmentation can be performed to divide the first media resource into multiple shot segments.

[0313] Optionally, before implementing this scheme, the electronic device iteratively trains an initial deep learning model based on a video dataset labeled with lens boundaries to obtain the aforementioned lens boundary detection model.

[0314] It is worth noting that in other possible implementations, the electronic device may also use other types of methods to complete the lens detection and lens segmentation process, such as pixel comparison method, edge comparison method, etc., which are not limited in this application embodiment.

[0315] Steps (1-2): The electronic device performs content understanding on multiple shot segments to obtain the semantic content of the multiple shot segments.

[0316] In some embodiments, the content semantics of multiple shot segments may include the content semantics of the audio and the content semantics of the video in each shot segment.

[0317] In some possible implementations, for each of the multiple shot segments, the audio track of that shot segment is separated to obtain its audio. Then, the semantic content of the audio of that shot segment is obtained. Optionally, the electronic device can perform the above audio track separation process by calling a video editing tool (moviepy) from a third-party library.

[0318] Optionally, the process of obtaining the content semantics of audio in each shot segment may include at least one of the following: for each shot segment among multiple shot segments, perform speaker recognition and speech recognition on the audio of that shot segment to obtain the recognition result of the audio of that shot segment, which is used as the content semantics of the audio of that shot segment; or, for each shot segment among multiple shot segments, perform speaker recognition and speech recognition on the audio of that shot segment to obtain the recognition result of the audio of that shot segment, and extract keywords from the audio recognition result as the content semantics of the audio of that shot segment; or, for each shot segment among multiple shot segments, perform speaker recognition on the audio of that shot segment and perform subtitle recognition on the video frame corresponding to the audio to obtain the recognition result of the audio of that shot segment, which is used as the content semantics of the audio of that shot segment; or, for each shot segment among multiple shot segments, perform speaker recognition on the audio of that shot segment and perform subtitle recognition on the video frame corresponding to the audio to obtain the recognition result of the audio of that shot segment, and extract keywords from the audio recognition result as the content semantics of the audio of that shot segment. For specific implementation details, please refer to the above process for obtaining the content semantics of audio, which will not be repeated here. In this embodiment of the application, keywords extracted from the semantic content of the audio in the shot clip can be used as shot tags.

[0319] Optionally, the process of obtaining the content semantics of the video in each shot segment may include at least one of the following: for each shot segment among multiple shot segments, performing at least one of entity recognition, scene recognition, face detection, and behavior detection on the video frames included in the shot segment to obtain the video frame recognition result of the shot segment, which serves as the content semantics of the video in that shot segment; or, for each shot segment among multiple shot segments, extracting keyframes from the video frames included in the shot segment, and performing at least one of entity recognition, scene recognition, face detection, and behavior detection on the keyframes of the shot segment to obtain the keyframe recognition result of the shot segment, which serves as the content semantics of the video in that shot segment. For specific implementation details, please refer to the above process for obtaining the content semantics of the video, which will not be repeated here.

[0320] Regarding the aforementioned behavior detection process for video frames of the first media resource, some possible implementations include performing behavior detection on entities in the video frame related to the second media resource, or entities in the video frame with a preset content type, to obtain behavior detection results for entities related to the second media resource, or entities in the video frame with a preset content type. Entities related to the second media resource can be those that are the same as or similar to keywords (i.e., advertising themes) extracted from the semantic content of the audio of the second media resource.

[0321] Thus, based on the above steps (1-1) and (1-2), multiple shot clips of the first media resource and the semantic text of each shot clip can be generated, such as (shot clip, semantic text) resource pairs. For example, Figure 9 This is a schematic diagram illustrating an audio / video parsing process provided in an embodiment of this application. See also... Figure 9 The process of obtaining resource pairs (shot clips, semantic text content) from a first media resource can include the following three steps: ① Video segmentation: Perform shot detection and segmentation on the first media resource to obtain multiple shot clips. ② Audio separation: Perform audio separation on the multiple shot clips of the first media resource to obtain the audio of each shot clip. ③ Speech recognition: Perform speaker recognition and speech recognition on the audio of the multiple shot clips to obtain the audio recognition results of each shot clip. In this way, resource pairs (shot clips, semantic text content) from the first media resource can be obtained.

[0322] For example, taking a TV drama video involving a vehicle collision as the first media resource, the metadata of the resource pair (shot fragment, content semantic text) is shown below using consecutive shots (shot 10, shot 11 and shot 12) involving a vehicle collision as an example.

[0323] Scene 10:

[0324] {

[0325] "start_time": "00:10:58"

[0326] "end_time": "00:11:19",

[0327] "text": "Female voice 1: Oh my god, it's still rising.\nMale voice 1: Dad's driving, you should hang up now.\nFemale voice 1: Wait, I'm still happy.\nMale voice 1: Dad's driving, this car is borrowed from the repair shop, it doesn't have a spare, it's dangerous, you should hang up now.\nFemale voice 1: No, no.\nMale voice 1: This child."

[0328] }

[0329] Scene 11:

[0330] {

[0331] "start_time": "00:11:20",

[0332] "end_time": "00:11:25",

[0333] "text": "..."

[0334] }

[0335] Scene 12:

[0336] {

[0337] "start_time": "00:11:25"

[0338] "end_time": "00:11:42",

[0339] "text": "Male voice 1: I was wondering if there's something wrong with the passenger airbag. I'm fine, but how did you get such a bad neck injury?\nMale voice 2: It shouldn't be, it's just that my seatbelt strained a bit."

[0340] }

[0341] For example, Figure 10 This is a schematic diagram of a keyframe in a shot segment provided in an embodiment of this application. See also... Figure 10 Taking a TV series video involving vehicle collisions as an example of the first media resource, and taking the key frame in the first media resource as the key frame of shot 11 as an example, after performing the above entity recognition and behavior detection on the key frame, the entity recognition result can be "car" or "road", and the behavior detection result can be "collision".

[0342] In the embodiments shown in steps (1-1) to (1-2) above, a method for determining the content semantics of a first media resource is provided. This method involves segmenting the first media resource into multiple shot fragments through shot parsing, and then obtaining the content semantics of these multiple shot fragments through content understanding, thus obtaining the content semantics of the first media resource. In some possible implementations, the content semantics of the first media resource can be in the form of tag phrases, such as phrases including entity tags, scene tags, expression tags, behavior tags, and shot tags.

[0343] In other embodiments, the content semantics of the first media resource includes the content semantics of the target shot segment. The process of acquiring the content semantics of the target shot segment includes the following steps (2-1) to (2-3).

[0344] Step (2-1): The electronic device extracts keywords from the text information of the second media resource.

[0345] The text information of the second media resource can be the advertising copy of the second media resource. For example, if the second media resource is a video resource, the text information can be the semantic content of the audio of the second media resource, that is, the advertising copy obtained from audio recognition. As another example, if the second media resource is an image resource, the text information can be a preset advertising copy, such as the preset advertising copy of an advertising resource.

[0346] In some possible implementations, the electronic device extracts keywords from the text information of the second media resource based on a topic extraction model. Based on these keywords, it performs preliminary localization of the locations in the target video text containing the keywords or synonyms, for use in subsequent shot selection.

[0347] Step (2-2): The electronic device selects the target shot segment from multiple shot segments, including the theme word or synonym.

[0348] The similarity between synonyms and keywords reaches the similarity threshold.

[0349] Step (2-3): The electronic device performs content understanding on the target shot fragment to obtain the content semantics of the target shot fragment.

[0350] In the embodiments shown in steps (2-1) to (2-3) above, a method for determining the content semantics of a first media resource is provided. This method involves identifying target shot segments, including those containing the keywords or synonyms, based on the keywords of the second media resource. This allows for the identification of shot segments in the first media resource that are related to the second media resource. This not only identifies valid shot segments but also reduces the processing load on electronic devices, thus improving the processing efficiency of media resources. It is understood that since the first media resource may contain numerous shots with extremely short durations, the aforementioned keyword-based range positioning can reduce the output of invalid (shot segments, content semantic text) resource pairs, focusing more on relatively valid segments.

[0351] Optionally, the electronic device first performs shot analysis on the first media resource to obtain multiple shot fragments of the first media resource, and then obtains the semantic content of the audio of the multiple shot fragments. Next, it extracts keywords from the text information of the second media resource to obtain the keywords of the second media resource. From the multiple shot fragments, a target shot fragment containing the keywords or synonyms is selected.

[0352] Alternatively, the electronic device first acquires the complete audio recognition result of the first media resource. It then extracts keywords from the text information of the second media resource to obtain the keywords of the second media resource. From the complete audio recognition result of the first media resource, it determines the video frame corresponding to the audio recognition result including the keywords or synonyms. Next, it performs shot analysis on the first media resource to obtain multiple shot segments of the first media resource. Finally, from the multiple shot segments, it selects the target shot segment containing the determined video frame.

[0353] This application does not limit the execution order of shot analysis followed by keyword extraction, or keyword extraction followed by shot analysis.

[0354] The following section introduces the process of acquiring the content semantics of second media resources.

[0355] In some embodiments, when the second media resource is a video resource, the process by which the electronic device acquires the content semantics of the second media resource may include the following steps (3-1) to (3-2).

[0356] Step (3-1): The electronic device acquires the semantic content of the audio of the second media resource.

[0357] The process by which the electronic device acquires the semantic content of the audio from the second media resource is the same as the process by which it acquires the semantic content of the audio from each of the multiple shot segments. This can be referred to in steps (1-2) above, where the process of acquiring the semantic content of the audio from the shot segments is described, and will not be repeated here. In this embodiment, the keywords extracted from the semantic content of the audio from the second media resource can be used as advertising tags.

[0358] Step (3-2): The electronic device acquires the semantic content of the video from the second media resource.

[0359] The process by which the electronic device acquires the content semantics of the video of the second media resource is the same as the process by which it acquires the content semantics of the video of each of the multiple shot segments. Please refer to the process of acquiring the content semantics of the video of the shot segments in steps (1-3) above, which will not be repeated here.

[0360] Based on the above steps (3-1) to (3-2), the content semantics of the second media resource can be obtained. The content semantics of the second media resource can be in the form of tag phrases, such as phrases including entity tags, scene tags, emoticon tags, behavior tags, and advertising tags.

[0361] Understandably, for secondary media resources such as advertising resources to be embedded, since the length of secondary media resources is usually shorter than that of primary media resources such as long videos or derivative videos, in order to improve the processing efficiency of media resources, and considering that the processing details of secondary media resources focus more on the naturalness of the connection when the secondary media resources are embedded, there is no need to perform shot analysis processing on the secondary media resources. It is only necessary to obtain the content text of the secondary media resources through audio track separation.

[0362] Based on the above S804, by using content understanding technology such as video content understanding technology, it is possible to obtain tag phrases for multiple shot segments in the first media resource based on multimodal information, as well as tag phrases for the second media resource (such as advertising resources) based on multimodal information, so as to complete the fusion of media resources based on content semantics in the future.

[0363] S805, the electronic device determines candidate frames from the first media resource based on the content semantics of the first media resource and the content semantics of the second media resource.

[0364] Candidate frames are used to indicate video frames for which a second media resource is to be implanted.

[0365] In some embodiments, the process by which an electronic device determines candidate frames from a first media resource may include either of the following two implementations.

[0366] Implementation Method 1: The candidate frames and the second media resource satisfy the content semantic similarity condition. Accordingly, the process of determining candidate frames is as follows: Based on the content semantics of the first media resource and the second media resource, the electronic device determines video frames from the first media resource that satisfy the content semantic similarity condition with the second media resource, and uses these as candidate frames.

[0367] In one implementation, the electronic device determines candidate frames from the first media resource that meet the content semantic similarity condition with the second media resource, including the following steps (4-1) to (4-3).

[0368] Step (4-1): The electronic device determines the content similarity between the shot clip and the second media resource based on the content semantics of the shot clip in the first media resource and the content semantics of the second media resource.

[0369] Content similarity is used to measure the degree of similarity between feature vectors of two content semantics. For example, content similarity can be cosine similarity, Jaccard similarity, Manhattan distance, Hamming distance, Mahalanobis distance or other similarity metrics.

[0370] In some possible implementations, the electronic device acquires a first feature representation of the content semantics of each shot segment in the first media resource, and acquires a second feature representation of the content semantics of the second media resource, performs a similarity calculation between the first feature representation and the second feature representation, and obtains the content similarity between the first feature representation and the second feature representation.

[0371] Step (4-2): The electronic device identifies candidate shot segments from the first media resource.

[0372] Specifically, the content similarity between candidate shot clips and secondary media resources meets the similarity criteria. These criteria can be that the content similarity reaches a similarity threshold, or that the content similarity is ranked higher in a descending order of numerical values.

[0373] In some possible implementations, taking a similarity threshold as an example, the electronic device determines candidate shot segments from multiple shot segments whose content similarity reaches the threshold. Alternatively, in other possible implementations, taking a similarity criterion as ranking the content similarity higher in descending order of numerical value, the electronic device sorts the content similarity of multiple shot segments in descending order of numerical value and determines the candidate shot segments with the highest content similarity ranking.

[0374] Step (4-3): The electronic device determines candidate frames from the candidate shot segments.

[0375] Among them, the candidate frame is the last frame of the candidate shot segment or the first frame of the next shot segment.

[0376] In steps (4-1) to (4-3) above, a method for determining candidate frames based on the content semantics of shot fragments is provided. This method first determines candidate shot fragments that meet the similarity conditions, and then determines candidate frames from these candidate shot fragments, so that the determined candidate frames are candidate frames that match the plot content of the second media resource.

[0377] In another implementation, the electronic device determines candidate frames from the first media resource that meet the content semantic similarity condition with the second media resource, including the following steps (5-1) to (5-6).

[0378] Step (5-1): The electronic device performs clustering processing on the continuous shots in the first media resource to obtain at least one scene segment.

[0379] In some possible implementations, the electronic device clusters consecutive shots from multiple shot segments based on a scene segmentation model to obtain at least one scene segment. Optionally, the scene segmentation model is trained using a deep learning model based on a Transformer architecture.

[0380] The semantic content of each scene segment is composed of the semantic content of each shot segment in the same cluster.

[0381] Step (5-2): The electronic device determines the content similarity between at least one scene segment and the second media resource based on the content semantics of at least one scene segment in the first media resource and the content semantics of the second media resource.

[0382] In some embodiments, the process of determining the content similarity between at least one scene segment and a second media resource may include: determining the content semantics of at least one scene segment based on the content semantics of the shot segments included in the at least one scene segment; determining a first content similarity between at least one scene segment and the second media resource based on the content semantics of the at least one scene segment and the content semantics of the second media resource; generating a text summary of at least one scene segment based on the content semantics of the shot segments included in the at least one scene segment; determining a second content similarity between at least one scene segment and the second media resource based on the text summary of at least one scene segment and the text information of the second media resource; and determining the content similarity between at least one scene segment and the second media resource based on the first content similarity and the second content similarity.

[0383] In some possible implementations, the process of generating a text summary of at least one scene segment based on the content semantics of the shot segments included in at least one scene segment can be as follows: A text summary of at least one scene segment can be generated based on the audio content semantics of the shot segments included in at least one scene segment. Specifically, this process can include: for each scene segment in the at least one scene segment, concatenating the audio content semantics of multiple shot segments included in the scene segment to obtain the audio content semantics of the scene segment. Then, based on a large natural language model, the audio content semantics of the scene segment are processed to obtain a text summary of the scene segment.

[0384] In some possible implementations, determining the content similarity between at least one scene segment and the second media resource based on the first content similarity and the second content similarity includes: performing a weighted summation of the first content similarity and the second content similarity to obtain the content similarity between at least one scene segment and the second media resource.

[0385] The above implementation provides a way to determine the content similarity between a scene segment and a second media resource. This method not only considers the content semantics of the scene segment and the content semantics of the second media resource, but also the text summary of the scene segment and the text information of the second media resource. This increases the amount of information considered in determining the content similarity and improves the accuracy of determining the content similarity.

[0386] Step (5-3): The electronic device determines candidate scene segments from at least one scene segment.

[0387] Among them, the content similarity between candidate scene segments and second media resources meets the similarity condition.

[0388] For example, the semantic content of each scene segment can be represented by scene content tags, the semantic content of the second media resource can be represented by advertising content tags, and the audio recognition result of the second media resource is the advertising text. By calculating the content similarity between the scene content tags, text summaries and advertising content tags and advertising text of each scene segment, scene segments that meet the similarity conditions, such as the similarity reaching the similarity threshold, are selected as candidate scene segments.

[0389] Step (5-4): The electronic device determines the content similarity between the shot segments included in the candidate scene segment and the second media resource based on the content semantics of the shot segments included in the candidate scene segment and the content semantics of the second media resource.

[0390] Thus, after identifying candidate scene segments whose content similarity meets the similarity criteria, the similarity between the content semantics of each shot segment within the candidate scene segment and the content semantics of the second media resource can be further calculated in order to subsequently determine candidate shot segments.

[0391] Step (5-5): The electronic device determines the candidate shot segments from the shot segments included in the candidate scene segments.

[0392] Among them, the content similarity between candidate shot clips and second media resources meets the similarity condition.

[0393] Steps (5-6): The electronic device determines candidate frames from the candidate shot segments.

[0394] Among them, the candidate frame is the last frame of the candidate shot segment or the first frame of the next shot segment.

[0395] In steps (5-1) to (5-6) above, a method is provided to determine candidate frames based on the content semantics of scene segments and shot segments. This method first determines candidate scene segments that meet the similarity conditions, then determines candidate shot segments from the candidate scene segments, and then determines candidate frames from the candidate shot segments, so that the determined candidate frames are candidate frames that match the plot content of the second media resource.

[0396] The above embodiment illustrates the process of determining candidate frames based on video plot content by calculating content similarity with a second media resource. The process of determining candidate frames based on a preset template will be described below.

[0397] Implementation Method Two: The content semantics of candidate frames include one or more of entity type, scene type, expression type, or behavior type. Accordingly, the process of determining candidate frames is as follows: Based on the content semantics of the first media resource and the content semantics of the second media resource, the electronic device determines video frames from the first media resource whose content semantics include one or more of entity type, scene type, expression type, or behavior type, as candidate frames.

[0398] The entity type can be a preset entity type, the scene type can be a preset scene type, the expression type can be a preset expression type, and the behavior type can be a preset behavior type. For example, the electronic device can be configured with templates based on entities in video frames (used to determine if a preset entity type exists in the first media resource), templates based on scene transitions (used to determine if a preset scene type exists in the first media resource), templates based on facial expressions in video frames (used to determine if a preset expression type exists in the first media resource), and templates based on actions in video frames (used to determine if a preset behavior type exists in the first media resource). In this way, candidate frames based on preset templates can be determined.

[0399] In some possible implementations, the electronic device determines the content similarity between the semantic content of the first media resource and the content types (i.e., the aforementioned entity types, scene types, expression types, or behavior types) in each preset template, and acquires video frames whose content similarity reaches a similarity threshold as candidate frames. The content type can be in the form of a label, such as a content tag.

[0400] For example, taking a scene-transition-based template, the video frame corresponding to the end of a song in a singing scene can be used as a candidate frame. Similarly, taking a template based on facial expressions in video frames, the video frame corresponding to a character's blank expression can be used as a candidate frame. This allows for adaptation to a wider range of advertising placement scenarios, enriching the selection of advertising placement locations.

[0401] In the above embodiments, two implementation methods for determining candidate frames are provided. According to implementation method one, candidate frames based on video plot content can be determined, such as points determined based on the similarity between the plot content of a first media resource and the advertising content of a second media resource. According to implementation method two, candidate frames based on preset templates can be determined, such as points determined based on the plot content of the first media resource, including one or more content types such as entity type, scene type, expression type, or behavior type.

[0402] Furthermore, after identifying candidate frames for the first media resource, the electronic device displays these candidate frames. At this point, the user can confirm whether the candidate frames are suitable and perform a feedback operation. Correspondingly, the electronic device receives the user's feedback on the candidate frames and, in response, obtains first feedback information. This first feedback information instructs the user to confirm the candidate frames or the user-corrected video frames. Subsequently, the electronic device can perform subsequent media resource processing based on the first feedback information. Thus, by displaying candidate frames for the first media resource to the user, allowing them to choose whether to optimize or correct the candidate frames, the flexibility of media resource processing is improved, enabling the generation of more personalized target media resources based on user feedback.

[0403] S806. The electronic device obtains the linking text of the second media resource based on the determined candidate frames.

[0404] Among them, the text of the connecting words matches the semantic content of the first media resource and the semantic content of the second media resource.

[0405] In some embodiments, the connecting word text of the second media resource is generated based on the content semantics of the candidate frame, the shot clips before the candidate frame, the shot clips after the candidate frame, and the content semantics of the second media resource.

[0406] In some embodiments, the process of an electronic device generating the linking text of a second media resource may include the following steps (6-1) to (6-2).

[0407] Step (6-1): The electronic device generates plot summary text before the candidate frame and plot summary text after the candidate frame based on the content semantics of the shot segments before and after the candidate frame.

[0408] In some possible implementations, the electronic device generates plot summary text before the candidate frame and plot summary text after the candidate frame based on the content semantics of the audio of the shot segments before and after the candidate frame.

[0409] Optionally, the electronic device inputs the semantic content of the audio of the shot segments before and after the candidate frame into a natural language processing model. The natural language processing model processes the semantic content of the audio of the shot segments before and after the candidate frame to obtain the plot summary text before and after the candidate frame.

[0410] For example, when inputting the audio content semantics of the shot segments before and after the candidate frame into a natural language big data model, a built-in prompt can be: "[Please summarize the plot summary text based on ${the audio content semantics of the shot segments before and after the candidate frame}]".

[0411] Taking a TV series video involving a vehicle collision as an example, assuming the candidate frame is the video frame in shot 11, then the shot segment before the candidate frame is shot 10, and correspondingly, the plot summary text before the candidate frame is the plot text of shot 10. The shot segment after the candidate frame is shot 12, and correspondingly, the plot summary text after the candidate frame is the plot text of shot 12.

[0412] For example, the built-in prompt for the plot summary text used before generating candidate frames is: [Please summarize the plot summary text based on "Female voice 1: Oh my god, it's still rising. Male voice 1: Dad's driving, you should hang up now. Female voice 1: Wait, I'm still happy. Male voice 1: Dad's driving, this car is borrowed from the repair shop, it doesn't have a secure in-vehicle communication device, it's dangerous, you should hang up now. Female voice 1: No, no. Male voice 1: This child."]. Optionally, the plot summary text can be obtained as: [In a dialogue, a woman (daughter) expresses extreme excitement on the phone, while a man (father) is driving a vehicle borrowed from a repair shop. Due to the lack of a secure in-vehicle communication device, the father, for safety reasons, repeatedly asks his daughter to hang up. However, the daughter refuses this request and insists on continuing the call, leaving the father feeling helpless].

[0413] For example, a built-in prompt for generating the plot summary text after the candidate frame is: [Please summarize the plot summary text based on "Male Voice 1: I'm wondering if there's a problem with the passenger airbag. I'm fine, but how did you get such a neck injury? Male Voice 2: It shouldn't be, it's just that my seatbelt stretched."] Optionally, the plot summary text could be: [Two men are having a conversation after a possible car accident. One suspects there might be a problem with the passenger airbag because the other has a neck injury. But the injured man explains that the injury is more likely due to the seatbelt suddenly tightening in an emergency.]

[0414] Step (6-2): The electronic device generates the linking text of the second media resource based on the plot summary text before the candidate frame, the plot summary text after the candidate frame, the content semantics of the candidate frame, and the content semantics of the second media resource.

[0415] In some possible implementations, the electronic device generates the linking text of the second media resource based on the plot summary text before the candidate frame, the plot summary text after the candidate frame, the content semantics of the video of the candidate frame, and the content semantics of the audio of the second media resource.

[0416] Optionally, the electronic device inputs the plot summary text before the candidate frame, the plot summary text after the candidate frame, the content semantics of the video of the candidate frame, and the content semantics of the audio of the second media resource into a natural language big data model. The natural language big data model processes the plot summary text before the candidate frame, the plot summary text after the candidate frame, the content semantics of the video of the candidate frame, and the content semantics of the audio of the second media resource to obtain the linking word text of the second media resource.

[0417] For example, when inputting the plot summary text before the candidate frame, the plot summary text after the candidate frame, the content semantics of the candidate frame's video, and the content semantics of the second media resource's audio into the natural language processing model, a built-in prompt can be provided as follows: [The expected embedded advertising text is ${×content semantics of the second media resource's audio×}, the plot summary text before the candidate frame is ${×plot summary text×}, the plot summary text after the candidate frame is ${×plot summary text×}, and the content semantics of the candidate frame's video is ${×tag words×}. Please focus on generating a connecting text of less than ${expected word count} based on the video frame scene of the candidate frame to ensure that the connecting text conforms to the video frame scene of the first media resource and that the advertising text connects naturally with the preceding and following text]. The expected word count can be user-defined, with a default value of 50.

[0418] For example, the built-in prompt for generating the connecting text for the second media resource is: [The desired embedded advertising slogan is: ×× Intelligent Driving Edition, intelligent driving is safe and reliable, giving you a new travel experience.] The plot summary text before the candidate frame is: In a conversation, a woman (daughter) expresses extreme excitement on the phone, while a man (father) is driving a car borrowed from a repair shop. Due to the lack of secure in-vehicle communication equipment, the father repeatedly asks his daughter to hang up the phone for safety reasons. However, the daughter refuses this request and insists on continuing the call, leaving the father feeling helpless. The plot summary text after the candidate frame is: Two men are talking after experiencing a possible car accident. One of them suspects that the passenger-side airbag may be faulty because the other person has a neck injury. However, the injured person explains that the injury is more likely due to the seatbelt suddenly tightening in an emergency. The semantic content of the candidate frame video is: car, road, collision. Please focus on generating a connecting text of no more than 50 characters based on the video frame scene of the candidate frame, to ensure that the connecting text conforms to the video frame scene of the first media resource, and that the advertising slogan connects naturally with the preceding and following text.] Optionally, the text of the connecting words can be obtained as follows: A sudden collision tests not only the driver's reaction but also the safety performance of the car.

[0419] In the above embodiments, a method for generating transition text for a second media resource is provided. Specifically, by generating plot summary text before and after candidate frames, and using the determined plot summary text before and after candidate frames, transition text for the second media resource is generated. This method can generate transition text that is relevant to both the plot content of the first and second media resources.

[0420] In some possible implementations, the embedding duration of the second media resource can be determined based on the length of the first media resource and user preferences. For example, the playback duration of the second media resource can be scaled by 0.5 to 2 times to obtain the embedding duration. Optionally, the electronic device can utilize a large natural language model to summarize or expand the audio recognition results and connecting words of the second media resource to a specified number of words, so as to control the embedding duration of the second media resource by controlling the number of words in the text.

[0421] The above example uses a candidate frame based on video plot content to illustrate the process of obtaining the corresponding connecting word text for that candidate frame. In other embodiments, for candidate frames based on preset templates, the preset connecting word text can be directly invoked.

[0422] For example, taking a music variety show video unrelated to vehicles as the primary media resource, assuming candidate frames are determined based on a scene transition template, the obtained transition text would be: "Every note leap is a perfect encounter between the heart and the melody. During this break from listening to music, let me recommend an excellent product." In this example, the preset template is stored in the electronic device in the form of {scene tag, transition text} pairs. An example of a scene transition-based template is as follows:

[0423] {

[0424] "category":"scene_based",

[0425] "label":"singing, performing, playing instruments, concert",

[0426] "caption": "A song has finished being performed, played, and played."

[0427] "connecting_words":"Every note leap is a perfect encounter between the soul and the melody. In the intervals between listening to music, we recommend {advertising subject type}."

[0428] "default":{"Advertiser Type":"An Excellent Product"}

[0429] }

[0430] For example, taking a music variety show video unrelated to vehicles as the primary media resource, suppose candidate frames are determined based on templates of facial expressions in the video frames, and the obtained connecting text is "While the protagonist is daydreaming, let me recommend an excellent product to everyone." In this example, the preset template is stored in the electronic device in the form of {facial expression tag, connecting text} pairs. An example of a template based on facial expressions in the video frames is as follows:

[0431] {

[0432] "category":"emotion_based",

[0433] "label":"Staring blankly",

[0434] "caption": "The character has a blank expression."

[0435] "connecting_words":"While I'm spacing out, I'd like to recommend {ad body type}",

[0436] "default":{"Advertiser Type":"An Excellent Product"}

[0437] }

[0438] It is worth noting that the linker text in the preset template can consist of preset text (i.e., the content shown in "{}") and optional filler text (i.e., the filler text for "{advertising subject type}"). In some possible implementations, this can be obtained by performing natural language processing (e.g., entity recognition) on the audio recognition results (e.g., advertising copy) of the second media resource, and generating quantifiers based on the entity recognition results. For example, taking a car company advertisement as an example, the extracted advertising subject type is "a car". For cases where recognition fails, the default field in the template is used directly for filling.

[0439] In this embodiment of the application, the connecting word text matches the content semantics of the first media resource and the content semantics of the second media resource. That is to say, the connecting word text not only connects with the plot content of the first media resource but also with the plot content of the second media resource, thereby improving the display effect of the target media resource.

[0440] Furthermore, after determining the linking text for the second media resource, the electronic device displays the linking text. At this point, the user can confirm whether the linking text is appropriate and perform a feedback operation. Correspondingly, the electronic device receives the user's feedback on the linking text and, in response, obtains second feedback information. This second feedback information instructs the user to confirm the linking text or the user-corrected linking text. Then, the electronic device can perform subsequent media resource processing based on the second feedback information. Thus, by displaying the linking text of the second media resource to the user, allowing them to choose whether to optimize or correct the linking text, the flexibility of media resource processing is improved, enabling the generation of more personalized target media resources based on user feedback.

[0441] S807: The electronic device obtains the aspect ratio and candidate resource area of ​​the second media resource based on the determined candidate frames.

[0442] Among them, the candidate resource area is the screen area in the first media resource used to display the second media resource, such as the screen area in the video resource used to embed advertising resources.

[0443] In some embodiments, the electronic device performs foreground detection on the video frames of the candidate frames to obtain the foreground detection results of the video frames. Based on the foreground detection results of the video frames, the aspect ratio of the video frames, the edge regions of the video frames, and the aspect ratio of the second media resource, the aspect ratio of the second media resource and the candidate resource region are determined.

[0444] Foreground detection results are used to indicate the foreground and background regions of a video frame. For example, the foreground detection result can be a trimap, a binary image indicating whether each pixel in the video frame belongs to the foreground, background, or an undefined region.

[0445] In some embodiments, the aspect ratio of the second media resource is obtained. Based on the foreground detection results of the video frame and a connected component analysis algorithm, multiple independent connected components are determined from the video frame. From these multiple independent connected components, connected components whose bounding box aspect ratios satisfy a specific aspect ratio and a specific area requirement are determined as background regions. Furthermore, within the background region, with the maximum aspect ratio of the second media resource as the optimization objective and the constraint that the second media resource does not touch the edge region or the foreground region of the video frame, the aspect ratio scaling ratio of the second media resource and candidate resource regions are determined based on the foreground detection results of the video frame, the aspect ratio of the video frame, the edge region of the video frame, and the aspect ratio of the second media resource.

[0446] For example, suppose the aspect ratio (i.e., width × height) of a video frame is W. k ×H k The aspect ratio of the second media resource is W. ad ×H ad Given a frame scaling factor of S, the actual width of the embedded second media resource is calculated to be W. ad '=S·W ad The actual height of the implanted second media resource is H. ad '=S·H ad The foreground detection result of a video frame can be represented by matrix M, where M(i,j) = 1 indicates that position (i,j) is the foreground region, and M(i,j) = 0 indicates that position (i,j) is the background region. A buffer region c is set to represent the edge region of the video frame. For aesthetic purposes, the constraint is set to not touch the edge region and the foreground region of the video frame. It is understandable that since a buffer region c is set, and the size of the buffer region c determines the position of matrix M, for example, c = 1 represents the outermost edge of the video frame, c = 2 represents the outermost two edges of the video frame, then matrix M can be transformed into M... c =dilate(M, c), where dilate is an expansion coefficient.

[0447] Based on the set constraints, with the maximum frame size of the second media resource as the optimization objective, the candidate resource region (x, y) and the frame scaling ratio S that satisfy the constraints are determined based on the following formula (1). Wherein, (x, y) refers to the coordinates of the upper left corner of the candidate resource region.

[0448]

[0449] In the formula, A represents the frame area of ​​the second media resource. According to formula (1), for each possible frame scaling ratio S, the possible (x, y) are traversed in descending order to determine whether the constraint condition is met. If the constraint condition is met, the traversal stops. Optionally, the traversal process can be accelerated and position preferences can be set by selecting the initial (x, y). In this way, the frame scaling ratio and candidate resource area that maximize the frame size of the second media resource can be determined to ensure the display effect of the subsequent second media resource.

[0450] For example, Figure 11 This is a schematic diagram of a candidate resource area for a second media resource provided in an embodiment of this application. Taking a TV drama video involving a vehicle collision as an example, the candidate frames of this TV drama video are the last frame of shot 11 or the first frame of shot 12. See also... Figure 11 The diagram shows the candidate resource region in the last frame of shot 11 and the candidate resource region in the first frame of shot 12. Since the area of ​​the candidate resource region in the first frame of shot 12 is larger, the first frame of shot 12 is selected as the final candidate frame.

[0451] Furthermore, after identifying candidate resource areas for the second media resource, the electronic device displays these candidate resource areas. At this point, the user can confirm whether the candidate resource areas are suitable and provide feedback. Correspondingly, the electronic device receives the user's feedback on the candidate resource areas and, in response, obtains third feedback information. This third feedback information instructs the user to confirm the candidate resource areas or to modify them. The electronic device can then perform subsequent media resource processing based on this third feedback information. Thus, by displaying candidate resource areas for the second media resource to the user, allowing them to choose whether to optimize or modify these areas, the flexibility of media resource processing is improved, enabling the generation of more personalized target media resources based on user feedback.

[0452] S808: The electronic device acquires the target audio of the second media resource based on the determined candidate frames.

[0453] Implementation Method 1: The electronic device extracts the audio features of an audio segment of a preset duration prior to the candidate frame. The audio content of the second media resource and the audio features of the audio segment are input into a speech synthesizer to obtain the target audio of the second media resource.

[0454] The preset duration is a fixed, pre-defined duration, such as 1 minute. The process by which an electronic device extracts audio features from an audio segment of the preset duration preceding a candidate frame is as follows: A human voice track of the preset duration preceding the candidate frame is acquired; noise reduction, echo cancellation, and pre-emphasis are applied to the acquired human voice track to obtain an audio segment. Based on a pre-trained speech feature extraction model, the audio segment is processed to extract its audio features. For example, the audio features of the audio segment include, but are not limited to, spectral envelope, fundamental frequency, formants, and prosodic patterns.

[0455] The target audio has the same audio features as the audio segment. Thus, by simulating the original audio of the first media resource and generating a target audio with the same features as the original audio, it is possible to ensure that the video style of the first media resource and the second media resource are consistent, and to ensure that the target media resource has a stronger sense of integration in terms of listening experience.

[0456] In some embodiments, the electronic device concatenates the audio content of the second media resource (such as the semantic content of the audio or the semantic content of a preset advertising audio) with the text of connecting words. Based on a pre-trained text feature extraction model, the electronic device processes the audio content and text of the second media resource to extract text features such as text phoneme features, semantic coding features, and BERT text features. Then, the extracted text features and the audio features of the aforementioned audio segment are input into a text-to-speech (TTS) synthesizer to generate a Mel spectrogram of the audio content and the text of connecting words. Finally, a vocoder converts the Mel spectrogram of the audio content and the text of connecting words into audio output, thus obtaining the target audio of the second media resource.

[0457] Alternatively, the text feature extraction model, speech synthesizer, and vocoder can all be trained independently on publicly available speech datasets. To reduce training costs, fine-tuning can also be performed based on open-source speech cloning and speech synthesis models, which can improve the consistency of style and timbre of the cloned speech across different sentences.

[0458] Method 2: The electronic device extracts the audio features of the preset audio, and inputs the audio content of the second media resource and the audio features of the preset audio into the speech synthesizer to obtain the target audio of the second media resource.

[0459] In this method, the target audio has the same audio features as the preset audio. Thus, by simulating the preset audio, a target audio with the same audio features as the preset audio is generated, enabling the generation of audio in the desired style and achieving personalized generation of the target audio.

[0460] S809. The electronic device adds the second media resource to the first media resource based on the candidate frames of the first media resource, the connecting word text of the second media resource, the candidate resource region of the second media resource, and the target audio of the second media resource, to obtain the target media resource.

[0461] In some possible implementations, the electronic device embeds the target audio of the second media resource into the first media resource, blending it with the original audio track of the first media resource. Furthermore, the electronic device embeds the second media resource into the candidate resource region of the first media resource in a picture-in-picture format, according to the aspect ratio.

[0462] Optionally, during audio track merging, the background audio track of the first media resource is copied, and the target audio track of the second media resource is mixed with the background audio track to obtain a mixed audio track. This mixed audio track is then embedded into the original audio track of the first media resource. Alternatively, during audio track merging, the background audio track of the first media resource is copied, its volume is first reduced, and then the audio track of the second media resource is mixed with the reduced background audio track to obtain a mixed audio track. This mixed audio track is then embedded into the original audio track of the first media resource. An audio track mixer can be used to perform the above mixing process.

[0463] Optionally, during video embedding, the duration of embedding the video frame corresponding to the candidate frame into the audio track is fixed, the aspect ratio of the second media resource is reduced and adjusted based on the aspect ratio of the second media resource, and the second media resource is embedded into the candidate resource region of the candidate frame of the first media resource using masking technology.

[0464] In some possible implementations, the electronic device can adjust the transparency of the second media resource and smooth its edges so that the second media resource blends naturally with the background area of ​​the first media resource.

[0465] In some possible implementations, the electronic device can generate captions of connecting words and set the captions to a preset state. This preset state could include a preset color, a preset style, etc.

[0466] S810, electronic devices play target media resources.

[0467] In some possible embodiments, when candidate frames of a first media resource are determined, the electronic device playing the target media resource can be replaced by playing a second media resource based on the candidate frames within the first media resource. Taking an image resource as an example, the corresponding process could be: playing the second media resource within the candidate frames of the first media resource. Alternatively, taking a video resource as an example, the corresponding process could be: playing the second media resource within multiple video frames of the first media resource, each starting with a candidate frame. It is understood that the duration of the multiple video frames starting with a candidate frame is the same as the duration of the second media resource.

[0468] In some other possible embodiments, when candidate frames of the first media resource and connecting words of the second media resource are determined, the electronic device playing the target media resource can be replaced by: playing the second media resource and its connecting words based on the candidate frames in the first media resource. Taking the second media resource as an image resource as an example, the corresponding process can be: playing the second media resource and its connecting words in the candidate frames of the first media resource. Or, taking the second media resource as a video resource as an example, the corresponding process can be: playing the second media resource and its connecting words in multiple video frames in the first media resource, with the candidate frames as the starting frames.

[0469] In other possible embodiments, given that candidate frames of the first media resource, the aspect ratio of the second media resource, and candidate resource regions are determined, the electronic device playing the target media resource can be replaced by: playing the second media resource within the candidate resource regions determined based on the candidate frames in the first media resource, according to the aspect ratio of the second media resource. Taking an image resource as an example, the corresponding process could be: playing the second media resource within the candidate resource regions of the candidate frames of the first media resource, according to the aspect ratio of the second media resource. Alternatively, taking a video resource as an example, the corresponding process could be: playing the second media resource within the candidate resource regions of multiple video frames in the first media resource, starting with the candidate frames, according to the aspect ratio of the second media resource.

[0470] In this way, the target media resource can be played automatically, allowing users to preview the effect of adding a second media resource to the first media resource.

[0471] In other embodiments, the electronic device may display a file of the target media resource, and the user may trigger the electronic device to play the target media resource by performing a playback operation on the file of the target media resource, so that the user can preview the effect of adding a second media resource to the first media resource.

[0472] The technical solution provided in this application can automatically determine candidate frames from the first media resource that match the plot content of the second media resource, eliminating the need for video creators to manually understand the plot content of the first media resource to determine candidate frames, thus greatly reducing labor and time costs. Furthermore, by adding the second media resource to the first media resource based on candidate frames to obtain the target media resource, the second media resource can be added naturally and smoothly to the first media resource in a more efficient and economical way, resulting in better integration of the first and second media resources and ensuring a natural and smooth fusion of media resources.

[0473] The following section uses an advertiser's decision-making scenario as an example to explain the process of media resource processing in detail. Figure 12 This is a flowchart illustrating a resource playback method in an advertiser decision-making scenario, as provided in an embodiment of this application. See also... Figure 12 Taking the interaction process between the terminal and the server as an example, the method includes the following S1201-S1213.

[0474] S1201, The development terminal displays the first interface.

[0475] S1202, The development terminal receives the first media resource and the second media resource input by the user on the first interface.

[0476] For example, this application embodiment takes a car manufacturer advertiser decision-making scenario as an example. The car manufacturer advertiser can select multiple primary media resources on the video platform of the development terminal. Optionally, the car manufacturer advertiser can perform preliminary filtering from the media resources provided by the video platform based on the content tags set by the video platform, such as "car", to obtain multiple primary media resources related to "car". Alternatively, the car manufacturer advertiser can select videos from top creators and other creators that meet the advertiser's requirements from the media resources provided by the video platform as primary media resources.

[0477] In the context of automotive advertisers' decision-making, the primary media resource selected by an advertiser can be a TV drama video involving vehicle collisions, or a music variety show video unrelated to vehicles. This application will subsequently use these two primary media resources as examples to illustrate the solution.

[0478] In the context of car manufacturers' advertising decisions, the second media resource can be advertising videos used to promote electric vehicles. The content of these videos presents intelligent driving electric vehicles, with the tagline "XX (electric vehicle brand) Intelligent Driving Edition, intelligent driving is safe and reliable, giving you a new travel experience."

[0479] In this embodiment, for each first media resource selected by the advertiser, a second media resource needs to be embedded into the first media resource using the resource playback method provided in this embodiment, thereby generating a target media resource, i.e., the embedded effect video. Then, the order of the association between the target media resources generated from each first media resource and the advertiser's advertising content is returned, thereby assisting the advertiser in making placement decisions. This embodiment will subsequently use the embedding of a second media resource into a single first media resource as an example to describe the media resource processing flow. It is understood that the embedding of multiple first media resources can be achieved by repeatedly executing the process shown in this embodiment or by executing the process shown in this embodiment in parallel to obtain the target media resources generated for each first media resource.

[0480] S1203: The development terminal responds to the input first media resource and second media resource, performs a legality check on the first media resource and the second media resource, and if both the first media resource and the second media resource pass the check, then execute S1204.

[0481] The content of S1203 is the same as that shown in S803 above, and will not be repeated here.

[0482] S1204. The development terminal sends a request to the server to add a second media resource to the first media resource.

[0483] S1205. In response to the request, the server obtains the content semantics of the first media resource and the content semantics of the second media resource.

[0484] The content of S1205 is the same as that shown in S804 above, and will not be repeated here.

[0485] S1206. The server determines candidate frames from the first media resource based on the content semantics of the first media resource and the content semantics of the second media resource.

[0486] The content of S1206 is the same as that shown in S805 above, and will not be repeated here.

[0487] For example, in the decision-making scenario of car manufacturers' advertisers, taking a TV drama video involving vehicle collisions as the primary media resource, since the TV drama video contains vehicle-related content, there may be candidate frames based on the video's plot. It is worth noting that in the car manufacturer's advertiser decision-making scenario, the position with the highest content similarity score is selected for insertion by default.

[0488] In some possible implementations, if there are no candidate frames in the first media resource that meet the similarity condition with the second media resource, the server can determine video frames of a preset content type from the first media resource as candidate frames based on the content semantics of the first media resource.

[0489] For example, in the decision-making scenario of car company advertisers, taking music variety shows unrelated to vehicles as the primary media resource, each song can be clustered into a scene based on a scene transition template. The end of the song marks the scene transition point, which can match the scene transition template in the preset template. Therefore, the last frame of the scene to which the song ends can be used as a candidate frame. Thus, for cases where the content similarity is too low to generate candidate frames, a common candidate frame in the preset template can be matched.

[0490] Understandably, since candidate frames based on video story content are usually more natural and relevant than candidate frames based on preset templates, in the decision-making scenario of car advertisers, when there are candidate frames based on video story content in the primary media resource, there is no need to determine candidate frames based on preset templates.

[0491] S1207. Based on the determined candidate frames, the server obtains the linking text of the second media resource.

[0492] The content of S1207 is the same as that shown in S806 above, and will not be repeated here.

[0493] S1208. Based on the determined candidate frames, the server obtains the aspect ratio and candidate resource area of ​​the second media resource.

[0494] The content of S1208 is the same as that shown in S807 above, and will not be repeated here.

[0495] S1209. The server obtains the target audio of the second media resource based on the determined candidate frames.

[0496] In the decision-making scenario of car manufacturers' advertisers, the audio features of audio segments of a preset duration preceding candidate frames can be extracted. The audio content of the second media resource, the text of connecting words, and the audio features of the audio segments are then input into a speech synthesizer to obtain the target audio for the second media resource. In this way, by generating target audio with the same audio features as the audio segments in the first media resource, the integration of the first and second media resources is improved.

[0497] The content of S1209 is the same as that shown in S808 above, and will not be repeated here.

[0498] S1210. The server adds the second media resource to the first media resource based on the candidate frames of the first media resource, the connecting words text of the second media resource, the candidate resource area of ​​the second media resource, and the target audio of the second media resource, to obtain the target media resource.

[0499] In some embodiments, after the server adds a second media resource to the first media resource, it re-encodes the video frames after adding the second media resource to obtain a new video file, which is the target media resource, so that it can be returned to the development terminal for users to select and use.

[0500] S1211, The server sends the target media resources to the development terminal.

[0501] S1212, The development terminal receives target media resources from the server.

[0502] S1213, Develop the terminal to play target media resources.

[0503] In the above Figure 12 In the illustrated embodiment, after processing the multiple first media resources selected by the car manufacturer advertiser, the target media resources generated corresponding to the multiple first media resources can be played on the development terminal. This means playing the effect of embedding car manufacturer advertisements in different video resources, allowing the car manufacturer advertiser to view the processing results and make a selection. For example, Figure 13 This is a schematic diagram illustrating an example of media resource embedding provided in an embodiment of this application. See also... Figure 13 For TV drama videos involving vehicle collisions, see [link to relevant documentation]. The effectiveness of media placement in these videos is discussed below. Figure 13 As shown in (13-1), candidate video 1 refers to a TV drama video involving a vehicle collision; the advertising placement point refers to the candidate frame (e.g., the first frame of shot 12 in candidate video 1); the advertising linking phrase refers to the linking text of the second media resource (e.g., "A sudden collision tests not only the driver's reaction but also the car's safety performance"); the advertising embedding area refers to the candidate resource area of ​​the second media resource; and the advertising slogan refers to the text information of the second media resource (e.g., "×× Intelligent Driving Edition, intelligent driving, safe and reliable, giving you a new travel experience"). For music variety shows unrelated to vehicles, the effect of media resource placement is shown in [reference needed]. Figure 13As shown in (13-2), candidate video 2 refers to music variety shows unrelated to vehicles; advertising placement points refer to candidate frames (such as the video frame corresponding to the end of a song in a singing scene); advertising linking words refer to the linking text of the second media resource (such as "Every note leap is a perfect encounter between the heart and the melody. In the interval of listening to music, we recommend an excellent product"); advertising embedding area refers to the candidate resource area of ​​the second media resource; and advertising slogans refer to the text information of the second media resource (such as "×× Intelligent Driving Edition, intelligent driving, safe and reliable, giving you a new travel experience"). It can be seen that the placement effect of TV drama videos involving vehicle collisions is better. Therefore, car advertisers can use TV drama videos involving vehicle collisions as the videos for successful placement.

[0504] For example, Figure 14 This is a flowchart illustrating the module interaction in an advertiser's decision-making scenario, as provided in an embodiment of this application. See also... Figure 14 Taking the interaction between the interaction module, audio / video parsing module, content understanding module, audio generation module, and audio / video embedding module as an example, the interaction process includes: the advertiser uploads video resources and advertising resources through the interaction module. The video resources are the first media resources, and the advertising resources are the second media resources. The interaction module then verifies the legality of both the first and second media resources. If both the first and second media resources pass verification, the interaction module triggers the audio / video parsing module to initiate the process of adding the second media resource to the first media resource. Next, the audio / video parsing module performs audio / video parsing on the first and second media resources to obtain the audio content semantics of the first and second media resources, and sends the audio content semantics of the first and second media resources to the content understanding module. The content understanding module performs video content understanding-related processing, identifying the ad insertion point (i.e., candidate frame), ad linking words (i.e., linking word text), and ad embedding area (i.e., candidate resource area). The content understanding module then sends the ad insertion point, ad linking words, and ad embedding area to the audio / video embedding module. Furthermore, the content understanding module also sends the advertising audio content (i.e., the audio content of the second media resource, such as the semantic meaning of the audio) and advertising linking words to the audio generation module, triggering the audio generation module to generate target audio including the advertising audio content and advertising linking words through speech cloning and speech synthesis. Then, the audio and video embedding module performs advertising embedding based on the advertising insertion point, advertising linking words, advertising embedding area, and target audio, obtaining the effect video after advertising embedding, and returns the effect video after advertising embedding to the development terminal.

[0505] The technical solution provided in this application can automatically determine candidate frames from the first media resource that match the plot content of the second media resource, eliminating the need for video creators to manually understand the plot content of the first media resource to determine candidate frames, thus greatly reducing labor and time costs. Furthermore, by adding the second media resource to the first media resource based on candidate frames to obtain the target media resource, the second media resource can be added naturally and smoothly to the first media resource in a more efficient and economical way, resulting in better integration of the first and second media resources and ensuring a natural and smooth fusion of media resources.

[0506] The following section uses the scenario of video creators as an example to explain the process of media resource processing in detail. Figure 15 This is a flowchart illustrating a resource playback method in a video creator scenario, provided as an embodiment of this application. See also... Figure 15 Taking the interaction process between the terminal and the server as an example, the method includes the following S1501-S1518.

[0507] S1501, The development terminal displays the first interface.

[0508] S1502, The development terminal receives the first media resource and the second media resource input by the user on the first interface.

[0509] S1503: The development terminal responds to the input first media resource and second media resource, performs a legality check on the first media resource and the second media resource. If both the first media resource and the second media resource pass the check, then S1504 is executed.

[0510] S1504. The development terminal sends a request to the server to add a second media resource to the first media resource.

[0511] S1505. In response to the request, the server obtains the content semantics of the first media resource and the content semantics of the second media resource.

[0512] S1506. The server determines candidate frames from the first media resource based on the content semantics of the first media resource and the content semantics of the second media resource.

[0513] It is worth noting that in the context of video creators, it is possible to determine both candidate frames based on the video's plot content and candidate frames based on preset templates, thus giving video creators more creative freedom.

[0514] S1507. Based on the determined candidate frames, the server obtains the linking text of the second media resource.

[0515] S1508. Based on the determined candidate frames, the server obtains the aspect ratio and candidate resource area of ​​the second media resource.

[0516] S1509, The server sends the candidate frames of the first media resource, the linking text of the second media resource, and the candidate resource area of ​​the second media resource to the development terminal.

[0517] S1510, The development terminal receives and displays candidate frames of the first media resource, the connecting word text of the second media resource, and the candidate resource area of ​​the second media resource from the server.

[0518] In some possible implementations, there are multiple candidate frames. When displaying candidate frames from the first media resource, the development terminal displays multiple candidate frames in descending order of their matching degree with the second media resource. This helps video creators decide whether to optimize and modify the candidate frames.

[0519] The development terminal can be configured with a function option to display matching degree sorting. If the user enables this function, multiple candidate frames can be displayed in descending order of matching degree. If the user disables this function, the first candidate frame in the descending order of matching degree will be selected by default, and the subsequent media resource processing flow will be executed.

[0520] The above implementation provides a method for displaying multiple candidate frames based on their matching degree. Displaying candidate frames in descending order of their matching degree with the second media resource allows users to quickly identify the most suitable candidate frames, improving human-computer interaction efficiency.

[0521] In some possible implementations, there are multiple linking texts. When displaying linking texts for the second media resource, the development terminal displays multiple linking texts in descending order of their semantic match with the content of the first and second media resources. This helps video creators decide whether to optimize or modify the linking texts.

[0522] The development terminal can be configured with a function option to display matching degree sorting. If the user enables this function, multiple linking word texts can be displayed in descending order of matching degree. If the user disables this function, the first linking word text in the descending order of matching degree will be selected by default, and the subsequent media resource processing flow will be executed.

[0523] Among the above implementation methods, a way to display multiple linking word texts based on matching degree is provided. Specifically, displaying the linking word texts in descending order of matching degree with the content semantics of the first media resource and the second media resource allows users to quickly identify the more matching linking word texts, improving human-computer interaction efficiency.

[0524] In some possible implementations, there are multiple candidate resource regions. When displaying candidate resource regions of the second media resource, the development terminal displays multiple candidate resource regions in descending order of their aspect ratio.

[0525] The development terminal can be configured with a function option to sort and display image sizes. If the user enables this function, multiple candidate resource areas can be displayed in descending order of image size. If the user disables this function, the first candidate resource area in the descending order of image size will be selected by default, and the subsequent media resource processing flow will be executed.

[0526] Among the above implementation methods, a way to display multiple candidate resource areas based on the image size is provided. Displaying the candidate resource areas in descending order of image size allows users to quickly identify the larger candidate resource areas, improving human-computer interaction efficiency.

[0527] Furthermore, the terminal categorizes and displays candidate frames based on different types of video plot content and candidate frames based on preset templates. In the display area for candidate frames based on video plot content, multiple candidate frames are displayed in descending order of similarity to the second media resource. Similarly, in the display area for candidate frames based on preset templates, multiple candidate frames are displayed in descending order of similarity to the second media resource.

[0528] For example, Figure 16 This is a schematic diagram illustrating the interface display in a video creator scenario, provided as an embodiment of this application. See also... Figure 16 This demonstrates the development terminal used to display candidate frames (such as...). Figure 16 The advertised placement locations shown), and the connecting text (such as...) Figure 16 (shown as conjunctions) and candidate resource areas (such as...) Figure 16 The interface of the ad embedding area shown. In some embodiments, when displaying the ad embedding point, the development terminal can display the ad embedding point on the time track of the first media resource, such as... Figure 16 The "Location 1601" shown can be displayed, or the timestamp corresponding to the ad placement location can be displayed directly, such as... Figure 16The displayed "Ad placement location: 00:23:25" indicates the ad placement location. In some embodiments, the development terminal can display the placement type of each ad placement location, such as... Figure 16 The options shown are "Implantation Point Type: Plot Content", "Implantation Point Type: Scene Template", and "Implantation Point Type: Expression Template". In some embodiments, the development terminal can display the candidate frames, connecting text, and candidate resource areas in descending order of their semantic matching with the content of the second media resource. For example, if the matching degree is... Figure 16 The "relevance" is shown. Furthermore, a recommendation index can be displayed to the user, such as using the number of stars to indicate the recommendation index.

[0529] In this way, users can modify the candidate frames, linking word text, and candidate resource areas displayed on the development terminal, gaining greater autonomy and enabling more interactivity and personalization. It is worth noting that the function of displaying candidate frames, linking word text, and candidate resource areas can be adaptively enabled or disabled based on user needs.

[0530] S1511. The development terminal receives user feedback operations on candidate frames of the first media resource, connecting words of the second media resource, and candidate resource areas of the second media resource, and obtains user feedback information.

[0531] In some embodiments, in response to a feedback operation on a candidate frame of a first media resource, the development terminal obtains first feedback information, which is used to instruct the user to confirm the candidate frame or the video frame corrected by the user. Subsequently, the media resource processing flow can be executed based on the first feedback information.

[0532] In some embodiments, in response to feedback operations on the linking text of the second media resource, the development terminal obtains second feedback information, which is used to instruct the user to confirm the linking text or the user-corrected text. Subsequently, the media resource processing flow can be executed based on the second feedback information.

[0533] In some embodiments, in response to feedback operations on candidate resource regions of a second media resource, the development terminal obtains third feedback information, which is used to instruct the user to confirm the candidate resource region or the region corrected by the user. Subsequently, the media resource processing flow can be executed based on the third feedback information.

[0534] S1512, The development terminal returns user feedback information to the server.

[0535] S1513. The server receives user feedback information returned by the development terminal.

[0536] S1514. The server obtains the target audio of the second media resource based on the determined candidate frames.

[0537] In some embodiments, the server can extract audio features of a preset audio source, input the audio recognition result of the second media resource and the audio features of the preset audio source into a speech synthesizer to obtain the target audio of the second media resource. For example, it supports voice cloning based on the voiceprint features of the video creator to generate personalized advertising voice. As another example, it supports video creators uploading a 1-minute audio file (such as a WAVE file) for specific voice cloning. Before performing voice cloning, a pop-up prompts the user whether to upload their personal audio; if the user clicks upload, the uploaded audio is used for voice cloning.

[0538] S1515, The server adds the second media resource to the first media resource based on the candidate frames of the first media resource, the connecting words text of the second media resource, the candidate resource region of the second media resource, and the target audio of the second media resource, to obtain the target media resource.

[0539] S1516. The server sends the target media resources to the development terminal.

[0540] S1517. The development terminal receives target media resources from the server.

[0541] S1518, Develop a terminal to play target media resources.

[0542] In some embodiments, the development terminal plays a second media resource based on a target frame within the first media resource, based on the user's feedback operation on candidate frames of the first media resource. Where the feedback operation is used to instruct the user to confirm a candidate frame, the target frame is the candidate frame. Where the feedback operation is used to instruct the user to correct a video frame, the target frame is the video frame corrected by the user.

[0543] In some embodiments, the development terminal plays the target text in the first media resource based on the user's feedback operation on the linking text of the second media resource. Wherein, if the feedback operation is used to instruct the user to confirm the linking text, the target text is the linking text; if the feedback operation is used to instruct the user to correct the text, the target text is the text corrected by the user.

[0544] In some embodiments, the development terminal plays the second media resource in the target area of ​​the first media resource, according to the scaling ratio of the second media resource, based on the user's feedback operation on the candidate resource area of ​​the second media resource. Wherein, when the feedback operation is used to instruct the user to confirm the candidate resource area, the target area is the candidate resource area; when the feedback operation is used to instruct the user to correct the area, the target area is the area corrected by the user.

[0545] For example, Figure 17 This is a flowchart illustrating the module interaction in a video creator scenario, as provided in an embodiment of this application. See also... Figure 17 Taking the interaction between the interaction module, audio / video parsing module, content understanding module, audio generation module, and audio / video embedding module as an example, the interaction process includes: The video creator uploads video resources and advertising resources through the interaction module. The video resources are the first media resources, and the advertising resources are the second media resources. The interaction module then verifies the legality of both the first and second media resources. If both pass verification, the interaction module triggers the audio / video parsing module to add the second media resource to the first media resource. Next, the audio / video parsing module performs audio / video parsing on the first and second media resources to obtain the audio content semantics of both resources, and sends this semantics to the content understanding module. The content understanding module performs video content understanding-related processing, identifying the advertising insertion point (candidate frame), advertising linking words (linking word text), and advertising embedding area (candidate resource area). The content understanding module then sends the advertising insertion point, advertising linking words, and advertising embedding area to the interaction module. The interaction module receives feedback from video creators, obtains user feedback information, and sends this information to the content understanding module. The content understanding module updates its settings based on this feedback, resulting in updated ad placement locations, ad linking keywords, and ad embedding areas. This updated information is then sent to the audio / video embedding module. Furthermore, the content understanding module sends the updated ad audio content (i.e., the audio content of the secondary media resource, such as the semantic meaning of the audio) and ad linking keywords to the audio generation module. This triggers the audio generation module to generate target audio, including the ad audio content and ad linking keywords, through speech cloning and speech synthesis. Finally, the audio / video embedding module executes ad placement based on the updated placement locations, ad linking keywords, ad embedding areas, and target audio, producing a post-ad placement video, which is then returned to the development terminal.

[0546] The technical solution provided in this application can automatically determine candidate frames from the first media resource that match the plot content of the second media resource, eliminating the need for video creators to manually understand the plot content of the first media resource to determine candidate frames, thus greatly reducing labor and time costs. Furthermore, by adding the second media resource to the first media resource based on candidate frames to obtain the target media resource, the second media resource can be added naturally and smoothly to the first media resource in a more efficient and economical way, resulting in better integration of the first and second media resources and ensuring a natural and smooth fusion of media resources.

[0547] Compared with related technologies, the embodiments of this application have several advantages. Firstly, in the process of embedding advertising resources into video resources, one implementation does not provide advertising placement points related to the video's plot, resulting in low relevance between the advertising placement and the video content. Secondly, another implementation requires manual intervention to understand the video content and select suitable advertising placement points, leading to high labor and time costs. In this application, the embodiments automatically identify advertising placement points based on video content understanding technology and generate advertising linking words that connect with the plot, ensuring a logical connection between the advertising placement and the original video in terms of content or plot, while also reducing the cost of manual screening.

[0548] On the other hand, during the process of embedding advertising resources into video resources, the inconsistency in style and content between the advertising video and the original video may cause viewers to experience a sense of disconnect or discomfort. For example, the integration of the advertising with the original video is poor in terms of auditory perception. In this embodiment, based on speech cloning and speech synthesis technology, advertising audio that is consistent with or has the expected style of the original video is generated, which can reduce the discomfort or abruptness caused by advertising embedding to video viewers.

[0549] On the other hand, the integration of video resources and advertising resources in related technologies is poor, with issues such as obscuring or directly embedding advertisements into the original video, affecting the continuity of the original video. Furthermore, the size of the embedded advertising resources cannot be adaptively adjusted, failing to fully and rationally utilize the space of the video resources. In this embodiment, the advertising resources are embedded in a picture-in-picture manner, without obscuring the original video subject, thus rationally utilizing the space of the video resources, increasing the exposure area while minimizing the impact on the original video.

[0550] Figure 18 This is a flowchart illustrating another resource playback method provided in an embodiment of this application. In some possible implementations, this resource playback method can be executed by the application terminal in the above system architecture, see [link to relevant documentation]. Figure 18 Using the application terminal as the execution subject, the method includes the following S1801-S1803.

[0551] S1801, The application terminal displays the identifier of the target film and television resource.

[0552] Here, "target video resource" refers to the video resource to be played. For example, the target video resource is the video resource displayed in a video playback application, such as the video resource displayed on the main interface of the video playback application. The identifier of the target video resource can be a text identifier, such as the video title, or an image identifier, such as a cover image. It is understood that the target video resource is based on the above... Figure 7 , Figure 8 , Figure 12 or Figure 15 The steps shown in the embodiment are generated. For example, through the above... Figure 7 , Figure 8 , Figure 12 or Figure 15 After generating the target media resource as shown in the embodiment, the generated target media resource can be uploaded to the video playback application. Then, the video playback application can display the target media resource and provide the playback function of the target media resource.

[0553] The target film and television resource includes a first segment and a second segment. The second segment is either the segment preceding or following the first segment. That is, the first and second segments are two adjacent segments. In this embodiment, a segment refers to a video clip of the target film and television resource, which can be a video clip related to the plot, such as a shot clip, scene clip, or transition clip. A shot clip refers to a clip containing the content of the scene within the same camera's perspective. A scene clip refers to a clip used to show the location and environment where a relatively complete storyline or event takes place. It is understood that a scene clip is composed of multiple shot clips. A transition clip refers to the transition section used to connect two different scene clips or shot clips. It is understood that... Figure 18 The segments mentioned in the embodiments are segments of the main video in the target film and television resource. The main video here can correspond to the above. Figure 7 , Figure 8 , Figure 12 or Figure 15 The first media resource mentioned in the embodiment. The first segment of the target film / television resource can be based on the above. Figure 7 , Figure 8 , Figure 12 or Figure 15 The candidate frames mentioned in the embodiments are determined as segments, such as segments determined by using candidate frames as starting frames.

[0554] The first segment embeds video. This video showcases product features that match the storyline of either the first or second segment. Exemplarily, the video can be a dynamic video or a still image with audio. In this embodiment, product features refer to the special properties or functions possessed by a product, used to describe the product's uniqueness and advantages. Exemplarily, product features can be physical properties, such as the product's size, shape, color, or material. Also exemplarily, product features can be functional, meaning the product has specific functions or performance that meet consumer needs. It is understood that... Figure 18 The video mentioned in the embodiment is an advertising video, which can correspond to the above. Figure 7 , Figure 8 , Figure 12 or Figure 15 The second media resource mentioned in the embodiments.

[0555] In some possible implementations, video is embedded in the background area of ​​the first segment. The background area refers to the portion of the first segment's frame excluding the foreground area (i.e., the region of interest). Thus, by embedding video in the background area of ​​the first segment, obstruction of the foreground area is avoided, allowing users to watch the embedded video without affecting their viewing of the first segment, thereby improving the viewing experience. Understandably, Figure 18 The background area mentioned in the embodiment may correspond to the above. Figure 8 , Figure 12 or Figure 15 The background area mentioned in the embodiment. The embedding location of the video can be based on the above. Figure 8 , Figure 12 or Figure 15 The candidate resource regions mentioned in the embodiments were determined.

[0556] S1802, The application terminal receives the user's playback operation for the identifier.

[0557] In some embodiments, the playback operation may be a trigger operation on the identifier, such as a single click or double click on the identifier. In other embodiments, the identifier may include a playback control, and the playback operation may be a trigger operation on the playback control, such as a single click, long press, or double click on the playback control.

[0558] S1803: The application terminal responds to the playback operation and plays the target video resource.

[0559] The first video clip and the embedded video are played synchronously, meaning that the video plays simultaneously with the clip. This ensures that both the first video clip and the embedded video are playing, preventing one from playing while the other is frozen. This allows users to flexibly choose to watch either the first video clip or the embedded video, avoiding a degraded viewing experience caused by the embedded video freezing the original clip.

[0560] In some embodiments, when the application terminal plays the first segment, it simultaneously plays the embedded video in a picture-in-picture format. This picture-in-picture playback avoids obscuring the original subject in the first segment and reduces the impact on the viewer's experience, thus improving the viewing experience.

[0561] It is worth noting that, Figure 18 The illustrated embodiment uses a picture-in-picture format, where video is embedded within the first segment, to illustrate the resource playback method. In other embodiments, the application terminal can also play the target video resource by inserting video between the first and second segments.

[0562] In addition, in some other embodiments, the application terminal also displays target text. This target text is used to introduce product features in conjunction with the plot of the first or second segment. That is, the content of the target text not only matches the plot of the first or second segment but also matches the product features. This allows for the introduction of product features while connecting the first or second segment, effectively improving the viewing experience for video viewers. For example, the application terminal can display the target text while playing the first segment. Alternatively, the application terminal can display the target text before playing the first segment. It is understood that the target text can be based on the above... Figure 8 , Figure 12 or Figure 15 The text of the connectors mentioned in the examples is determined.

[0563] Furthermore, the application terminal also plays the audio corresponding to the target text. For example, the audio characteristics of this audio are the same as those of the first or second segment. Thus, by playing audio with the same audio characteristics as the first or second segment, the smoothness of the target video resource's listening experience can be ensured. Again, for example, the audio characteristics of this audio are the same as those of a preset audio. This allows for the generation of audio in the desired style, achieving personalized audio generation. It is understood that the playback of the audio corresponding to the target text is synchronized with the display of the target text. It is understood that the audio corresponding to the target text can be based on the above... Figure 8 , Figure 12 or Figure 15 The target audio mentioned in the example was determined.

[0564] The technical solution provided in this application embodiment can simultaneously play the first segment of the target film and television resource and the video embedded in the first segment. On the one hand, since the product characteristics displayed in the video match the plot of the first or second segment (i.e., the adjacent segment of the first segment), the smoothness of the playback of the target film and television resource can be ensured, avoiding a disjointed feeling for the user when watching the film and television resource. On the other hand, by embedding the video in the frame of the first segment and playing the first segment and the video simultaneously, it can be ensured that the first segment and the embedded video play synchronously, so that the user can watch the embedded video without affecting the viewing of the first segment, thereby improving the viewing experience of the video viewer.

[0565] It should be noted that the above description is for the purpose of more clearly explaining the resource playback method described in the embodiments of this disclosure, and should not be construed as a limitation on the specific implementation of this application.

[0566] The above mainly describes the solutions provided by the embodiments of this application from the perspective of processing flow. Correspondingly, the embodiments of this application also provide a resource playback device for implementing the various methods described above. This resource playback device can be one of the devices described in the above method embodiments, or include the aforementioned devices, or be a usable component. It is understood that, in order to achieve the above functions, the resource playback device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0567] This application embodiment can divide the resource playback device into functional modules according to the above method embodiment. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be understood that the module division in this application embodiment is illustrative and is only a logical functional division. In actual implementation, there may be other division methods.

[0568] For example, Figure 19 This is a schematic diagram of a resource playback device provided in an embodiment of this application. See also... Figure 19The resource playback device includes a display module 1901, an input module 1902, and a playback module 1903. Wherein:

[0569] Display module 1901 is used to perform the above. Figure 7 The S701 shown or Figure 8 The S801 or shown Figure 12 S1201 or shown Figure 15 S1501 as shown;

[0570] Input module 1902 is used to perform the above. Figure 7 The S702 or shown Figure 8 The S802 or shown Figure 12 S1202 or shown Figure 15 S1502 as shown;

[0571] Playback module 1903 is used to perform the above. Figure 7 The S703 or shown Figure 8 The S810 or shown Figure 12 S1213 or shown Figure 15 S1518 is shown.

[0572] In some possible implementations, the resource playback device also includes a verification module for performing the above-mentioned... Figure 8 The S803 or shown Figure 12 S1203 or shown Figure 15 S1503 is shown.

[0573] In some possible implementations, the resource playback device also includes a content understanding module for performing the above-mentioned functions. Figure 8 The S804 or shown Figure 12 S1205 or shown Figure 15 S1505 is shown.

[0574] In some possible implementations, the resource playback device further includes a candidate frame determination module for performing the above-mentioned... Figure 8 The S805 or shown Figure 12 S1206 or shown Figure 15 S1506 is shown.

[0575] In some possible implementations, the resource playback device also includes a text determination module for performing the above-mentioned... Figure 8 The S806 or shown Figure 12 S1207 or shown Figure 15 S1507 is shown.

[0576] In some possible implementations, the resource playback device further includes a region determination module for performing the above-mentioned... Figure 8The S807 or shown Figure 12 S1208 or shown Figure 15 S1508 is shown.

[0577] In some possible implementations, the resource playback device further includes an audio determination module for performing the above-mentioned... Figure 8 The S808 or shown Figure 12 S1209 or shown Figure 15 S1514 is shown.

[0578] In some possible implementations, the resource playback device further includes a resource generation module for performing the above-mentioned... Figure 8 S809 or shown Figure 12 S1210 or shown Figure 15 S1515 is shown.

[0579] For example, Figure 20 This is a schematic diagram of a resource playback device provided in an embodiment of this application. See also... Figure 20 The resource playback device includes a display module 2001, a receiving module 2002, and a playback module 2003. Wherein:

[0580] Display module 2001 is used to perform the above. Figure 18 S1801 as shown;

[0581] Receiver module 2002 is used to perform the above. Figure 18 S1802 as shown;

[0582] Playback module 2003 is used to perform the above. Figure 18 S1803 is shown.

[0583] For a detailed description of the above-mentioned optional methods, please refer to the foregoing method embodiments, which will not be repeated here. Furthermore, the explanation of any of the resource playback devices provided above and the description of their beneficial effects can be found in the corresponding method embodiments described above, which will not be repeated here.

[0584] As an example, combined Figure 5 The above Figure 19 The display module 1901, input module 1902, and playback module 1903 in the resource playback device shown, or the aforementioned Figure 20 The functions implemented by some or all of the display module 2001, receiving module 2002, and playback module 2003 in the resource playback device shown can be achieved through... Figure 5 The processor 510 in the middle executes Figure 5 The computer executes instructions in the internal memory 521.

[0585] In this embodiment, the resource playback device is presented in an integrated manner, divided into various functional modules. Here, "module" can refer to a specific ASIC, circuitry, a processor and memory executing one or more software or firmware programs, integrated logic circuitry, and / or other devices that can provide the aforementioned functions. In a simplified embodiment, those skilled in the art will understand that the resource playback device can employ... Figure 5 The terminal shown is in this form.

[0586] for example, Figure 5 The processor 510 in the terminal shown can cause the electronic device to execute the resource playback method in the above method embodiment by calling computer execution instructions stored in the internal memory 521.

[0587] As an example, combined Figure 6 The above Figure 19 The display module 1901, input module 1902, and playback module 1903 in the resource playback device shown, or the aforementioned Figure 20 The functions implemented by some or all of the display module 2001, receiving module 2002, and playback module 2003 in the resource playback device shown can be achieved through... Figure 6 Processor 601 in the middle executes Figure 6 The computer executes instructions in memory 602.

[0588] In this embodiment, the resource playback device is presented in an integrated manner, divided into various functional modules. Here, "module" can refer to a specific ASIC, circuitry, a processor and memory executing one or more software or firmware programs, integrated logic circuitry, and / or other devices that can provide the aforementioned functions. In a simplified embodiment, those skilled in the art will understand that the resource playback device can employ... Figure 6 The server configuration shown.

[0589] for example, Figure 6 The processor 601 in the server shown can cause the electronic device to execute the resource playback method in the above method embodiment by calling computer execution instructions stored in the memory 602.

[0590] Since the resource playback device provided in this application embodiment can execute the above-described resource playback method, the technical effects it can achieve can be referred to the above-described method embodiment, and will not be repeated here.

[0591] It should be understood that one or more of the above modules or units can be implemented by software, hardware, or a combination of both. When any of the above modules or units are implemented by software, the software exists as computer program instructions and is stored in memory. The processor can be used to execute the program instructions and implement the above method flow. The processor can be built into a SoC (System-on-a-Chip) or ASIC, or it can be a separate semiconductor chip. In addition to the core that executes software instructions for computation or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), PLDs (Programmable Logic Devices), or logic circuits that implement dedicated logic operations.

[0592] When the above modules or units are implemented in hardware, the hardware can be any one or any combination of a microprocessor, digital signal processing (DSP) chip, microcontroller unit (MCU), artificial intelligence processor, ASIC, SoC, FPGA, PLD, application-specific digital circuit, hardware accelerator, or non-integrated discrete device, which can run the necessary software or perform the above method flow independently of software.

[0593] Optionally, embodiments of this application also provide an electronic device (e.g., the electronic device may be a chip or a chip system), which includes a processor for implementing the methods executed by the electronic device in any of the above method embodiments. In one possible design, the electronic device further includes a memory. The memory is used to store necessary program instructions and data, and the processor can call the program code stored in the memory to instruct the electronic device to execute the methods in any of the above method embodiments. Of course, the memory may not be present in the electronic device. When the electronic device is a chip system, it may be composed of chips or may include chips and other discrete devices; embodiments of this application do not specifically limit this.

[0594] This application also provides a computer-readable storage medium storing computer-executable instructions that, when executed on an electronic device, cause the electronic device to perform the method executed by any of the resource playback devices provided above.

[0595] For explanations of the relevant content and descriptions of the beneficial effects in any of the computer-readable storage media provided above, please refer to the corresponding embodiments described above, which will not be repeated here.

[0596] This application also provides a chip. This chip integrates a control circuit for implementing the functions of the aforementioned resource playback device and one or more ports. Optionally, the functions supported by this chip can be referred to above, and will not be repeated here. Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, random access memory, etc. The aforementioned processing unit or processor can be a central processing unit, a general-purpose processor, an application-specific integrated circuit (ASIC), a microprocessor (digital signal processor, DSP), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0597] This application also provides a computer program product containing computer-executable instructions, which, when executed on an electronic device, cause the electronic device to perform any of the methods described in the above embodiments. The computer program product includes one or more computer-executable instructions. When these computer-executable instructions are loaded and executed on the electronic device, all or part of the flow or function according to the embodiments of this application is generated. The electronic device may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0598] Computer-executable instructions can be stored in or transmitted from one computer-readable storage medium to another. For example, computer-executable instructions can be transmitted from one website, computer, electronic device, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium accessible to an electronic device or a data storage device including one or more electronic devices, data centers, etc., that can be integrated with such media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0599] It should be noted that the devices for storing computer instructions or computer programs provided in the embodiments of this application, such as but not limited to the memory, computer-readable storage medium and communication chip, are all non-transitory.

[0600] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product.

[0601] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0602] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

Claims

1. A method for playing back a resource, characterized in that, The method includes: Display the first interface; Receive the first media resource and the second media resource input by the user on the first interface; In response to the input first media resource and second media resource, a target media resource is played. The target media resource is a media resource obtained by adding the second media resource to the first media resource based on candidate frames. The content semantics of the candidate frames match the content semantics of the second media resource.

2. The method according to claim 1, characterized in that, The target media resources for playback include: The second media resource is played in a picture-in-picture format within the first media resource.

3. The method according to claim 1 or 2, characterized in that, Before playing the target media resource, the method further includes: Obtain the content semantics of the first media resource and the content semantics of the second media resource; Based on the content semantics of the first media resource and the content semantics of the second media resource, the candidate frame is determined from the first media resource; wherein the candidate frame and the content semantics of the second media resource satisfy a similarity condition, or the content semantics of the candidate frame includes one or more of entity type, scene type, expression type or behavior type.

4. The method according to claim 3, characterized in that, The content semantics of the first media resource includes the content semantics of multiple shot clips, and the process of acquiring the content semantics of the multiple shot clips includes: The first media resource is analyzed to obtain multiple shot segments; the content of each of the multiple shot segments is understood to obtain the content semantics of the multiple shot segments. or, The content semantics of the first media resource includes the content semantics of the target shot segment, and the process of acquiring the content semantics of the target shot segment includes: Extract keywords from the text information of the second media resource; select target shot segments that include the keywords or synonyms from the multiple shot segments, where the similarity between the synonyms and the keywords reaches a similarity threshold; perform content understanding on the target shot segments to obtain the content semantics of the target shot segments.

5. The method according to claim 3 or 4, characterized in that, Determining the candidate frame from the first media resource includes: Based on the content semantics of the shot fragments in the first media resource and the content semantics of the second media resource, candidate shot fragments are determined from the first media resource, and the content similarity between the candidate shot fragments and the second media resource satisfies the similarity condition; From the candidate shot segments, the candidate frame is determined, wherein the candidate frame is the last frame of the candidate shot segment or the first frame of the next shot segment.

6. The method according to claim 3 or 4, characterized in that, Determining the candidate frame from the first media resource includes: Cluster the continuous shots in the first media resource to obtain at least one scene segment; Based on the content semantics of at least one scene segment in the first media resource and the content semantics of the second media resource, candidate scene segments are determined from the at least one scene segment, and the content similarity between the candidate scene segment and the second media resource satisfies the similarity condition; Based on the content semantics of the shot segments included in the candidate scene segments and the content semantics of the second media resource, candidate shot segments are determined from the shot segments included in the candidate scene segments, and the content similarity between the candidate shot segments and the second media resource satisfies the similarity condition; The candidate frame is determined from the candidate shot segments, wherein the candidate frame is the last frame of the candidate shot segment or the first frame of the next shot segment.

7. The method according to any one of claims 3-6, characterized in that, The first media resource is a video resource, and the content semantics of the first media resource includes the content semantics of audio and the content semantics of video. The second media resource is a video resource or an image resource; when the second media resource is a video resource, the content semantics of the second media resource includes the content semantics of audio and the content semantics of video; when the second media resource is an image resource, the content semantics of the second media resource includes the content semantics of image.

8. The method according to any one of claims 1-7, characterized in that, Before playing the target media resource, the method further includes: Display candidate frames of the first media resource; Receive user feedback on the candidate frame, the feedback being used to instruct the user to confirm the candidate frame or the video frame corrected by the user; The target media resource is the media resource obtained by adding the second media resource to the first media resource based on the target frame; wherein, when the feedback operation is used to instruct the user to confirm the candidate frame, the target frame is the candidate frame; when the feedback operation is used to instruct the user to correct the video frame, the target frame is the video frame corrected by the user.

9. The method according to claim 8, characterized in that, The number of candidate frames is multiple, and the candidate frames for displaying the first media resource include: Multiple candidate frames are displayed in descending order of their matching degree with the second media resource.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: The first media resource displays the linking text of the second media resource, and the linking text matches the content semantics of the first media resource and the content semantics of the second media resource.

11. The method according to claim 10, characterized in that, The method further includes: Based on the semantic content of the shot segments before and after the candidate frame, plot summary texts before and after the candidate frame are generated. Based on the plot summary text before the candidate frame, the plot summary text after the candidate frame, the content semantics of the candidate frame, and the content semantics of the second media resource, a linking word text for the second media resource is generated.

12. The method according to claim 10 or 11, characterized in that, Before playing the target media resource, the method further includes: Display the text of the connecting words in the second media resource; Receive user feedback on the linking word text, the feedback being used to instruct the user to confirm the linking word text or the user-corrected text; The step of displaying the linking text of the second media resource in the first media resource includes: Based on the feedback operation, the target text is displayed in the first media resource; Wherein, when the feedback operation is used to instruct the user to confirm the conjunction text, the target text is the conjunction text; when the feedback operation is used to instruct the user to correct the text, the target text is the text corrected by the user.

13. The method according to any one of claims 1-12, characterized in that, The target media resources for playback include: In the candidate resource area of ​​the first media resource, the second media resource is played according to the aspect ratio of the second media resource. The candidate resource area is the screen area in the first media resource used to display the second media resource.

14. The method according to claim 13, characterized in that, Before playing the target media resource, the method further includes: Display the candidate resource area for the second media resource; Receive user feedback on the candidate resource region, the feedback operation being used to instruct the user to confirm the candidate resource region or the region corrected by the user; Playing the second media resource in the candidate resource area of ​​the first media resource according to the aspect ratio scaling ratio of the second media resource includes: Based on the feedback operation, the second media resource is played in the target area of ​​the first media resource according to the aspect ratio of the second media resource; Wherein, when the feedback operation is used to instruct the user to confirm the candidate resource region, the target region is the candidate resource region; when the feedback operation is used to instruct the user to correct the region, the target region is the region corrected by the user.

15. The method according to any one of claims 1-14, characterized in that, The method further includes: Play the target audio from the second media resource in the first media resource.

16. The method according to claim 15, characterized in that, The method further includes: Extract the audio features of an audio segment of a preset duration preceding the candidate frame; input the audio content of the second media resource and the audio features of the audio segment into a speech synthesizer to obtain the target audio of the second media resource, wherein the audio features of the target audio are the same as the audio features of the audio segment; or, The audio features of the preset audio are extracted, and the audio content of the second media resource and the audio features of the preset audio are input into a speech synthesizer to obtain the target audio of the second media resource. The audio features of the target audio are the same as the audio features of the preset audio.

17. A method for playing back a resource, characterized in that, The method includes: The target film and television resource is identified, and the target film and television resource includes a first segment and a second segment. The first segment contains a video embedded in the frame. The video is used to showcase product features. The product features match the plot of the first segment or the second segment. The second segment is the segment before or after the first segment. Receive the user's playback command for the identifier; In response to the playback operation, the target video resource is played, wherein the first segment is played synchronously with the embedded video.

18. The method according to claim 17, characterized in that, The video is embedded in the background area of ​​the first segment.

19. The method according to claim 17 or 18, characterized in that, The method further includes: Display target text, which is used to describe the product features in conjunction with the plot of the first or second segment.

20. The method according to claim 19, characterized in that, The method further includes: Play the audio corresponding to the target text, wherein the audio features are the same as the audio features of the first segment or the second segment, or the audio features are the same as the audio features of a preset audio.

21. A resource playback device, characterized in that, The device includes: The display module is used to display the first interface. The input module is used to receive the first media resource and the second media resource input by the user on the first interface; A playback module is used to play a target media resource in response to the input first media resource and second media resource. The target media resource is a media resource obtained by adding the second media resource to the first media resource based on candidate frames. The content semantics of the candidate frames match the content semantics of the second media resource.

22. A resource playback device, characterized in that, The device includes: The display module is used to display the identifier of the target film and television resource, which includes a first segment and a second segment. The first segment contains a video embedded in its frame. The video is used to showcase product features. The product features match the plot of the first segment or the second segment. The second segment is the segment before or after the first segment. A receiving module is used to receive the user's playback operation on the identifier; A playback module is used to play the target video resource in response to the playback operation, wherein the first segment is played synchronously with the embedded video.

23. An electronic device, characterized in that, The device includes a memory and a processor, the memory and the processor being connected; the memory is used to store computer-executable instructions; the processor is used to invoke the computer-executable instructions to perform the method as described in any one of claims 1-16 or 17-20.

24. A computer-readable storage medium, characterized in that, Includes computer execution instructions, which, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-16 or 17-20.

25. A computer program product, characterized in that, Includes computer-executable instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-16 or 17-20.