Display device and media asset display method

By optimizing multi-dimensional similarity calculations and user behavior feedback, the problems of single-dimensional association and misassociation in multimedia content linkage have been solved, achieving accurate matching and personalized recommendations.

CN121985170APending Publication Date: 2026-05-05VIDAA (NETHERLANDS) INT HLDG LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIDAA (NETHERLANDS) INT HLDG LTD
Filing Date
2026-01-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing technologies, multimedia content linkage methods rely on manual annotation or single text tag matching, resulting in a single association dimension, easy semantic ambiguity and misassociation, and difficulty in efficiently managing a large-scale, dynamically growing content library.

Method used

By calculating the multi-dimensional similarity between the media asset being displayed and the media asset to be associated, including text, generation time, generation location, and scene similarity, a comprehensive similarity is calculated to determine the associated media asset, and the weight model is optimized by combining user behavior feedback.

Benefits of technology

It achieves accurate matching of multimedia content, reduces false associations, improves the transparency of content association and the effect of personalized recommendations, and conforms to user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985170A_ABST
    Figure CN121985170A_ABST
Patent Text Reader

Abstract

Some embodiments of the present application show a display device and a media asset display method, the method comprising: in response to an instruction for displaying a first media asset, displaying the first media asset in a first page, and obtaining first media asset information of the first media asset and second media asset information respectively corresponding to a plurality of second media assets; calculating a first similarity, a second similarity, a third similarity and a fourth similarity; calculating a comprehensive similarity according to the first similarity, the second similarity, the third similarity and the fourth similarity; determining the second media assets with the comprehensive similarity greater than a first threshold value as associated media assets; and displaying the associated media assets on the first page. According to the embodiment of the invention, the multi-dimensional similarity calculation is carried out by utilizing the media asset information of the media assets being displayed and the media assets to be associated, the comprehensive similarity is calculated based on the multi-dimensional similarity, the associated media assets are determined according to the comprehensive similarity, the ambiguity problem of single text matching is solved, and accurate matching is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and in particular to a display device and a media asset display method. Background Technology

[0002] With the significant improvement in hardware performance of smart TVs, set-top boxes, and home multimedia terminals, these devices are generally equipped with large-capacity storage, enabling users to store large amounts of high-resolution images and videos locally or in the cloud for extended periods. Simultaneously, the development of the Internet of Things (IoT) ecosystem allows media data from mobile devices to be seamlessly synchronized to televisions, further exacerbating the exponential growth of multimedia content in the home environment. Against this backdrop, there is a need for efficient and intelligent connections between these heterogeneous and massive amounts of visual content.

[0003] Multimedia content linkage methods mainly rely on manual annotation or matching based on single text tags. Manual annotation is costly and difficult to cover a large-scale, dynamically growing content library. Single text tag matching can be done by matching only name keywords, but it is still limited to the comparison of surface features, resulting in an overly simplistic association dimension and susceptibility to misassociations due to semantic ambiguity. Summary of the Invention

[0004] Some embodiments of this application provide a display device and a media asset display method, which calculates multi-dimensional similarity using media asset information of the currently displayed media asset and the media asset to be associated, and calculates a comprehensive similarity based on the multi-dimensional similarity, and determines the associated media asset based on the comprehensive similarity, thereby solving the ambiguity problem of single text matching and achieving accurate matching.

[0005] In a first aspect, some embodiments of this application provide a display device, including: monitor; The controller is configured as follows: In response to an instruction to display a first media asset, the first media asset is displayed on a first page, and the first media asset information and the second media asset information corresponding to multiple second media assets are obtained respectively; the first media asset information includes multiple first texts extracted from the name of the first media asset, the generation time of the first media asset, the generation location of the first media asset, and the scene of the first media asset, wherein the multiple first texts correspond to different semantic categories; the second media asset information includes second texts extracted from the name of the second media asset, the generation time of the second media asset, the generation location of the second media asset, and the scene of the second media asset, wherein the media asset types of the first media asset and the second media asset are different; Based on multiple first texts, the semantic categories corresponding to the multiple first texts, and the second texts, calculate the first similarity; calculate the second similarity between the generation time of the first media asset and the generation time of the second media asset; calculate the third similarity between the generation location of the first media asset and the generation location of the second media asset; and calculate the fourth similarity between the scene of the first media asset and the scene of the second media asset. Calculate the overall similarity between the second media asset and the first media asset based on the first similarity, second similarity, third similarity, and fourth similarity. Among multiple secondary media assets, those with a comprehensive similarity greater than the first threshold are identified as related media assets; Related media assets are displayed on the first page.

[0006] The above technical solution has the following advantages or beneficial effects: it uses the media asset information of the currently displayed media asset and the media asset to be associated to perform multi-dimensional similarity calculation, and calculates the comprehensive similarity based on the multi-dimensional similarity, and uses the comprehensive similarity to determine the associated media asset, thereby solving the ambiguity problem of single text matching and achieving accurate matching.

[0007] In some embodiments, the controller performs a first similarity calculation based on a plurality of first texts, semantic categories corresponding to the plurality of first texts respectively, and second texts, which is further configured to: Based on the correspondence between semantic categories and semantic importance levels, and the semantic category corresponding to each first text, the semantic importance level corresponding to each first text is obtained; Based on the correspondence between semantic importance level and text weight, and the semantic importance level of each first text, the text weight corresponding to each first text is obtained; Calculate the text similarity between each first text and the second text; Calculate the weighted text similarity of each first text and the second text based on the text similarity between each first text and the text weight corresponding to each first text. The weighted text similarity scores of each first text are summed to obtain the first similarity score.

[0008] The above technical solution has the following advantages or beneficial effects: by introducing a multi-level semantic text structure and combining it with a weighted fusion mechanism to calculate the comprehensive similarity with media asset text, it can capture multi-level semantic associations from local to global and from surface to deep, avoiding information omissions or deviations caused by single-granularity matching.

[0009] In some embodiments, the controller performs a comprehensive similarity calculation between the second media asset and the first media asset based on a first similarity, a second similarity, a third similarity, and a fourth similarity, which is further configured to: Based on the correspondence between media asset scenarios and similarity weights, the first similarity weight, second similarity weight, third similarity weight, and fourth similarity weight corresponding to the first media asset scenario are obtained. Calculate the first weighted similarity based on the first similarity and the first similarity weight; calculate the second weighted similarity based on the second similarity and the second similarity weight; calculate the third weighted similarity based on the third similarity and the third similarity weight; and calculate the fourth weighted similarity based on the fourth similarity and the fourth similarity weight. The comprehensive similarity between the second media asset and the first media asset is calculated based on the first weighted similarity, the second weighted similarity, the third weighted similarity, and the fourth weighted similarity.

[0010] The above technical solution has the following advantages or beneficial effects: different media asset scenarios have different focuses on similarity. By establishing a mapping relationship between media asset scenarios and similarity weights, the most suitable weight combination can be automatically selected for the current scenario, achieving more accurate media asset matching.

[0011] In some embodiments, the controller's execution of displaying associated media assets on the first page is further configured to: Rank the related media assets based on their overall similarity. The display position of related media assets is determined based on their ranking. The relevant media assets are displayed in the designated location on the first page.

[0012] The above technical solution has the following advantages or beneficial effects: it prioritizes media assets with high overall similarity, ensuring that the most relevant and matching content is presented first, which meets user expectations.

[0013] In some embodiments, after the controller displays the associated media assets on the first page, it is further configured to: In response to the user's instruction to open the associated media asset, the user is redirected from the first page to the second page, where the associated media asset is displayed, and a timer is started; In response to the user's instruction to close the associated media asset, the system redirects from the second page to the first page and displays the first media asset on the first page, and retrieves the timer's duration. The overall similarity of related media assets is adjusted based on the timing. The adjusted overall similarity of related media assets is ranked; Update the display position of related media assets based on their adjusted rankings; Update and display the relevant related media assets in the display position on the first page.

[0014] The above technical solution has the following advantages or beneficial effects: Recording the dwell time based on users' active opening / closing of associated media assets is a strong signal for measuring content relevance and attractiveness. Longer dwell times indicate user interest, which can improve the overall similarity of the media asset; shorter dwell times indicate potential irrelevance or low quality, which can reduce its overall similarity. Compared to relying solely on static semantic matching, this mechanism can capture users' true intentions, making recommendations closer to actual preferences.

[0015] In some embodiments, after the controller identifies a second media asset with a comprehensive similarity greater than a preset value as an associated media asset among a plurality of second media assets, it is further configured to: Obtain the maximum weighted similarity among the first, second, third, and fourth weighted similarities of the related media assets; Obtain the similarity category corresponding to the maximum weighted similarity, and obtain the matching text corresponding to the similarity category. The matching text is used to represent the basis for the association between the second media asset and the first media asset. When displaying associated media assets on the first page, the controller is also configured as follows: Display the matching text on the associated media asset.

[0016] The above technical solution has the following advantages or beneficial effects: by making the most critical matching criteria in the internal multidimensional similarity calculation explicit into user-readable text, the transparency and credibility of content association can be improved, and user participation and satisfaction can be enhanced.

[0017] In some embodiments, after the controller displays the associated media assets on the first page, it is further configured to: In response to the user's command to open the associated media asset, the system redirects from the first page to the second page, displays the associated media asset on the second page, and starts a timer; In response to the user's input command to close the associated media asset, the system redirects from the second page to the first page and displays the first media asset on the first page, and retrieves the timer's duration; The similarity weight of the target similarity category is updated based on the timing time. The target similarity category is the similarity category corresponding to the matching text on the associated media asset.

[0018] The above technical solution has the following advantages or beneficial effects: it introduces a user behavior feedback mechanism, continuously iterates the weight model based on user behavior feedback, continuously optimizes the association rules, makes the linkage results more in line with user habits, and realizes personalized linkage.

[0019] In some embodiments, the controller performs the extraction of multiple first texts from the name of the first media asset, which is further configured to: The name of the first media asset is segmented into multiple first texts using preset delimiters, preset character lengths, or semantic recognition models.

[0020] The above technical solutions have the following advantages or beneficial effects: The segmentation method using preset delimiters or preset character lengths is suitable for media asset names with standardized formats and clear structures, and has low computational overhead and fast response. The segmentation method using semantic recognition models can make the segmentation results more accurate.

[0021] In some embodiments, the controller performs the task of obtaining the semantic categories corresponding to the multiple first texts, which is further configured as follows: Multiple first texts are input into the semantic recognition model, so that the semantic recognition model outputs the semantic categories corresponding to the multiple first texts respectively.

[0022] The above technical solution has the following advantages or beneficial effects: it performs fine-grained text segmentation on the first media asset name and uses a speech recognition model to generate the corresponding semantic category, providing a structured and hierarchical semantic representation for subsequent media asset matching.

[0023] Secondly, some embodiments of this application provide a media asset display method, including: In response to an instruction to display a first media asset, the first media asset is displayed on a first page, and the first media asset information and the second media asset information corresponding to multiple second media assets are obtained respectively; the first media asset information includes multiple first texts extracted from the name of the first media asset, the generation time of the first media asset, the generation location of the first media asset, and the scene of the first media asset, wherein the multiple first texts correspond to different semantic categories; the second media asset information includes second texts extracted from the name of the second media asset, the generation time of the second media asset, the generation location of the second media asset, and the scene of the second media asset, wherein the media asset types of the first media asset and the second media asset are different; Based on multiple first texts, the semantic categories corresponding to the multiple first texts, and the second texts, calculate the first similarity; calculate the second similarity between the generation time of the first media asset and the generation time of the second media asset; calculate the third similarity between the generation location of the first media asset and the generation location of the second media asset; and calculate the fourth similarity between the scene of the first media asset and the scene of the second media asset. Calculate the overall similarity between the second media asset and the first media asset based on the first similarity, second similarity, third similarity, and fourth similarity. Among multiple secondary media assets, those with a comprehensive similarity greater than the first threshold are identified as related media assets; Related media assets are displayed on the first page.

[0024] The above technical solution has the following advantages or beneficial effects: it uses the media asset information of the currently displayed media asset and the media asset to be associated to perform multi-dimensional similarity calculation, and calculates the comprehensive similarity based on the multi-dimensional similarity, and uses the comprehensive similarity to determine the associated media asset, thereby solving the ambiguity problem of single text matching and achieving accurate matching.

[0025] This application embodiment can respond to an instruction to display a first media asset, display a first page, and obtain first media asset information of the first media asset and second media asset information corresponding to multiple second media assets respectively; calculate a first similarity, a second similarity, a third similarity, and a fourth similarity; calculate a comprehensive similarity based on the first similarity, the second similarity, the third similarity, and the fourth similarity; determine the second media asset with a comprehensive similarity greater than a first threshold as an associated media asset; and display the associated media asset on the first page. This application embodiment utilizes the media asset information of the currently displayed media asset and the media asset to be associated to perform multi-dimensional similarity calculation, and calculates a comprehensive similarity based on the multi-dimensional similarity, using the comprehensive similarity to determine the associated media asset, thus resolving the ambiguity problem of single text matching and achieving accurate matching. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application; Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application; Figure 3 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application; Figure 4 A flowchart illustrating a media asset display method provided in some embodiments of this application; Figure 5 A schematic diagram illustrating a first type of associated image display page provided for some embodiments of this application; Figure 6 A schematic diagram illustrating a second type of associated image display page provided for some embodiments of this application; Figure 7 A schematic diagram of an image display page provided for some embodiments of this application; Figure 8 A schematic diagram illustrating a third type of associated image display page provided in some embodiments of this application; Figure 9 A flowchart illustrating another media asset display method provided in some embodiments of this application; Figure 10 This is a timing diagram of a media asset display method provided for some embodiments of this application. Detailed Implementation

[0028] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0029] In this embodiment, display device 200 generally refers to a device with screen display and data processing capabilities. For example, display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0030] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, users can operate the display device 200 via touch operation, mobile terminal 300, and control device 100. For example, control device 100 can be a remote control, stylus, gamepad, etc.

[0031] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.

[0032] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. This allows the display device 200 to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0033] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support, including but not limited to network television, smart television and Internet Protocol television.

[0034] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of display device 200.

[0035] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface 280.

[0036] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0037] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0038] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, the communication device 220 may include a WiFi module. When the display device 200 supports Bluetooth connection communication, the communication device 220 may include a Bluetooth module.

[0039] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0040] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0041] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0042] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0043] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0044] In some embodiments, the user input interface 280 can be used to receive instructions from user input. For example, the user input interface 280 can receive text information entered by the user in the user interface. The user input interface 280 can also receive confirmation instructions from the user regarding controls in the user interface. The user input interface 280 can also receive voice instructions entered by the user.

[0045] In some embodiments, to enable user interaction, the display device 200 may run an operating system. An operating system is a computer program that manages and controls the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface; for example, the operating system can directly control the display device to provide a user interface, or it can provide a user interface by running applications. The operating system also allows users to interact with the display device 200.

[0046] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices.

[0047] like Figure 3 As shown, the display device system is divided into three layers, from top to bottom: the application layer, the middleware layer, and the hardware layer.

[0048] The application layer mainly includes commonly used applications on TV, as well as the application framework. The commonly used applications are mainly browser-based applications, such as HTML5 apps, and native apps.

[0049] In this embodiment, the application layer includes an interaction layer, which includes a linkage execution module. The linkage execution module is used to display the media asset page and control the jump to the media asset page.

[0050] An application framework is a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interface for these functions (toolbar, status bar, menu, dialog box).

[0051] Native apps can support online or offline access, push notifications, or access to local resources.

[0052] The middleware layer includes various television protocols, multimedia protocols, and system components. Middleware can use the basic services (functions) provided by system software to connect different parts of application systems or different applications on the network, achieving resource sharing and function sharing.

[0053] In this embodiment, the middleware layer includes a data layer, a feature processing layer, a matching layer, and an optimization layer. The data layer stores video and image files and metadata (time, location, tags, etc.). The feature processing layer includes a video feature extraction module and an image feature parsing module. The video feature extraction module extracts video information, and the image feature parsing module extracts image information. The matching layer includes a cross-modal matching module and a scene adaptation module. The cross-modal matching module calculates the overall similarity, and the scene adaptation module provides weights that are appropriate for the scene. The optimization layer includes a user feedback learning module. The user feedback learning module updates the weights based on user actions.

[0054] The hardware layer mainly includes the HAL interface, hardware, and drivers. The HAL interface is a unified interface for all TV chips, with the specific logic implemented by each chip. The drivers mainly include: audio drivers, display drivers, Bluetooth drivers, camera drivers, Wi-Fi drivers, USB drivers, HDMI drivers, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers.

[0055] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.

[0056] With the significant improvement in hardware performance of smart TVs, set-top boxes, and home multimedia terminals, these devices are generally equipped with large-capacity storage, enabling users to store large amounts of high-resolution images and videos locally or in the cloud for extended periods. Simultaneously, the development of the Internet of Things (IoT) ecosystem allows media data from mobile devices to be seamlessly synchronized to televisions, further exacerbating the exponential growth of multimedia content in the home environment. Against this backdrop, there is a need for efficient and intelligent connections between these heterogeneous and massive amounts of visual content.

[0057] Multimedia content linkage methods mainly rely on manual annotation or matching based on single text tags. Manual annotation is costly and difficult to cover a large-scale, dynamically growing content library. Single text tag matching can be done by matching only name keywords, but it is still limited to the comparison of surface features, resulting in an overly simplistic association dimension and susceptibility to misassociations due to semantic ambiguity. For example, "Sanya sunset" and "Sanya sunrise" have similar texts but different scenes.

[0058] To avoid mis-associations due to overly simplistic association dimensions, this application provides a display device 200. The structure and functions of each part of the display device 200 can be found in the above embodiments. Furthermore, based on the display device 200 shown in the above embodiments, this embodiment further improves some functions of the display device 200. For example… Figure 4 As shown, controller 250 is configured to perform the following steps: Step S401: In response to the instruction to display the first media asset, display the first media asset on the first page, and obtain the first media asset information of the first media asset and the second media asset information corresponding to the multiple second media assets respectively.

[0059] In some embodiments, in response to an instruction to display the first media asset, the state of the linkage switch can also be obtained. If the linkage switch is in the "on" state, the first media asset can be displayed on the first page, and the first media asset information and the second media asset information corresponding to each of the multiple second media assets can be obtained. If the linkage switch is in the "off" state, only the first media asset can be displayed on the first page, without needing to obtain the first media asset information and the second media asset information corresponding to each of the multiple second media assets. The state of the linkage switch is used to indicate whether the associated media asset needs to be displayed on the media asset page.

[0060] In some embodiments, one implementation of obtaining second media asset information corresponding to multiple second media assets may include: obtaining second media assets from the same terminal or the same storage path of the same terminal as the first media assets, and then obtaining the second media asset information of the second media assets.

[0061] In some embodiments, one implementation of obtaining second media asset information corresponding to multiple second media assets may include: obtaining the application corresponding to displaying the first media asset, then obtaining the second media asset from the server corresponding to the application, and then obtaining the second media asset information of the second media asset.

[0062] In some embodiments, one implementation of obtaining second media asset information corresponding to multiple second media assets may include: obtaining the user account corresponding to the first media asset, then obtaining the second media asset from the media asset corresponding to the user account on the server, and then obtaining the second media asset information of the second media asset.

[0063] The first media asset and the second media asset have different media asset types, including videos and images. For example, in response to an instruction to display a first video, a playback page for the first video is displayed, and video information for the first video and image information corresponding to each of the multiple first images are obtained. Alternatively, in response to an instruction to display a second image, a playback page for the second image is displayed, and image information for the second image and video information corresponding to each of the multiple second videos are obtained.

[0064] In some embodiments, the first media asset information includes multiple first texts extracted from the name of the first media asset, the generation time of the first media asset, the generation location of the first media asset, and the scene of the first media asset, with the multiple first texts corresponding to different semantic categories.

[0065] In some embodiments, one implementation of splitting the name of a first media asset into multiple first texts may include: splitting the name of the first media asset into multiple first texts using a preset delimiter. The preset delimiter may be "_", "-", or a space, etc.

[0066] For example, the first media asset name is "2023-10_Yunnan Lijiang_Yulong Snow Mountain Hiking.mp4". After being truncated with "_", multiple first texts are obtained as ["2023-10", "Yunnan Lijiang", "Yulong Snow Mountain Hiking"].

[0067] In some embodiments, another implementation of segmenting a first media asset name into multiple first texts may include: inputting the first media asset name into a semantic recognition model to segment the first media asset name into multiple first texts according to the text type using the semantic recognition model.

[0068] For example, the semantic recognition model can be Natural Language Processing (NLP). The NLP model can identify time words ("2023-10"), place names ("Yunnan Lijiang"), etc., and automatically remove meaningless suffixes, such as ".mp4" and "fragment X".

[0069] In some embodiments, another implementation of dividing the first media asset name into multiple first texts may include dividing the first media asset name into multiple texts according to a preset length. For example, starting from the first character of the media asset name, every ten characters are divided into one first text.

[0070] In some embodiments, one implementation of obtaining semantic categories corresponding to multiple first texts may include: inputting multiple first texts into a semantic recognition model, so as to output semantic categories corresponding to the multiple first texts through the semantic recognition model. The semantic categories include core, theme, scene, object, behavior, and attribute, etc.

[0071] For example, the first media asset is named "Yunnan Lijiang_Yulong Snow Mountain Hiking.mp4". The segmented first text can include Yunnan Lijiang, Yulong Snow Mountain, and hiking. Among them, the semantic category corresponding to "Yunnan Lijiang" can be the core theme, the semantic category corresponding to "Yulong Snow Mountain" can be the scene, and the semantic category corresponding to "hiking" can be the behavior.

[0072] In some embodiments, one implementation of obtaining the semantic categories corresponding to multiple first texts may include: inputting the name of a first media asset into a semantic recognition model, so as to output multiple first texts and the semantic categories corresponding to the multiple first texts through the semantic recognition model.

[0073] For example, if the first media asset is named "XX Earth 2_Liu X Qiang_Segment 3.mp4", and is input into the semantic recognition model, the resulting first text may include "XX Earth 2", "Liu X Qiang", and "Segment 3". The semantic category corresponding to "XX Earth 2" can be a core topic, the semantic category corresponding to "Liu X Qiang" can be an object, and the semantic category corresponding to "Segment 3" can be an attribute.

[0074] In some embodiments, the generation time of the media asset may include the shooting time, upload time, or download time. The generation location of the media asset may include the shooting location, upload location, or download location, and the location may be a Global Positioning System (GPS) location. The scene of the media asset may include landscapes, movies, and home scenes.

[0075] In some embodiments, one implementation of obtaining the generation time, generation location, and scene of the first media asset may include: obtaining first metadata of the first media asset data, and extracting the generation time, generation location, and scene of the first media asset from the first metadata. Metadata refers to key information describing the content, attributes, source, structure, and other relevant information of a media asset, used for identifying, managing, retrieving, processing, and understanding media asset files. Metadata may include Exchangeable Image File Format (EXIF) and user tags. User tags are used to characterize user identity.

[0076] For example, taking the first media asset as an image, we obtain the image's EXIF ​​information, parse the EXIF ​​information, and extract keywords, shooting time, shooting location, and image scene from the image's name.

[0077] In some embodiments, one implementation of obtaining the generation time, generation location, and scene of the first media asset may include: obtaining the name of the first media asset, and then extracting the generation time, generation location, and scene of the first media asset from the name of the first media asset.

[0078] For example, in the media asset name "2023-10_Yunnan Lijiang_Yulong Snow Mountain Scenery.mp4", we can extract the generation time of the media asset as 2023-10, the generation location of the media asset as Lijiang, Yunnan, and the scene of the media asset as scenery.

[0079] In some embodiments, in order to better organize video information and facilitate subsequent matching with image information, after obtaining multiple first texts, the generation time of the first media asset, the generation location of the first media asset, and the scene of the first media asset, the multiple first texts, the generation time of the first media asset, the generation location of the first media asset, and the scene of the first media asset can be combined to form a first media asset feature matrix F_v = [F_v1, F_v2, F_v3, T_v, G_v, Type_v], where F_v1-F_v3 are the three first texts of the video, T_v is the shooting time of the video, i.e., the timestamp, G_v is the shooting location of the video, i.e., latitude and longitude, and Type_v is the scene of the video.

[0080] In some embodiments, the second media asset information includes second text extracted from the name of the second media asset, the generation time of the second media asset, the generation location of the second media asset, and the scene of the second media asset. For example, the second text may be keywords from the media asset name / tag. The generation time, generation location, and scene of the second media asset are obtained in the same way as the generation time, generation location, and scene of the first media asset, and will not be described again here.

[0081] It should be noted that if the first or second media asset is an image, the media asset scene in the image can be extracted using a lightweight convolutional neural network (CNN) model. The dominant color tone in the image can also be extracted using this model, and this dominant color tone is used to assist in scene matching.

[0082] In some embodiments, to better organize image information and facilitate subsequent matching with video information, after obtaining the second text, the generation time of the second media asset, the generation location of the second media asset, and the scene of the second media asset, the second text, the generation time of the second media asset, the generation location of the second media asset, and the scene of the second media asset can be combined to form a second media asset feature matrix F_p = [P_text, T_p, G_p, C_p, S_p], where T_p is the image capture time, i.e., timestamp, G_p is the image capture location, i.e., latitude and longitude, C_p is the main color tone, and S_p is the image scene.

[0083] Step S402: Calculate the first similarity based on multiple first texts, the semantic categories corresponding to the multiple first texts, and the second text; calculate the second similarity between the generation time of the first media asset and the generation time of the second media asset; calculate the third similarity between the generation location of the first media asset and the generation location of the second media asset; and calculate the fourth similarity between the scene of the first media asset and the scene of the second media asset.

[0084] In some embodiments, an implementation method for calculating a first similarity based on multiple first texts, semantic categories corresponding to the multiple first texts, and second texts may include: obtaining the semantic importance level corresponding to each first text based on the correspondence between semantic categories and semantic importance levels, and the semantic category corresponding to each first text; then obtaining the text weight corresponding to each first text based on the correspondence between semantic importance levels and text weights, and the semantic importance level corresponding to each first text; then calculating the text similarity between each first text and the second text, and calculating the weighted text similarity of each first text based on the text similarity between each first text and the second text, and the text weight corresponding to each first text; finally, summing the weighted text similarities of each first text to obtain the first similarity (Sim_text).

[0085] For example, the importance level corresponding to the core topic is level one, the importance level corresponding to the scene / object is level two, and the importance level corresponding to the behavior / attribute is level three. The text weight corresponding to importance level one is w_text1, the text weight corresponding to importance level two is w_text2, and the text weight corresponding to importance level three is w_text3. The similarity between video text 1 (core topic) and image text is Sim_text1, the similarity between video text 2 (scene) and image text is Sim_text2, and the similarity between video text 3 (behavior) and image text is Sim_text3. The first similarity Sim_text = Sim_text1 × w_text1 + Sim_text2 × w_text2 + Sim_text3 × w_text3.

[0086] The improved cosine similarity method is used to calculate the text similarity between the first and second texts. Cosine similarity is one of the most commonly used methods for calculating text similarity. It represents two text segments as vectors and then measures their semantic or lexical similarity by calculating the cosine of the angle between these two vectors. The importance levels corresponding to different first texts can be the same or different.

[0087] In some embodiments, one implementation of calculating the second similarity between the generation time of the first media asset and the generation time of the second media asset may include: calculating the second similarity between the generation time of the first media asset and the generation time of the second media asset based on the time difference, namely, time similarity (Sim_time).

[0088] For example, taking a video as the first media asset and an image as the second media asset, the second similarity Sim_time = 1 - min(|T_v-T_p| / 24h, 1), where T_v is the video recording time and T_p is the image recording time. In this formula, when the time difference (|T_v-T_p|) < 1 hour, Sim_time is 0.96-1; when the time difference (|T_v-T_p|) > 24 hours, Sim_time is 0.

[0089] In some embodiments, one implementation of calculating the third similarity between the generation location of the first media asset and the generation location of the second media asset may include: calculating the third similarity between the generation location of the first media asset and the generation location of the second media asset based on latitude and longitude distance, namely, location similarity (Sim_gps).

[0090] For example, the third similarity Sim_gps = 1 - min(distance / 1km, 1), where distance is the latitude and longitude distance between the locations of the first and second media asset shooting locations. In this formula, when the latitude and longitude distance is <100m, Sim_gps is 0.9-1, and when the latitude and longitude distance is >1km, Sim_gps is 0.

[0091] In some embodiments, one implementation of calculating the fourth similarity between the scene of the first media asset and the scene of the second media asset may include: determining whether the scene of the first media asset and the scene of the second media asset are the same. If the scene of the first media asset and the scene of the second media asset are the same, then the fourth similarity, i.e., the visual similarity (Sim_visual), is determined to be 1; if the scene of the first media asset and the scene of the second media asset are different, then the fourth similarity is determined to be 0.

[0092] For example, if the first media asset scenario is a snow mountain, and the second media asset scenario is a snow mountain, the fourth similarity is 1; if the second media asset scenario is a beach, the fourth similarity is 0.

[0093] In some embodiments, another implementation of calculating the fourth similarity between the scene of the first media asset and the scene of the second media asset may include: inputting the scene of the first media asset and the scene of the second media asset into a scene matching model, so as to output the similarity between the scene of the first media asset and the scene of the second media asset through the scene matching model. If the scene of the first media asset or the scene of the second media asset is an image, the dominant color tone may also be input into the scene matching model.

[0094] For example, the first media asset is set in a seaside setting, and the second media asset is set in a sandy beach setting. The scene matching model outputs a fourth similarity score of 0.8.

[0095] Step S403: Calculate the comprehensive similarity between the second media asset and the first media asset based on the first similarity, the second similarity, the third similarity, and the fourth similarity.

[0096] In some embodiments, an implementation of calculating the comprehensive similarity between the second media asset and the first media asset based on the first similarity, the second similarity, the third similarity, and the fourth similarity may include: adding the first similarity, the second similarity, the third similarity, and the fourth similarity to obtain the comprehensive similarity between the second media asset and the first media asset.

[0097] In some embodiments, one implementation of calculating the comprehensive similarity between the second media asset and the first media asset based on a first similarity, a second similarity, a third similarity, and a fourth similarity may include: obtaining a first similarity weight corresponding to the first similarity, a second similarity weight corresponding to the second similarity, a third similarity weight corresponding to the third similarity, and a fourth similarity weight corresponding to the fourth similarity. Then, based on the first similarity and the first similarity weight, a first weighted similarity is calculated; based on the second similarity and the second similarity weight, a second weighted similarity is calculated; based on the third similarity and the third similarity weight, a third weighted similarity is calculated; and based on the fourth similarity and the fourth similarity weight, a fourth weighted similarity is calculated. Finally, based on the first weighted similarity, the second weighted similarity, the third weighted similarity, and the fourth weighted similarity, the comprehensive similarity (Total_Sim) between the second media asset and the first media asset is calculated.

[0098] For example, the total similarity Total_Sim = Sim_text × w1 + Sim_time × w2 + Sim_gps × w3 + Sim_visual × w4. Where Sim_text is the first similarity score, w1 is the weight of the first similarity score, Sim_time is the second similarity score, w2 is the weight of the second similarity score, Sim_gps is the third similarity score, w3 is the weight of the third similarity score, Sim_visual is the fourth similarity score, and w4 is the weight of the fourth similarity score.

[0099] In some embodiments, the first similarity weight, the second similarity weight, the third similarity weight, and the fourth similarity weight are fixed values.

[0100] In some embodiments, one implementation of obtaining the first similarity weight, the second similarity weight, the third similarity weight, and the fourth similarity weight may include: obtaining the first similarity weight, the second similarity weight, the third similarity weight, and the fourth similarity weight corresponding to the first media asset scenario based on the correspondence between media asset scenarios and similarity weights. Different media asset scenarios correspond to different similarity weights.

[0101] For example, when the media asset scenario is scenery, w1=a1, w2=a2, w3=a3, w4=a4; the overall similarity Total_Sim = a1×Sim_text + a2×Sim_time + a3×Sim_gps + a4×Sim_visual. When the media asset scenario is film / video, w1=b1, w2=b2, w3=b3, w4=b4, where b2 and b3 can be 0 to weaken the influence of time and location factors; the overall similarity Total_Sim = b1×Sim_text + b4×Sim_visual. When the media asset scenario is home video, w1=c1, w2=c2, w3=c3, w4=c4, where c4 can be 0 to weaken the influence of visual factors; the overall similarity Total_Sim = c1×Sim_text + c2×Sim_time + c3×Sim_gps.

[0102] It should be noted that the overall similarity of the second media assets can be saved in the cache so that they can be retrieved and modified.

[0103] Step S404: Identify the second media assets with a comprehensive similarity greater than the first threshold as related media assets.

[0104] For example, the first threshold can be 0.6. If the overall similarity is 0.7, the media asset is determined to be a related media asset. If the overall similarity is 0.4, the media asset is determined not to be a related media asset.

[0105] Step S405: Display related media assets on the first page.

[0106] Related media assets can be displayed in the upper layer of the first page, or they can be displayed at the edge of the first page, such as the top bar, bottom bar, or side bar.

[0107] If the associated media asset is an image, it can be displayed as an image thumbnail. If the associated media asset is a video, it can be displayed as a thumbnail of the first frame of the video or a specified frame image.

[0108] In some embodiments, one implementation of displaying related media assets on the first page may include: sorting the overall similarity of the related media assets, for example from largest to smallest, then determining the display position of the related media assets based on the ranking of the related media assets, and finally displaying the corresponding related media assets at the display position on the first page.

[0109] In some embodiments, determining the display position of associated media assets based on their order may include: obtaining the display position corresponding to the associated media assets based on the correspondence between ranking and display position.

[0110] For example, the first media asset is a video, and the second media asset is an image. The images associated with the currently playing video 1 are sorted by their overall similarity from highest to lowest as image 1, image 2, image 3, and image 4. These associated images are displayed in the sidebar of the video 1 playback page. The ranking and display position correspond one-to-one with the overall similarity ranking and the position number in the sidebar from top to bottom. The page displaying associated images can be as follows: Figure 5 As shown, Images 1, 2, 3, and 4 are displayed sequentially from top to bottom in the sidebar. If the number of associated images exceeds the number of sidebar positions, the last position in the sidebar can be displayed as a "More Controls" section, allowing users to select the control to reveal images not currently displayed in the sidebar.

[0111] In some embodiments, another implementation of displaying associated media assets on the first page may include: identifying associated media assets with a comprehensive similarity greater than a second threshold as core associated media assets, and identifying associated media assets other than the core associated media assets as secondary associated media assets. Then, the core associated media assets and more controls are displayed on the first page. After receiving confirmation from the user regarding the more controls, the secondary associated media assets are displayed on the first page. The second threshold is greater than the first threshold.

[0112] For example, the second threshold is 0.8. The first media asset is a video, and the second media asset is an image. The associated images of video 1 are image 1, image 2, image 3, and image 4. Among them, image 1 has a comprehensive similarity of 0.9 and is the core associated image; the comprehensive similarities of images 2, 3, and 4 are 0.77, 0.75, and 0.65, respectively, and are secondary associated images. Figure 6 As shown, the page displaying associated images includes Image 1 control 61 and more controls 62. Upon receiving user confirmation input for more control 62, the following can be displayed: Figure 5 The associated image display page is shown.

[0113] In some embodiments, where the first media asset is a video and the second media asset is an image, if a pause command is received while the video is playing, the video playback is paused, and a core associated photo is displayed in a floating layer above the video to cover the paused video screen. An animation effect can be presented, such as enlarging the core associated image in the sidebar and centering it on the screen.

[0114] For example, in Figure 6 If the system receives a user's command to pause playback, it can pause the video playback and display a message such as... Figure 7 The image display page shown. Key related images are enlarged and displayed on the page.

[0115] In some embodiments, where the first media asset is an image and the second media asset is a video, after determining the associated video or core associated video of the image, keyframes in the associated video or core associated video are obtained, and the similarity between the keyframes and the image is determined sequentially. The keyframe with the highest similarity is determined as the core keyframe. The timestamp of the core keyframe is obtained and saved.

[0116] When the first page displays images, the associated video control or core associated video control is displayed in the sidebar. Upon receiving user confirmation of the associated video control or core associated video control, the video playback page is displayed, and the playback progress is redirected to the time stamp corresponding to the core keyframe to continue playback.

[0117] This application's embodiments achieve an upgrade of multimedia content from shallow text association to deep semantics + scene adaptation through multi-level feature fusion and dynamic rule adaptation, significantly improving the accuracy of multimedia interaction and user experience.

[0118] In some embodiments, after displaying associated media assets on the first page, in response to a user's instruction to open associated media assets, the system navigates from the first page to a second page where the associated media assets are displayed, and a timer is started. In response to a user's instruction to close associated media assets, the system navigates from the second page to the first page where the first media assets are displayed, the timer is stopped, and the timer duration is obtained. Then, the overall similarity of the associated media assets is adjusted based on the time duration. The adjusted overall similarity of the associated media assets is sorted, and the display position of the associated media assets is updated according to their adjusted ranking. Finally, the corresponding associated media assets are updated and displayed at their designated positions on the first page.

[0119] It should be noted that navigating from the first page to the second page can either cancel the first page and display the second page, or the second page can be displayed above the first page. Similarly, navigating from the second page to the first page can either cancel the second page and display the first page, or the second page can be canceled so that the first page, which is located below the second page, is shown to the user. The first page displays thumbnails of the associated media assets, while the second page displays the standard size of the associated media assets. The standard size is larger than the thumbnail size.

[0120] One implementation of adjusting the overall similarity of associated media assets based on the timing period may include: if the timing period is greater than a preset duration, adding the overall similarity of the associated media assets to the preset similarity to obtain the adjusted overall similarity of the associated media assets. If the timing period is less than or equal to the preset duration, subtracting the overall similarity of the associated media assets from the preset similarity to obtain the adjusted overall similarity of the associated media assets. Different preset durations can be set for images and videos.

[0121] For example, with a preset similarity of 0.1, the combined similarities of images 1, 2, 3, and 4 are 0.86, 0.77, 0.75, and 0.65, respectively. Figure 5 Upon receiving user confirmation for image 1, the system displays the image 1 page and starts a timer. Upon receiving user input to close image 1, the system stops the timer and retrieves the elapsed time. If the elapsed time is less than the preset duration t, the overall similarity of image 1 is updated to 0.86 - 0.1 = 0.76, and the image is displayed as shown below. Figure 8 The displayed page shows the images arranged in four positions: Image 2, Image 1, Image 3, and Image 4. If the timeout exceeds the preset duration t, the overall similarity of Image 1 is updated to 0.86 + 0.1 = 0.96, and it still displays as shown. Figure 5 The page shown is the display page.

[0122] In some embodiments, after identifying a second media asset with a comprehensive similarity greater than a preset value as an associated media asset from among multiple second media assets, the maximum weighted similarity among the first, second, third, and fourth weighted similarities of the associated media asset can be obtained. Then, the similarity category corresponding to the maximum weighted similarity is obtained, and the matching text corresponding to the similarity category is obtained. The matching text is used to characterize the association between the second media asset and the first media asset. When the associated media asset is displayed on the first page, the matching text is displayed on the associated media asset.

[0123] The similarity categories include text, time, location, and visual similarity. There is a correspondence between the similarity category and the matching text. For example, the matching text for text is "text match," the matching text for time is "same time," the matching text for location is "same location," and the matching text for visual similarity is "visual match," etc.

[0124] For example, Image 1 has the highest weighted similarity, and the text "text matching" can be displayed on Image 1.

[0125] In other embodiments, after identifying second media assets with a comprehensive similarity greater than a preset value as associated media assets from among multiple second media assets, the weighted similarity among the first, second, third, and fourth weighted similarities of the associated media assets that is greater than the preset weighted similarity can be determined as the target weighted similarity. Then, the similarity category corresponding to the target weighted similarity is obtained, and the matching text corresponding to the similarity category is obtained. The matching text is used to characterize the association basis between the second media asset and the first media asset. When the associated media asset is displayed on the first page, the matching text is displayed on the associated media asset. In the embodiments of this application, the matching text can be one or more.

[0126] For example, if the first weighted similarity and the third weighted similarity of image 1 are both greater than the preset weighted similarity, the text "text matching + same location" can be displayed on image 1.

[0127] In some embodiments, after displaying associated media assets on the first page, in response to a user's instruction to open the associated media assets, the system navigates from the first page to a second page where the associated media assets are displayed, and a timer is started. In response to a user's instruction to close the associated media assets, the system navigates from the second page to the first page where the first media assets are displayed, and the timer duration is obtained. The similarity weight of the target similarity category is updated based on the timer duration, and the updated similarity weight is used to calculate the overall similarity subsequently. The target similarity category is the similarity category corresponding to the matching text on the associated media assets.

[0128] In some embodiments, one implementation of adjusting the similarity weight of the target similarity category based on the timing time may include: if the timing time is greater than a preset duration, adding the similarity weight of the target similarity category to a preset weight to obtain the adjusted similarity weight of the target similarity category. If the timing time is less than or equal to the preset duration, subtracting the similarity weight of the target similarity category from the preset weight to obtain the adjusted similarity weight of the target similarity category. The similarity weight corresponding to the first media asset scenario may be updated.

[0129] For example, the preset weight is x, and the similarity weights for text, time, location, and vision are d1, d2, d3, and d4, respectively. Figure 5 Upon receiving user confirmation of image 1, the system displays the image 1 page and starts a timer. Upon receiving user input to close image 1, the system stops the timer and retrieves the timeout value. If the timeout is less than the preset duration t, the text similarity weight is updated to d1-x. If the timeout is greater than the preset duration t, the text similarity weight is updated to d1+x.

[0130] In some embodiments, one implementation of adjusting the similarity weight of the target similarity category based on the timing period may include: incrementing the count of the target similarity category by 1 if the timing period is greater than a preset duration; decrementing the count of the target similarity category by 1 if the timing period is less than or equal to the preset duration; and periodically increasing or decreasing the similarity weight of the target similarity category based on the cumulative count of the target similarity category.

[0131] For example, the adjusted similarity weight W = w + [n / N] × x, where w is the original similarity weight, n is the number of times the statistics are performed, [] is the rounding function, N is a fixed value, and x is the weight adjustment range after each number of statistics reaches N.

[0132] In some embodiments, the flowchart of the media asset display method may be as follows: Figure 9 As shown. After triggering a linkage event (displaying media assets or turning on a linkage switch), the event type is obtained. If the event type is video playback, video metadata is obtained, video information is extracted from the video metadata, and a video feature matrix is ​​constructed based on the video information. If the event type is image browsing, image metadata is obtained, image information is parsed from the metadata, and an image feature matrix is ​​constructed based on the image information. Weights are determined according to the scenario, and then multi-dimensional similarity with multiple images (playing videos) or multiple videos (browsing images) is calculated. Then, a comprehensive similarity is calculated by weighting the weights and multi-dimensional similarities. If the comprehensive similarity is less than or equal to the first threshold 'a', the linkage fails, meaning there are no media assets linked with the currently playing video or browsed image. If the comprehensive similarity is greater than the first threshold, the images are sorted according to the comprehensive similarity. If the current event is video playback, associated images are displayed in sorted order; if the current scenario is image browsing, associated videos are displayed in sorted order. User interaction data can also be recorded to update the similarity weights for different categories.

[0133] In some embodiments, taking a video as the first media asset and an image as the second media asset as an example, the timing diagram of the media asset display method can be as follows: Figure 10 As shown. Upon receiving the instruction to display a video, the linkage execution module displays the video playback page and retrieves the video's metadata and metadata for multiple images from the data layer. It then extracts video and image information from the metadata, sending the video information to the video feature extraction module and the image information to the image feature parsing module. The video feature extraction module generates a video feature matrix and sends it to the cross-modal matching module. The image feature parsing module generates an image feature matrix and sends it to the cross-modal matching module. The cross-modal matching module calculates the first, second, third, and fourth similarities, and retrieves the first, second, third, and fourth similarity weights corresponding to the video scene from the scene adaptation module, then calculates the comprehensive similarity. Images with a comprehensive similarity greater than a first threshold are identified as associated images and sent to the linkage execution module. The linkage execution module displays thumbnails of the associated images in the sidebar of the video playback page.

[0134] Upon receiving a user's instruction to select an associated image, the image browsing page is displayed. Upon receiving a user's instruction to exit the associated image browsing, the video playback page is displayed. The duration of the displayed image browsing page is obtained and sent to the user feedback learning module. The user feedback learning module updates the similarity weight and sends the updated similarity weight to the scene adaptation module. The scene adaptation module updates the similarity weight.

[0135] This application's embodiments can extract text corresponding to multiple semantic categories by truncating video names, construct a feature matrix by combining video metadata (time, location), and perform multi-dimensional analysis of image metadata (name, tags, shooting time, GPS, visual features) to form an image feature matrix. Feature weights are dynamically adjusted based on scene type (movie / home video / landscape, etc.), and accurate matching is achieved through a weighted similarity algorithm. A user behavior feedback mechanism is introduced to continuously optimize association rules and achieve personalized interaction.

[0136] For example, if the video being played is “2023-10-01_Yunnan Lijiang_Yulong Snow Mountain_Hiking.mp4”, the extracted video information is: F_v1=Yunnan Lijiang, F_v2=Yulong Snow Mountain, F_v3=Hiking, Type_v=Scenery, T_v=2023-10-01 09:30, G_v=Yulong Snow Mountain coordinates.

[0137] Image A is: Yulong Snow Mountain Hiking.jpg. The extracted image information is: P_text=Yulong Snow Mountain Hiking, T_p=2023-10-01 10:15, G_p=Yulong Snow Mountain coordinates, S_p=“Snow Mountain”; Similarity calculation: Sim_text = 0.9 (full matching of three-level features), Sim_time = 0.9 (time difference of 45 minutes), Sim_gps = 1 (same location), Sim_visual = 1 (scene matching); Weight adaptation (landscape category): Total_Sim = 0.3×0.9 + 0.2×0.9 + 0.3×1 + 0.2×1 = 0.95 → core association.

[0138] Image B is named "Lijiang Old Town.jpg". The extracted image information is: P_text = Lijiang Old Town; T_p = 2023-09-28; G_p = coordinates of Lijiang Old Town.

[0139] Similarity calculation: Sim_text = 0.6 (matching only first-level features), Sim_time = 0, Sim_gps = 0.1 (distance 50km), Sim_visual = 0; Total_Sim = 0.3×0.6 + 0 + 0.3×0.1 + 0 = 0.21<0.6 → no association.

[0140] When playing the video, the core related image A is displayed first in the sidebar, labeled with "text + time + location + scene matching".

[0141] For example, image C is "Liu Xqiang character poster.jpg", and the image information is: P_text = Liu Xqiang character poster; S_p = film and television character.

[0142] The video is titled "《The Wandering Earth 2》_Liu Xiqiang_Moon Crisis Clip.mp4", and the video information is: Type_v = "Film", F_v1 = "The Wandering Earth 2", F_v2 = "Liu Xiqiang", F_v3 = "Moon Crisis".

[0143] Similarity calculation: Sim_text = 0.85 (secondary feature matching), Sim_visual = 1 (scene matching); Weight adaptation (movie / TV category): Total_Sim = 0.7 × 0.85 + 0.3 × 1 = 0.895 → core association.

[0144] When browsing image C, an entry for "Related Video: The Wandering Earth 2: Lunar Crisis Clip" is displayed. Clicking on it will take you to the corresponding video clip.

[0145] This application's embodiments resolve the ambiguity problem of single-text matching through multi-level semantic feature fusion and cross-modal metadata, significantly improving the depth and accuracy of association and reducing the false association rate by more than 60%. It also dynamically adjusts rules for different content types such as scenery, movies, and home videos, enhancing scene adaptability and catering to diverse user needs. This application's embodiments can continuously iterate the weight model based on user behavior feedback, making the linkage results more aligned with user habits. This application's embodiments can improve user understanding and trust in the linkage logic through hierarchical display and association basis annotation, increasing operational efficiency by 40%.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A display device, characterized in that, include: monitor; The controller is configured as follows: In response to an instruction to display a first media asset, the first media asset is displayed on a first page, and first media asset information and second media asset information corresponding to multiple second media assets are obtained; the first media asset information includes multiple first texts extracted from the name of the first media asset, the generation time of the first media asset, the generation location of the first media asset, and the scene of the first media asset, wherein the multiple first texts correspond to different semantic categories; the second media asset information includes second texts extracted from the name of the second media asset, the generation time of the second media asset, the generation location of the second media asset, and the scene of the second media asset, wherein the media asset types of the first media asset and the second media asset are different; Based on the multiple first texts, the semantic categories corresponding to the multiple first texts respectively, and the second text, calculate a first similarity; calculate a second similarity between the generation time of the first media asset and the generation time of the second media asset; calculate a third similarity between the generation location of the first media asset and the generation location of the second media asset; and calculate a fourth similarity between the scene of the first media asset and the scene of the second media asset. Based on the first similarity, the second similarity, the third similarity, and the fourth similarity, calculate the comprehensive similarity between the second media asset and the first media asset; The second media assets whose overall similarity is greater than a first threshold among the plurality of second media assets are identified as associated media assets; The associated media assets are displayed on the first page.

2. The display device according to claim 1, characterized in that, The controller performs a first similarity calculation based on multiple first texts, the semantic categories corresponding to the multiple first texts respectively, and the second text, and is further configured to: Based on the correspondence between semantic categories and semantic importance levels, and the semantic category corresponding to each first text, the semantic importance level corresponding to each first text is obtained; Based on the correspondence between semantic importance level and text weight, and the semantic importance level of each first text, the text weight corresponding to each first text is obtained; Calculate the text similarity between each of the first text and the second text; Based on the text similarity between each first text and the second text, and the text weight corresponding to each first text, calculate the weighted text similarity of each first text; The weighted text similarity scores of each of the first texts are summed to obtain the first similarity score.

3. The display device according to claim 1, characterized in that, The controller is further configured to calculate the comprehensive similarity between the second media asset and the first media asset based on the first similarity, the second similarity, the third similarity, and the fourth similarity, as follows: Based on the correspondence between media asset scenarios and similarity weights, the first similarity weight, second similarity weight, third similarity weight and fourth similarity weight corresponding to the first media asset scenario are obtained; Calculate the first weighted similarity based on the first similarity and the first similarity weight; A second weighted similarity is calculated based on the second similarity and the second similarity weight; a third weighted similarity is calculated based on the third similarity and the third similarity weight; and a fourth weighted similarity is calculated based on the fourth similarity and the fourth similarity weight. The comprehensive similarity between the second media asset and the first media asset is calculated based on the first weighted similarity, the second weighted similarity, the third weighted similarity, and the fourth weighted similarity.

4. The display device according to claim 1, characterized in that, The controller, which executes the display of the associated media assets on the first page, is further configured to: Sort the overall similarity of the associated media assets; The display position of the associated media assets is determined based on their ranking. The corresponding associated media assets are displayed at the specified display location on the first page.

5. The display device according to claim 4, characterized in that, After the controller displays the associated media assets on the first page, it is further configured to: In response to the user's instruction to open the associated media asset, the user is redirected from the first page to the second page and the associated media asset is displayed on the second page, and a timer is started; In response to a user's instruction to close the associated media asset, the user is redirected from the second page to the first page and the first media asset is displayed on the first page, and the timer's duration is obtained; The overall similarity of the associated media assets is adjusted based on the timing period. The adjusted overall similarity of the associated media assets is then sorted. Update the display position of the associated media assets based on their adjusted ranking; The corresponding associated media assets are updated and displayed at the specified display location on the first page.

6. The display device according to claim 3, characterized in that, After the controller identifies the second media asset with a comprehensive similarity greater than a preset value as the associated media asset among the plurality of second media assets, it is further configured to: Obtain the maximum weighted similarity among the first weighted similarity, the second weighted similarity, the third weighted similarity, and the fourth weighted similarity of the associated media assets; Obtain the similarity category corresponding to the maximum weighted similarity, and obtain the matching text corresponding to the similarity category. The matching text is used to characterize the association basis between the second media asset and the first media asset. When the associated media assets are displayed on the first page, the controller is also configured to: Display the matched text on the associated media asset.

7. The display device according to claim 6, characterized in that, After the controller displays the associated media assets on the first page, it is further configured to: In response to a user's instruction to open the associated media asset, the system redirects from the first page to the second page, displays the associated media asset on the second page, and starts a timer. In response to a user's input command to close the associated media asset, the user is redirected from the second page to the first page and the first media asset is displayed on the first page, and the timer's duration is obtained; The similarity weight of the target similarity category is updated according to the time interval, where the target similarity category is the similarity category corresponding to the matching text on the associated media asset.

8. The display device according to claim 1, characterized in that, The controller is further configured to extract multiple first texts from the name of the first media asset. The name of the first media asset is segmented into multiple first texts using a preset delimiter, preset character length, or semantic recognition model.

9. The display device according to claim 8, characterized in that, The controller executes the process of obtaining the semantic categories corresponding to the multiple first texts, and is further configured as follows: The multiple first texts are input into a semantic recognition model, and the semantic recognition model outputs the semantic categories corresponding to the multiple first texts respectively.

10. A method for displaying media assets, characterized in that, include: In response to an instruction to display a first media asset, the first media asset is displayed on a first page, and first media asset information of the first media asset and second media asset information corresponding to multiple second media assets are obtained respectively; the first media asset information includes multiple first texts extracted from the name of the first media asset, the generation time of the first media asset, the generation location of the first media asset, and the scene of the first media asset, wherein the multiple first texts correspond to different semantic categories; the second media asset information includes second texts extracted from the name of the second media asset, the generation time of the second media asset, the generation location of the second media asset, and the scene of the second media asset, wherein the media asset types of the first media asset and the second media asset are different; Based on the multiple first texts, the semantic categories corresponding to the multiple first texts respectively, and the second text, calculate a first similarity; calculate a second similarity between the generation time of the first media asset and the generation time of the second media asset; calculate a third similarity between the generation location of the first media asset and the generation location of the second media asset; and calculate a fourth similarity between the scene of the first media asset and the scene of the second media asset. Based on the first similarity, the second similarity, the third similarity, and the fourth similarity, calculate the comprehensive similarity between the second media asset and the first media asset; The second media assets whose overall similarity is greater than a first threshold among the plurality of second media assets are identified as associated media assets; The associated media assets are displayed on the first page.