Content display method and apparatus, computer device, storage medium, and computer program product
By displaying the video feed of the interactive object in real time and identifying events of interest in the environment during video interaction, and automatically displaying related content, the problem of content sharing affecting continuity and efficiency in video interaction is solved, thus achieving efficient content display.
Patent Information
- Application Number
- PCT/CN2025/103899
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-15
- Filing Date
- 2025-06-26
- Publication Date
- 2026-02-19
AI Technical Summary
During video interaction, content sharing affects the continuity of video interaction and display efficiency. How can we display content more efficiently while ensuring the continuity of interaction?
In video interaction scenarios, the video feed of the interactive object is displayed in real time, and the selected content is identified through environmental content attention events. The content screen associated with the selected content is automatically displayed, avoiding additional sharing operations.
It improves the efficiency of content display, ensures the continuity of video interaction, and reduces interruptions caused by sharing operations.
Smart Images

Figure CN2025103899_19022026_PF_FP_ABST
Abstract
Description
Method, device, computer device, storage medium and computer program product for content display
[0001] Related applications
[0002] The present application claims priority to the Chinese patent application No. 202411135820.1, filed on August 15, 2024, and entitled "Method, device, computer device, storage medium and computer program product for content display", the contents of which are hereby incorporated by reference in its entirety. TECHNICAL FIELD
[0003] The present application relates to the technical field of Internet, and in particular to a method, device, computer device, storage medium and computer program product for content display. BACKGROUND
[0004] With the continuous development of technology and Internet, more and more application programs appear, and different application programs can interact with corresponding content information. For example, the video, news and text content in application program 1 are shared to application program 2. In the video interaction process, if the video interaction party 1 sees the content of interest and wants to share it with the video interaction party 2, the video interaction party 1 needs to share based on the sharing function of the application to which the content of interest belongs, and the video interaction party 2 needs to open the corresponding application to display and watch the content. That is, when video interaction is performed, content sharing will affect the continuity of video interaction between the two parties and the display efficiency of content display. Therefore, how to more efficiently display content and ensure the continuity of video interaction in the case of video interaction is a problem to be solved. SUMMARY
[0005] Therefore, it is necessary to provide a method, device, computer device, storage medium and computer program product for content display to solve the above technical problems.
[0006] In a first aspect, the present application provides a method for content display, which is executed by a computer device, and the method comprises:
[0007] In a video interaction scenario involving at least two interaction objects, a video screen captured in real time by a login terminal of at least one of the remaining interaction objects is displayed in a video interaction page of one of the interaction objects;
[0008] In response to an environment content attention event triggered by an environment content in any one of the video screens, the environment content matched with the environment content attention event is selected as selected content; and
[0009] In the video interaction page, a content screen associated with the selected content is displayed.
[0010] In a second aspect, the present application provides a content display device. The device comprises:
[0011] a video picture display module configured to display, in a video interaction page of one of the at least two interactive objects, a video picture collected in real time by a login terminal of the other at least one interactive object in a video interaction scenario in which the at least two interactive objects participate;
[0012] an event triggering module configured to, in response to an environmental content attention event triggered by environmental content in any one of the video pictures, select environmental content matched with the environmental content attention event as selected content; and
[0013] a content picture display module configured to display, in the video interaction page, a content picture associated with the selected content.
[0014] In a third aspect, the present application provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0015] display, in a video interaction page of one of the at least two interactive objects, a video picture collected in real time by a login terminal of the other at least one interactive object in a video interaction scenario in which the at least two interactive objects participate;
[0016] in response to an environmental content attention event triggered by environmental content in any one of the video pictures, select environmental content matched with the environmental content attention event as selected content;
[0017] display, in the video interaction page, a content picture associated with the selected content.
[0018] In a fourth aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0019] display, in a video interaction page of one of the at least two interactive objects, a video picture collected in real time by a login terminal of the other at least one interactive object in a video interaction scenario in which the at least two interactive objects participate;
[0020] in response to an environmental content attention event triggered by environmental content in any one of the video pictures, select environmental content matched with the environmental content attention event as selected content;
[0021] display, in the video interaction page, a content picture associated with the selected content.
[0022] In a fifth aspect, the present application provides a computer program product. The computer program product comprises a computer program which, when executed by a processor, implements the following steps:
[0023] In the video interaction scenario involving at least two interactive objects, a video screen captured in real time by a login terminal of at least one interactive object is displayed in a video interaction page of another interactive object;
[0024] In response to an environmental content focus event triggered by environmental content in any one video screen, the environmental content matched with the environmental content focus event is selected as selected content;
[0025] In the video interaction page, a content screen associated with the selected content is displayed.
[0026] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of the disclosed drawings.
[0028] FIG. 1 is a diagram of an application environment of a content display method in an embodiment;
[0029] FIG. 2 is a logic diagram of a technical implementation of a content display method in an embodiment;
[0030] FIG. 3 is a flow diagram of a content display method in an embodiment;
[0031] FIG. 4 is an interface diagram of a screen display mode of a login terminal of a first interactive object in an embodiment;
[0032] FIG. 5 is an interface diagram of a video screen captured in real time by a login terminal of an interactive object in an embodiment;
[0033] FIG. 6 is an embodiment diagram of an environmental content focus operation in an embodiment;
[0034] FIG. 7 is an embodiment diagram of the presence of environmental content and the absence of environmental content in a target video screen in an embodiment;
[0035] FIG. 8 is a diagram of an interactive operation triggered by environmental content in an embodiment;
[0036] FIG. 9 is a schematic diagram of an interaction operation of gesture zooming for environmental content in one embodiment;
[0037] FIG. 10 is a schematic diagram of an interaction operation for detecting a control in one embodiment;
[0038] FIG. 11 is a schematic diagram of displaying a content page associated with selected content in one embodiment;
[0039] FIG. 12 is a schematic diagram of an interface of a content page associated with selected content in one embodiment;
[0040] FIG. 13 is a schematic diagram of an interface of a content page associated with selected content in another embodiment;
[0041] FIG. 14 is a schematic diagram of an interface of a content page associated with selected content in yet another embodiment;
[0042] FIG. 15 is a schematic diagram of an interface of displaying selected content in a preset angle in one embodiment;
[0043] FIG. 16 is a schematic diagram of an interface of displaying a global image of selected content in another embodiment;
[0044] FIG. 17 is a schematic diagram of an interface of displaying a content promotion page of selected content in one embodiment;
[0045] FIG. 18 is a schematic diagram of a flow of determining selected content in one embodiment;
[0046] FIG. 19 is a schematic diagram of a flow of displaying a content page associated with selected content in a video interaction page in one embodiment;
[0047] FIG. 20 is a schematic diagram of an interface of determining multimedia content as selected content, displaying a video page, and playing the multimedia content in one embodiment;
[0048] FIG. 21 is a schematic diagram of an interface of playing multimedia content in a program page in one embodiment;
[0049] FIG. 22 is a schematic diagram of an interface of invoking a program page of a target application in a video interaction page in one embodiment;
[0050] FIG. 23 is a schematic diagram of an interface of playing multimedia content in a target webpage in one embodiment;
[0051] FIG. 24 is a schematic diagram of an interface of a video interaction page displaying download prompt information for a target application in one embodiment;
[0052] FIG. 25 is a schematic diagram of an interface of displaying play confirmation information of playing multimedia content through a target application in one embodiment;
[0053] FIG. 26 is a schematic diagram of an interface showing device identification of a candidate delivery device of multimedia content in an embodiment;
[0054] FIG. 27 is a schematic diagram of an interface of jumping to a target application to play multimedia content in an embodiment;
[0055] FIG. 28 is a schematic diagram of a complete flow of a content presentation method in an embodiment;
[0056] FIG. 29 is a structural block diagram of a content presentation device in an embodiment;
[0057] FIG. 30 is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0059] With the continuous development of technology and the Internet, more and more application programs appear, and different application programs can interact with corresponding content information. For example, the video, news, and text content in application program 1 is shared to application program 2. In the video interaction process, if video interaction party 1 sees the content of interest and wants to share it with video interaction party 2, video interaction party 1 needs to share based on the sharing function of the application to which the content of interest belongs, and video interaction party 2 needs to open the corresponding application to present and view the content, that is, content sharing during video interaction will affect the continuity of video interaction between the two parties and will also affect the presentation efficiency of content presentation. Therefore, how to more efficiently present content and ensure the continuity of video interaction in the case of video interaction is a problem to be solved.
[0060] The embodiments of the present application provide a content presentation method capable of improving the efficiency of content presentation and ensuring the continuity of video interaction in the case of video interaction. The content presentation method provided by the embodiments of the present application can be applied to the application environment as shown in FIG. 1. Wherein, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required to be processed by the server 104. The data storage system can be integrated on the server 104, or placed on the cloud or other servers.
[0061] Specifically, taking the terminal 102 as an example, in a video interaction scene in which at least two interaction objects participate, in a video interaction page of one of the interaction objects, a video screen collected in real time by a login terminal of the remaining at least one interaction object is displayed, in response to an environmental content attention event triggered for environmental content in any one video screen, environmental content matched with the environmental content attention event is selected content, and a content screen associated with the selected content is displayed in the video interaction page. Therefore, without the interaction object performing a corresponding content sharing operation, the process of environmental content triggering and content display is used to automatically display the selected content, so that the continuity of the video interaction can be avoided in the video interaction process, and the selected content can be displayed without multiple sharing interactions, thereby improving the efficiency of content display.
[0062] The terminal 102 can be, but is not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, and the like. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, and the like. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers. The content display method provided by the application embodiments can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, and the like.
[0063] Secondly, the technical terms involved in the present application are introduced as follows:
[0064] I. Camera Application Programming Interface (API).
[0065] The camera API generally refers to a set of programming interfaces that allow developers to control and manage camera hardware in software applications. Through the camera API, functions such as taking photos, recording videos, adjusting camera settings (such as focal length, exposure, white balance, etc.), and accessing real-time data streams of the camera can be achieved.
[0066] II. Convolutional Neural Networks (CNNs).
[0067] CNNs is a deep learning architecture commonly used to process data with grid structures, such as images (two-dimensional grid) and audio (one-dimensional grid). CNNs have a wide range of applications in image recognition, video analysis, image classification, natural language processing, and other fields. The core concept of CNNs is the "convolutional layer", and CNNs extract features from input data through convolution operations.
[0068] Three, image recognition model.
[0069] An image recognition model is an algorithmic model that uses machine learning, especially deep learning techniques, to recognize and process information in images. These models can identify objects, faces, scenes, and even the style and emotion of an image. Image recognition models have a wide range of applications in medical image analysis, autonomous driving, security monitoring, social networks, and many other fields.
[0070] Four, Approximate Nearest Neighbor Search (ANN).
[0071] ANN is an algorithm that quickly finds the nearest point to a query point in a large-scale dataset. Compared with Exact Nearest Neighbor Search, ANN can guarantee higher search efficiency. In many practical applications such as image retrieval, recommendation systems, machine learning, etc., the data volume is usually very large and the dimension is also high. In these cases, the exact nearest neighbor search will be very time-consuming, therefore, ANN algorithm can greatly improve the search speed, making it possible to apply in real-time systems.
[0072] Five, Uniform Resource Identifier (URI).
[0073] The URI scheme is a standard for identifying strings of resources on the Internet. The purpose of URI is to provide a simple and flexible way to identify and access resources on the network. The URI scheme is an important part of the URI, which defines how to specify the type of resource and access method.
[0074] The technical implementation logic of the present application can be summarized as follows: user side, product side and background side, as shown in the technical implementation logic diagram of FIG. 2, for the user side, in the video interaction scene participated by at least two interaction objects, and the interaction objects participating in the video interaction at least include the first interaction object and the second interaction object, in the login terminal of the first interaction object, the video picture collected by the login terminal of the second interaction object is displayed, similarly, in the login terminal of the second interaction object, the video picture collected by the login terminal of the first interaction object is also displayed. The present application is introduced based on the login terminal of the first interaction object, the login terminal of the first interaction object can filter target video frames from the video stream data collected by the login terminal of the second interaction object, analyze the target video frames for application program matching, and then display the content picture associated with the selected content on the basis of the display video picture, at this time, the content picture is related to the target application program for displaying the selected content.
[0075] And in the product side, if it is needed to display the video screen collected in real time by the login terminal of the second interactive object in the login terminal of the first interactive object, at this time, considering that the selected content is the multimedia content played by the terminal device, at this time, the second interactive object can display the multimedia content played by the terminal device to the camera, so that the multimedia content played by the terminal device exists in the video screen collected in real time by the login terminal of the second interactive object, and the video screen collected in real time by the login terminal of the first interactive object includes the image of the first interactive object. In the case of the object visual line of the first interactive object, the visual line focus detection is performed to obtain the visual line focus result. When the visual line focus result represents that the object visual line of the first interactive object is focused on the environmental content in the video screen, it is judged whether the duration of focusing on the environmental content reaches the preset focusing duration. If not, the video screen collected in real time by the login terminal of the second interactive object is continuously displayed in the login terminal of the first interactive object, and no other content screen is displayed. If yes, the duration of focusing on the environmental content reaches the preset focusing duration, at this time, the application program corresponding to the selected content is analyzed, at this time, the product side needs to perform application program retrieval, and the application program corresponding to the selected content is also analyzed and matched with the application program on the background side, which will be described in detail in subsequent embodiments. If there is no application program corresponding to the selected content, that is, the user side will not display the application program. Secondly, considering that in actual application, there can also be a case that the selected content is a real object content, at this time, the user side can also be caused to display the video screen and the selected content displayed at a preset angle, or the global image of the complete display of the real object content, or the content promotion page for recommending the real object content, or the content detail page for introducing the real object content.
[0076] Based on this, if there is an application program corresponding to the selected content, that is, the application program identification matching is needed, in the case that the target application program exists in the login terminal of the first interactive object, the program page of the target application program is invoked, that is, at this time, the target application program can be opened and navigated to the selected content, so that the user side displays the video screen and the selected content. On the contrary, in the case that the target application program does not exist in the login terminal of the first interactive object, the selected content can be displayed through a webpage, that is, the user side displays the video screen and the selected content displayed based on the webpage. Or, the download prompt information of the target application program can also be displayed, if the user needs to download the target application program, at this time, the program page of the target application program can be displayed in the case that the target application program is successfully downloaded in response to the program download confirmation operation triggered for the download prompt information.
[0077] The background side is described in detail below. The background will detect object operations and video stream data in real time and filter target video frames, thereby performing content detection on the target video frames. The specific manner of triggering application retrieval through content detection results is as follows: when the content detection result indicates that specific multimedia content exists in the target video frame, first, candidate applications that can possibly support display of the multimedia content of the type (such as video, text, graphics, etc.) are filtered from the pre-constructed application database according to the type of the multimedia content. Then, the content features of the multimedia content are extracted using an image recognition model, and the interface features corresponding to each candidate application are extracted from the application feature library. Next, the content features and the interface features are represented as feature vectors, and efficient search algorithms (such as approximate nearest neighbor search) and optimized data structures (such as KD tree, ball tree, etc.) are used to calculate the distance (such as Euclidean distance, cosine distance, etc.) between the content feature vector and each interface feature vector. The smaller the distance, the higher the similarity. Hashing techniques (such as local sensitive hashing (LSH)) can also be used to further improve the query and filtering speed. Finally, the application corresponding to the interface feature with the smallest distance (highest similarity) is selected as the target application to which the selected content belongs. The candidate application refers to an application that can possibly support display of the multimedia content of the type, which is filtered from the pre-constructed application database according to the type of the multimedia content. For example, when the type of the multimedia content is text, applications that support text display are filtered from the application database as candidate applications; when the type of the multimedia content is video, applications that support video playback are filtered as candidate applications.
[0078] Therefore, in the case where the target application exists in the login terminal of the first interactive object, the selected content is directly navigated to, that is, the target application is opened in the product side and the selected content is navigated to, at this time, the video screen and the selected content are displayed on the user side. Similarly, in the case where the target application does not exist in the login terminal of the first interactive object, the selected content is synchronized in the webpage, that is, the webpage is opened in the product side to display the selected content, at this time, the video screen and the selected content displayed based on the webpage are displayed on the user side. In actual applications, the first interactive object can also be prompted to download the target application, and if the first interactive object needs to download the target application, the program page of the target application can be displayed in the case where the target application is successfully downloaded in response to the program download confirmation operation triggered by the first interactive object for the download prompt information.
[0079] Specifically, the following embodiments are provided for illustration: in one embodiment, as shown in FIG. 3, a method for content display is provided, which is applied to the terminal 102 in FIG. 1 for example, and it can be understood that the method can also be applied to the system including the terminal 102 and the server 104, and is realized through the interaction of the terminal 102 and the server 104. In this embodiment, the method includes the following steps:
[0080] Step 302, in the video interaction scene in which at least two interaction objects participate, a video screen collected in real time by a login terminal of the rest of the at least one interaction object is displayed in a video interaction page of one of the interaction objects.
[0081] Among them, the interaction object is an object in the video interaction scene, and there are at least two interaction objects in the video interaction scene, that is, the interaction objects at least include a first interaction object and a second interaction object, and can also include a third interaction object and a fourth interaction object and more objects, which are not specifically limited. Based on this, the interaction objects all have login terminals, that is, there is a login terminal of the first interaction object for the first interaction object, and there is a login terminal of the second interaction object for the second interaction object.
[0082] Since it is in the video interaction scene, the login terminal of the object participating in the video interaction will display a real-time video collection video screen for other interaction objects participating in the video interaction. The video screen is obtained by analyzing real-time video stream data, and the video stream data specifically includes video frames collected by an image collection device, and can also include audio data collected by an audio collection device. The image collection device can be built-in in the login terminal of the interaction object, and can be a camera or the like that can collect images and is externally connected to the login terminal of the interaction object. In this application, a camera API is used to control and manage the image collection device to realize the functions of taking pictures, recording videos, and adjusting the settings of the image collection device, and the real-time data stream of the image collection device is accessed to analyze the video frames. Similarly, the audio collection device can also be built-in in the login terminal of the interaction object, and can be a microphone or the like that can collect audio and is externally connected to the login terminal of the interaction object. Herein, no specific limitation is made.
[0083] Specifically, in the video interaction scenario in which at least two interaction objects participate, in the video interaction page of one of the interaction objects, a video picture collected in real time by a login terminal of the rest of the at least one interaction object is displayed. For example, in the case where the interaction objects include a first interaction object and a second interaction object, that is, the objects participating in the video interaction include at least the first interaction object and the second interaction object, in the login terminal of the first interaction object, a video picture collected in real time by a login terminal of the second interaction object can be displayed. The video picture collected in real time by the login terminal of the second interaction object can include the second interaction object and environmental content in a picture frame collected in real time by the login terminal of the second interaction object. The environmental content can be a video object, a text object, a graphic-text object, or the like, or can be a real object, and the video picture can be a picture without any object.
[0084] In addition, in the login terminal of the first interaction object, a video picture collected in real time by a login terminal of an interaction object participating in the same video interaction can also be displayed. Since the video interaction is performed, in actual application, to ensure that the first interaction object can know all the objects participating in the video interaction, in the login terminal of the first interaction object, a video picture collected in real time by the login terminal of the first interaction object can also be displayed.
[0085] Similarly, in the login terminal of the second interaction object, a video picture collected in real time by the login terminal of the first interaction object can also be displayed, and in actual application, in the login terminal of the second interaction object, a video picture collected in real time by the login terminal of the second interaction object can also be displayed, and a video picture collected in real time by a login terminal of an interaction object participating in the same video interaction can also be displayed, which will not be described herein again.
[0086] For ease of understanding, the display mode of the login terminal of the first interaction object is shown in FIG. 4. In (A) of FIG. 4, a video picture 402 collected in real time by a login terminal of the second interaction object is displayed in the login terminal of the first interaction object. In (B) of FIG. 4, a video picture 404 collected in real time by a login terminal of the second interaction object and a video picture 406 collected in real time by the login terminal of the first interaction object are displayed in the login terminal of the first interaction object.
[0087] As can be known from the foregoing, since the interactive objects participating in the video interaction can also include a third interactive object and a fourth interactive object, if the interactive objects participating in the video interaction also include the third interactive object and the fourth interactive object, the video pictures collected in real time by the login terminal of the second interactive object, the video pictures collected in real time by the login terminal of the first interactive object, the video pictures collected in real time by the terminal logged in by the third interactive object, and the video pictures collected in real time by the terminal logged in by the fourth interactive object can be displayed in the login terminal of the first interactive object. As shown in FIG. 5, in the login terminal of the first interactive object, the video picture 502 collected in real time by the login terminal of the first interactive object, the video picture 504 collected in real time by the login terminal of the second interactive object, the video picture 506 collected in real time by the login terminal of the third interactive object, and the video picture 508 collected in real time by the login terminal of the fourth interactive object are displayed.
[0088] In step 304, in response to the environment content attention event triggered for the environment content in any one video picture, the environment content matched with the environment content attention event is selected as the selected content.
[0089] The environment content refers to various elements contained in the video pictures collected in real time by the login terminal of the interactive object in the video interaction scene, which can be multimedia objects such as video objects, text objects, and graphic-text objects, or real object contents such as books, cups, and dolls, or pictures without any object. These contents are key elements triggering the environment content attention event, and through identification and processing of the elements, the selected content can be determined and the associated content picture can be displayed.
[0090] The environmental content attention event is a key event in a video interaction scene for triggering content recognition and detection of a video picture. The core role thereof is to accurately capture the attention diagram of an interactive object to environmental content in a video picture, thereby providing a basis for subsequent determination of selected content and display of related content pictures. The event can be triggered in at least the following ways: the object line of sight of the interactive object is focused on the environmental content in the video picture, and the focusing time of the focused environmental content reaches a preset focusing time; the interactive object triggers a long press operation on the video picture, and the long press time reaches a preset long press time; the interactive object inputs voice interaction content for the video picture; the interactive object triggers an interactive operation on the detection control; and a preset gesture operation is performed on the video picture. In addition, the event can be triggered by real-time content detection of the video picture based on actual application requirements, or by setting a trigger period to periodically detect the content of the video picture. The trigger period is a time setting method for triggering the environmental content attention event. The user can set a preset time interval according to actual application requirements, and the video picture will be periodically detected for content every preset time interval, thereby triggering the environmental content attention event. This method can realize periodic monitoring of the content of the video picture, and ensure that the environmental content that the interactive object may be interested in is discovered in time.
[0091] The selected content is environmental content determined to match the environmental content attention event in a video interaction scene when the login terminal of the interactive object detects the environmental content attention event triggered for the environmental content in any one video picture. It can be multimedia content played by the terminal device, such as a video object, a text object, a graphic object, etc. It can also be a real object content such as a book, a cup, a doll, etc. The selected content is the basis for displaying associated content pictures in the video interaction page.
[0092] Specifically, in the video interaction page of the login terminal of any interactive object, the video pictures collected in real time by the login terminals of the remaining at least one interactive object are displayed. If the login terminal of the interactive object detects the environmental content attention event triggered for the environmental content in any one of the video pictures, the login terminal can determine the environmental content matching the environmental content attention event in response to the environmental content attention event, and take the environmental content matching the environmental content attention event as the selected content.
[0093] The following introduces a manner of determining a target video picture from multiple video pictures: in an optional embodiment, in a case where the interactive objects participating in the video interaction are at least three, for example, the interactive objects participating in the video interaction are a first interactive object, a second interactive object, and a third interactive object. Based on this, the manner of triggering the environment content attention event based on the environment content in any one video picture includes: in response to the environment content attention operation triggered for one of the video pictures, taking the video picture triggered by the environment content attention operation as a target video picture; and triggering the environment content attention event for the environment content in the target video picture.
[0094] The environment content attention operation can be a long press operation for the video picture for a preset time length. The long press operation is an operation of long pressing the video picture displayed by the login terminal of any interactive object, that is, the long press operation needs to be detected by the touch control detection module of the login terminal (and the touch screen) of the interactive object. Based on this, the preset long press time length is a time threshold set when the interactive object triggers the environment content attention event by triggering the long press operation on the video picture. When the interactive object performs the long press operation on the video picture and the long press time length reaches the preset long press time length, the environment content attention event is triggered. The specific value of the time length needs to be determined according to actual conditions and application requirements, which is an important standard for judging whether the long press operation is valid, and ensures that the long press operation of the interactive object is an intentional content detection behavior. The preset long press time length can be 3 seconds, 5 seconds, or 7 seconds, and the specific preset long press time length needs to be flexibly determined based on actual scene requirements.
[0095] Specifically, in the video interaction page of the login terminal of any interactive object, based on displaying the video pictures collected in real time by the login terminals of the remaining at least one interactive object, in a case where the interactive objects participating in the video interaction are at least three, after detecting the environment content attention operation triggered for one of the video pictures, the login terminal responds to the environment content attention operation, first determines the target video picture indicated by the environment content attention operation, and then triggers the environment content attention event for the environment content in the target video picture. Taking the long press operation of the first interactive object on the login terminal of the first interactive object as an example, the login terminal of the first interactive object detects the long press operation on any video picture, and counts the time length of the long press operation. In a case where the long press time length of the long press operation reaches the preset long press time length, it is indicated that the first interactive object indeed has the operation on any video picture, rather than a false touch or other situations. At this time, the video picture triggered by the environment content attention operation is taken as the target video picture, and the environment content attention event for the environment content in the target video picture is triggered.
[0096] For example, the environmental content focus operation can be a long press operation for a video picture for a preset time length. The video interactive page displays video picture 1, video picture 2, and video picture 3. If the interactive object performs a long press operation for a preset time length on video picture 1, it can be determined that video picture 1 is the target video picture, and then an environmental content focus event for the environmental content in video picture 1 is triggered.
[0097] For ease of understanding, based on the example of the video interactive page shown in FIG. 5, as shown in FIG. 6, the interactive object performs an interactive operation of long pressing video picture 504 for a preset time length. At this time, it can be determined that video picture 504 is the target video picture.
[0098] The following describes a triggering manner of the environmental content focus event. In another optional embodiment, the interactive object participating in the video interaction includes at least a first interactive object and a second interactive object. Based on this, in a case where the video interactive page is a display page of the first interactive object, and the video picture on which the environmental content focus event is triggered is a target video picture corresponding to the second interactive object, that is, as shown in FIG. 6, the video interactive page is a display page of the first interactive object, and the interactive operation of long pressing for a preset time length is performed on video picture 504 corresponding to the second interactive object. At this time, video picture 504 is the target video picture.
[0099] Based on this, the triggering manner of the environmental content focus event includes any one of the following manners: in a case where the voice interactive content of the first interactive object and the second interactive object contains a preset keyword, the environmental content focus event is triggered; in a case where a preset gesture operation is triggered on the target video picture, the environmental content focus event is triggered; in a case where an interactive operation is triggered on the target video picture, the environmental content focus event is triggered.
[0100] The voice interactive content is voice information uttered by the first interactive object and the second interactive object in the video interaction process. When the voice interactive content contains a preset keyword (such as “content detection”, “start recognition”, “what is being watched”, etc.), a content detection event for the video picture is triggered, and then the environmental content focus event is triggered. It provides an interactive object with a way of triggering content detection and display through voice, and increases the diversity and convenience of interaction. The preset gesture operation is a specific gesture action performed by the interactive object on the environmental content in the video picture, such as gesture zooming in or gesture zooming out on a certain environmental content. When the interactive object triggers a preset gesture operation on the environmental content in the video picture, the environmental content focus event is triggered. This operation manner provides an interactive object with an intuitive and convenient means of triggering content detection and display, and enhances the interactivity between the user and the video picture.
[0101] Firstly, it can be understood that in the case that there is no environmental content in the target video screen displayed by the login terminal of the first interactive object, the subsequent steps are not performed. Then in the case that there is at least one environmental content in the target video screen displayed by the login terminal of the first interactive object, the response and processing for the environmental content attention event are triggered. For ease of understanding, as shown in FIG. 7, (A) of FIG. 7 illustrates an example in which there is no environmental content in the target video screen displayed by the login terminal of the first interactive object, and (B) of FIG. 7 illustrates the case that there are environmental content 702 and environmental content 704 in the target video screen displayed by the login terminal of the first interactive object. Based on this, if the first interactive object triggers an environmental content attention event on the video screen on the login terminal, at this time, the content detection and matching identification can be performed in response to the foregoing environmental content attention event, and then the environmental content matched with the environmental content attention event is selected as the selected content.
[0102] The foregoing triggering manner of the environmental content attention event will be introduced below. First, the manner of triggering the environmental content attention event in the case of triggering an interactive operation on the target video screen is introduced. At this time, the environmental content attention event is triggered in the case of triggering an interactive operation on the display area of the environmental content in the target video screen. The interactive operation can be a preset gesture on the display area of the environmental content, a long press on the display area of the environmental content, etc., which is not limited here. The display area of the environmental content is a preset area in the video screen, and the display area must include the complete environmental content in the video screen. The shape of the foregoing preset area can be a quadrilateral or an irregular shape.
[0103] Specifically, the first interactive object performs an interactive operation on the display area of one of the environmental contents in the target video screen. At this time, the login terminal of the first interactive object receives the interactive operation triggered on the environmental content and responds to the interactive operation, so that the login terminal of the first interactive object can determine the environmental content as the selected content. For ease of understanding, further introduction is made based on the long press operation and the example of (B) of FIG. 7. As shown in FIG. 8, the video screen displayed by the login terminal of the first interactive object in real time collected by the login terminal of the second interactive object exists environmental content 702 and environmental content 704, and the display area 802 of the environmental content 702 and the display area 804 of the environmental content 704. If the first interactive object performs a long press operation on the display area 802, at this time, the login terminal of the first interactive object can determine the environmental content 702 in the display area 802 as the selected content in response to the foregoing long press operation.
[0104] Secondly, the first interactive object can also perform a preset gesture operation on the video screen, so that the login terminal of the first interactive object performs content detection and matching recognition in response to the aforementioned preset gesture operation. Therefore, the manner of triggering the environmental content attention event in the case of triggering the preset gesture operation on the target video screen is introduced again:
[0105] Among them, the aforementioned preset gesture operation can be a gesture zoom-in or a gesture zoom-out on certain environmental content. Specifically, the first interactive object performs a preset gesture operation on any environmental content in the video screen, at which time the login terminal of the first interactive object receives the preset gesture operation triggered on the environmental content in the target video screen and responds to the preset gesture operation, at which time the login terminal of the first interactive object can directly determine the environmental content to which the preset gesture operation is performed as the selected content. For ease of understanding, the preset type is gesture zoom-in and the example of (B) in FIG. 7 is further introduced, as shown in FIG. 9, there are environmental content 702 and environmental content 704 in the video screen displayed by the login terminal of the second interactive object and collected in real time by the login terminal of the first interactive object, if the first interactive object performs a gesture zoom-in interaction operation on the environmental content 704, at which time the login terminal of the first interactive object responds to the gesture zoom-in interaction operation and determines the environmental content 704 as the selected content.
[0106] The manner of triggering the environmental content attention event in the case where the voice interaction content between the first interactive object and the second interactive object contains a preset keyword is introduced below:
[0107] Among them, the voice interaction content is the voice content uttered by the first interactive object and the second interactive object in the video interaction, and the preset keyword is a voice for triggering content detection. The preset keyword can be "content detection", "start recognition", etc., and can also be "what are you looking at" and other daily voice expressions, which are not limited here. Specifically, the login terminal of the first interactive object can perform real-time voice detection, that is, voice detection on the voice interaction content collected in real time by the login terminal of the first interactive object itself, and in the case where the voice interaction content contains a preset keyword (such as the aforementioned "content detection", "start recognition", and "what are you looking at", etc.), at which time the content detection event on the video screen is triggered.
[0108] The above describes methods for interactive objects to actively trigger environmental content attention events. In practical applications, if an interactive object is interested in a certain environmental content, its gaze will focus. Therefore, the environmental content attention event can be triggered based on the gaze focus result of the interactive object. In an optional embodiment, the triggering method for the environmental content attention event further includes: when the image of the first interactive object is included in the video frame captured in real time by the login terminal of the first interactive object, performing gaze focus detection on the object's gaze to obtain a gaze focus result; and when the gaze focus result indicates that the object's gaze is focused on the environmental content in the video frame captured in real time by the login terminal of the second interactive object, and the duration of focusing on the environmental content reaches a preset focus duration, triggering the environmental content attention event.
[0109] Among them, gaze focus detection is used to identify the focus area of an object's gaze. Therefore, the gaze focus result is used at least to characterize the focus area of the first interactive object's gaze, and can further determine the environmental content present in the focus area. The preset focus duration is a time standard set when an environmental content attention event is triggered based on the gaze focus result of the interactive object. When the first interactive object's gaze focuses on the environmental content in the video frame, and the focus duration on the environmental content reaches the preset focus duration, an environmental content attention event will be triggered. The specific value of this duration (such as 3 seconds, 5 seconds, or 7 seconds, etc.) needs to be flexibly determined according to the actual scenario requirements. Its function is to determine whether the interactive object's attention to the environmental content has reached a certain level, so as to determine whether to perform subsequent content detection and display operations.
[0110] The gaze focus detection of the first interactive object is performed as follows: The camera of the login terminal is used to capture real-time images of the first interactive object's eyes. Image processing techniques, such as grayscale conversion and filtering, are used to preprocess the eye images to improve image quality. Then, a deep learning-based eye feature extraction model is used to extract key eye features, such as pupil position and eyeball contour.
[0111] To calculate the eye's rotation angle and direction based on extracted eye features, we can use the following method: First, establish an eye coordinate system with the eyeball center as the origin. Then, calculate the pupil's offset relative to the origin using the detected pupil position. Let the horizontal offset of the pupil be x, and the vertical offset be y.
[0112] Horizontal rotation angle θ of the eyeball x It can be done through formula Calculate, where d is the distance from the center of the eyeball to the screen. Vertical rotation angle θ y It can be done through formula calculate.
[0113] According to the calculated rotation angle θ x and θ y , combined with the installation position and angle of the camera, and the size and position information of the screen, the focus area of the object's line of sight can be determined. Among them, x represents the horizontal offset of the pupil, y represents the vertical offset of the pupil, d represents the distance from the center of the eyeball to the screen, θ x represents the horizontal rotation angle of the eyeball, and θ y represents the vertical rotation angle of the eyeball.
[0114] Specifically, the login terminal of the first interactive object can perform content recognition on the real-time collected video stream data to obtain a second content recognition result. The login terminal of the first interactive object can also send the real-time collected video stream data to the server to enable the server to perform content recognition on the video stream data to obtain a second content recognition result. The specific device performing content recognition is not limited here. The purpose of content recognition on the video stream data is to identify the object and object type in the video frame in the video stream data. The second content recognition result is used to represent the object and object type in the video stream data. Since it needs to be considered whether the line of sight of the first interactive object is focused on a certain environmental content in the video picture collected by the login terminal of the second interactive object, the main purpose in judging whether to trigger a content detection event for the video picture in the present application is to identify whether the first interactive object exists in the video stream data. Therefore, the second content recognition result is specifically used to represent that the first interactive object exists in the video stream data, or that the first interactive object does not exist.
[0115] The video stream data is data collected by the image acquisition device and the audio acquisition device of the interactive object login terminal in a video interaction scene. The image acquisition device can be built-in in the terminal or an external camera or the like, and is used to acquire video frames. The audio acquisition device can also be built-in in the terminal or an external microphone, and is used to acquire audio data. The video stream data is the basis for displaying the video picture, and needs to be analyzed and processed to display the corresponding video picture on the login terminal of the interactive object.
[0116] Based on this, in the case that the second content recognition result represents that the image of the first interactive object is included in the video picture collected by the login terminal of the first interactive object in real time, it is indicated that the first interactive object is in the video picture collected by the login terminal of the first interactive object in real time. At this time, the line-of-sight focus detection of the object line-of-sight of the first interactive object can be further performed to obtain a line-of-sight focus result, that is, the focus area of the object line-of-sight of the first interactive object is determined. At this time, it is necessary to judge whether the focus area of the object line-of-sight exists any environmental content. If it exists, that is, the line-of-sight focus result represents that the object line-of-sight of the first interactive object focuses on the environmental content in the video picture, then the timing of the duration that the first interactive object focuses on the environmental content is started. In the case that the duration that the first interactive object focuses on the environmental content reaches the preset focus duration, it is indicated that the first interactive object focuses on the focus area for the preset focus duration, that is, the first interactive object may have interest in the environmental content in the object area. At this time, the content detection event of the video picture is triggered. At this time, the region of the content detection of the video picture is specifically the focus area of the environmental content in the video picture.
[0117] In actual application, the triggering of the environmental content attention event can also be performed by the interactive object actively. In an optional embodiment, the triggering manner of the environmental content attention event further includes: in the case that the first interactive object triggers the interactive operation on the detection control, the environmental content attention event is triggered.
[0118] The detection control is a component for detecting the interactive operation of the interactive object. It can be deployed in the touch screen of the login terminal of the first interactive object and displayed on the terminal, or it can be an external control of the login terminal. When the first interactive object triggers the interactive operation on the detection control, the environmental content attention event is triggered. The detection control provides a way for the interactive object to actively trigger the content detection, which is convenient for the user to operate according to the own demand. In addition, the detection control can also be an external control of the login terminal of the first interactive object, etc. Specifically, when the first interactive object performs the interactive operation on the detection control, the login terminal of the first interactive object detects the interactive operation on the detection control. At this time, the environmental content attention event of the video picture collected by the login terminal of the second interactive object in real time is triggered. Since the detection control is also displayed on the login terminal of the first interactive object, the detection control can be specifically displayed in the video picture collected by the login terminal of the second interactive object in real time. At this time, the detection control can be displayed in the video picture collected by the login terminal of the object participating in the video interaction. The first interactive object performs the interactive operation on the detection control in which video picture. At this time, the environmental content attention event of the video picture is triggered.
[0119] For ease of understanding, as shown in FIG. 10, the login terminal of the first interactive object displays a video screen 1002 collected in real time by the login terminal of the second interactive object, and the video screen 1002 has a detection control 1004. If the first interactive object performs an interactive operation on the detection control 1004, an environmental content attention event for the video screen 1002 is triggered.
[0120] It can be understood that, since the login terminal of the first interactive object can also display a video screen collected in real time by the login terminal of another object participating in the video interaction, the foregoing manner can also be used to trigger the detection of an environmental content attention event for the video screen, and details are not repeated here.
[0121] Step 306. In the video interaction page, a content screen associated with the selected content is displayed.
[0122] The selected content refers to environmental content determined to match an environmental content attention event triggered for any video screen in a video interaction scene after the login terminal of the interactive object detects the environmental content attention event. The selected content can be a real object, such as a book, a cup, and a doll. The selected content can also be multimedia content played by a terminal device, and the foregoing selected content can be a video object, a text object, a graphic object, and the like, which are not limited in type. In addition, the content screen is a screen related to the selected content. For ease of understanding, taking the selected content as a real object as an example, the content screen associated with the selected content can be a global image of the real object. The global image is an image including the complete real object, and the global image can include a front view of the real object, a top view of the real object, and a side view of the real object, and the like. Alternatively, the content screen associated with the selected content can also be a partial image of the real object. The partial image is an image not including the complete real object. For example, when the real object is a doll, the global image can be a front view including the entire doll, and the partial image can be a front view including only the head of the doll or a front view including only the half body of the doll. Considering that the real object can usually be purchased online, the content screen associated with the selected content can also be a content promotion page of the real object, which is not limited here.
[0123] If taking the multimedia content played by the terminal device as an example, the content picture associated with the selected content can be the multimedia content, or the content picture associated with the selected content can also be a target application program, which is an application program for displaying the multimedia content by the terminal device, or the content picture associated with the selected content can also be download prompt information for the target application program to which the multimedia content belongs, or the content picture associated with the selected content is a delivery selection area for the multimedia content, and the like, which is not limited herein.
[0124] Specifically, the terminal displays the content picture associated with the selected content on the video interaction page. The content picture associated with the selected content can be displayed simultaneously with the video picture collected by the terminal in real time, or can not be displayed simultaneously, which is not limited herein. Taking the simultaneous display as an example, the interaction object for the video interaction includes at least a first interaction object and a second interaction object, and the video interaction page is a display page of the first interaction object, and at this time, the content picture associated with the selected content can be displayed simultaneously with the video picture collected by the terminal of the other interaction object.
[0125] It can be understood that the terminal of the interaction object participating in the video interaction can also send the video stream data collected in real time to the server, and the server performs content detection and identification based on the video stream data sent by each terminal, so as to detect whether the selected content exists in the video picture collected by the terminal of each interaction object in real time. In the case where the selected content is determined in the video picture, when the video stream data of the terminal of the interaction object is sent to the terminal of the other interaction object, the related information indicating the selected content is also carried, so that the video picture and the content picture associated with the selected content are displayed simultaneously in the terminal of the interaction object.
[0126] For the convenience of understanding, taking the display of the terminal of the first interaction object as an example, as shown in FIG. 11, (A) of FIG. 11 illustrates that the video picture 1102 collected by the terminal of the second interaction object in real time and the content picture 1104 associated with the selected content are displayed simultaneously in the terminal of the first interaction object. While (B) of FIG. 11 illustrates that the video picture 1102 collected by the terminal of the second interaction object in real time, the video picture 1106 collected by the terminal of the first interaction object in real time, and the content picture 1104 associated with the selected content are displayed simultaneously in the terminal of the first interaction object.
[0127] Based on this, since the video screen in real time collected by the login terminal of the second interactive object needs to be displayed together with the content screen associated with the selected content, how to display the video screen and the content screen at the same time is as follows: the method further includes: analyzing the video stream data in real time collected by the login terminal of the second interactive object to obtain the video screen corresponding to the video stream data; performing content matching based on the selected content to determine the content screen associated with the selected content.
[0128] The second video stream data is video stream data in real time collected by the login terminal of the second interactive object, and the video stream data specifically includes video frames collected by an image collection device and can also include audio data collected by an audio collection device. The image collection device can be built-in in the login terminal of the interactive object or can be a device such as a camera externally connected to the login terminal of the interactive object and capable of collecting images. In addition, the content screen is similar to the foregoing embodiments, and is a screen related to the object type of the selected content, which will not be described herein again.
[0129] Specifically, the server can analyze the video stream data in real time collected by the login terminal of the second interactive object to obtain the video screen corresponding to the video stream data, and then the server can further distribute the video screen corresponding to the video stream data to the login terminal of the first interactive object. Similarly, the server can perform content matching on the selected content and the object type of the selected content to determine the content screen associated with the selected content. For example, when the selected content is a real object content, the content screen associated with the selected content can be a panoramic image of the real object content, or the content screen associated with the selected content can also be a partial image of the real object content, or the content screen associated with the selected content can also be a content promotion page of the real object content. Then the server can further distribute the content screen associated with the selected content to the login terminal of the first interactive object. Based on this, the login terminal of the first interactive object can simultaneously display the video screen and the content screen associated with the selected content.
[0130] In actual application, the server can also receive the video stream data in real time collected by the login terminal of the second interactive object, and then distribute the video stream data in real time collected by the login terminal of the second interactive object to the login terminal of the first interactive object. The login terminal of the first interactive object can analyze the video stream data by the foregoing similar method to obtain the video screen corresponding to the video stream data, and perform content matching on the selected content and the object type of the selected content to determine the content screen associated with the selected content. Thus, the video screen and the content screen associated with the selected content can be simultaneously displayed in the login terminal of the first interactive object.
[0131] The following introduces the way of displaying the content picture associated with the selected content: in one specific embodiment, on the video interactive page, the content picture associated with the selected content is displayed in any of the following ways: the content picture associated with the selected content is displayed in full screen mode, and the video picture is displayed on the content picture associated with the selected content in floating window mode; the video picture is displayed in full screen mode, and the content picture associated with the selected content is displayed on the video picture in floating window mode; the video picture is displayed through a first floating window, and the content picture associated with the selected content is displayed through a second floating window.
[0132] The following introduces specific display modes of displaying the content picture associated with the selected content:
[0133] 1. The content picture associated with the selected content is displayed in full screen mode, and the video picture is displayed on the content picture associated with the selected content in floating window mode. The floating window mode means that the picture floats on other pictures. Specifically, when the content picture associated with the selected content is displayed, the content picture associated with the selected content can be displayed in full screen mode in the login terminal of the interactive object, so as to ensure that the corresponding content picture can be clearly and intuitively displayed to the interactive object. At this time, since the aforementioned content picture has been displayed in full screen mode, in order to ensure that the video picture can also be displayed at the same time, the video picture is displayed on the content picture associated with the selected content in floating window mode. For ease of understanding, the login terminal of the first interactive object is taken as the display reference, as shown in FIG. 12, the content picture 1202 associated with the selected content is displayed in full screen mode in the login terminal of the first interactive object, and the video picture 1204 is displayed on the content picture 1202 associated with the selected content in floating window mode.
[0134] The floating window mode is a picture display mode, which means that one picture floats on other pictures. In the content display method of the present application, when the video picture and the content picture associated with the selected content need to be displayed at the same time, the floating window mode can be adopted. For example, when the content picture associated with the selected content is displayed in full screen mode, the video picture is displayed on the content picture in floating window mode; or when the video picture is displayed in full screen mode, the content picture associated with the selected content is displayed on the video picture in floating window mode; or the video picture can be displayed through a first floating window, and the content picture associated with the selected content can be displayed through a second floating window. This display mode can display other related content at the same time without affecting the display of the main picture, thereby improving the flexibility of display and the efficiency of information display.
[0135] 2. Display the video screen in full-screen mode, and display the content screen associated with the selected content on the video screen in a floating window. Specifically, when displaying both the video screen and the content screen simultaneously, the video screen can be displayed in full-screen mode on the login terminal of the interactive object. To ensure that the content screen associated with the selected content can also be displayed simultaneously, the content screen associated with the selected content is displayed on the video screen in a floating window. For ease of understanding, taking the login terminal of the first interactive object as the display reference, as shown in Figure 13, in the login terminal of the first interactive object, the video screen 1302 is displayed in full-screen mode, and the content screen 1304 associated with the selected content is displayed on the video screen 1302 in a floating window.
[0136] 3. Display the video image through a first floating window and display the content image associated with the selected content through a second floating window. Specifically, considering that the object can perform other interactive processing simultaneously while performing video interaction, displaying the image entirely in full-screen mode may not meet other interactive needs. This application also provides a display method in the login terminal of the interactive object, in which the video image is displayed through a first floating window and the content image associated with the selected content is displayed through a second floating window. For ease of understanding, taking the login terminal of the first interactive object as the display reference, as shown in Figure 14, in the login terminal of the first interactive object, the video image 1402 is displayed through the first floating window and the content image 1404 associated with the selected content is displayed through the second floating window. It can be understood that when there are multiple interactive objects, a similar display method can be used to display them simultaneously as shown in Figure 11. Also, the corresponding examples in the embodiments of this application are used to understand this solution, but should not be construed as specific limitations on this solution.
[0137] By displaying video and content images simultaneously using the aforementioned methods, it is possible to meet various display needs in practical applications and improve the feasibility and flexibility of content display.
[0138] The following describes how to display content images associated with the selected content when the selected content is a physical object. In one specific embodiment, displaying content images associated with the selected content on the video interaction page includes at least one of the following methods: displaying the selected content shown at a preset angle on the video interaction page; displaying a global image of the selected content on the video interaction page; displaying a promotional page for the selected content on the video interaction page; or displaying a detailed page for the selected content on the video interaction page.
[0139] The following sections describe the methods for displaying content on the screen:
[0140] 1. In the video interactive page, display the selected content displayed at a preset angle. The preset angle can be 90° or other angles, which is not limited here. The selected content displayed at the preset angle can be a partial image of the real object content, which is an image not including the complete real object content. For example, when the real object image is a doll, the partial image can be a front view including only the head of the doll, or a front view including only the half body of the doll, or a side view of the doll.
[0141] Specifically, in the video interactive page, display the selected content displayed at a preset angle. As known from the foregoing, the selected content displayed at the preset angle can be displayed in a full-screen mode in the video interactive page, and the video screen can be displayed on the content screen associated with the selected content in a floating window mode. Alternatively, the video screen can be displayed in a full-screen mode, and the selected content displayed at the preset angle can be displayed on the video screen in a floating window mode. Alternatively, the video screen can be displayed through a first floating window, and the selected content displayed at the preset angle can be displayed through a second floating window. The preset angle is an angle value set when the selected content is displayed at a specific angle in the video interactive page, such as 90° or other angles. When the selected content is a real object content, a partial image of the real object content can be displayed at the preset angle, such as a front view including only the head of the doll, a front view including only the half body of the doll, or a side view of the doll. The setting of the preset angle provides a diversified way to display the selected content, meeting different display requirements.
[0142] For ease of understanding, the display mode in which the content screen associated with the selected content is displayed in a full-screen mode, and the video screen is displayed on the content screen associated with the selected content in a floating window mode is taken as an example, and the selected content is specifically a color palette, which is taken as an example for introduction. As shown in FIG. 15, a partial image 1502 of the real object content is displayed in a full-screen mode, and a video screen 1504 is displayed on the partial image 1502 of the real object content in a floating window mode.
[0143] 2. In the video interactive page, display a global image of the selected content. The global image is an image including the complete real object content, and can include a front view of the real object content, a top view of the real object content, and a side view of the real object content, etc. Specifically, in the video interactive page, display the global image of the selected content, and the specific display mode is similar to the foregoing embodiments, which is not described here again. For ease of understanding, the display mode in which the content screen associated with the selected content is displayed in a full-screen mode, and the video screen is displayed on the content screen associated with the selected content in a floating window mode is taken as an example, and the selected content is specifically a color palette, which is taken as an example for introduction. As shown in FIG. 16, a global image 1602 of the real object content is displayed in a full-screen mode, and a video screen 1604 is displayed on the global image 1602 of the real object content in a floating window mode.
[0144] 3. On the video interaction page, display the content promotion page for the selected content. Considering that physical products can usually be purchased online, the content screen associated with the selected content can also be a content promotion page for the physical product. Therefore, the video screen and the content promotion page recommending physical products can be displayed simultaneously on the login terminal of the interactive object. The content promotion page is a page used to recommend physical products; therefore, it includes the physical product itself and descriptive text describing the physical product. In practical applications, since the content promotion page is used to recommend physical products, after the first interactive object interacts with the displayed content promotion page, it indicates that the interactive object may be interested in the physical product. At this point, it can also redirect to the shopping platform indicated by the content promotion page. Specific details regarding subsequent steps are not limited here. The content promotion page is a page related to the selected content (especially physical product) and is used to recommend the physical product. This page contains the physical product and text information describing the physical product. When the interactive object interacts with the displayed content promotion page, it may redirect to the promotional shopping platform indicated on the page, providing users with a way to purchase physical products and simultaneously promoting the physical products.
[0145] To facilitate understanding, we will use a display method where the video screen is displayed in full screen and the content screen associated with the selected content is displayed on the video screen in a floating window as an example. The selected content is specifically a color palette, as shown in Figure 17. The video screen 1702 is displayed in full screen and the content promotion page 1704 recommending physical content is displayed on the video screen 1702 in a floating window.
[0146] 4. On the video interaction page, display the content details page for the selected content. Similar to the content promotion page, the content details page for the selected content does not provide promotional redirects; it only displays detailed information about the selected content. The specific display method is similar to the aforementioned embodiments and will not be repeated here. The content details page is a page related to the selected content. Unlike the content promotion page, it does not provide promotional redirects; it only displays detailed information about the selected content. When the selected content is physical content or other types of content, the content details page can provide the interactive object with a comprehensive and detailed introduction to the selected content, helping users better understand the selected content.
[0147] It is understood that the corresponding examples in the embodiments of this application are used to understand this solution, but should not be construed as specific limitations on this solution.
[0148] The method shown in the foregoing content displays, in a video interaction scene in which at least two interactive objects participate, in a video interaction page of one of the interactive objects, a video picture collected in real time by a login terminal of the remaining at least one interactive object, to ensure real-time display of the picture during video interaction. Based on this, in response to an environmental content attention event triggered for environmental content in any one video picture, environmental content matched with the environmental content attention event is selected as selected content, and a content picture associated with the selected content is displayed in the video interaction page. Without a corresponding content sharing operation by the interactive object, the selected content is automatically displayed through the process of environmental content triggering and content display, so that the continuity of video interaction can be avoided during video interaction due to object operation, and the display of the selected content can be performed without multiple sharing interactions and the like, thereby improving the efficiency of content display.
[0149] The following describes how to determine the selected content. In one specific embodiment, as shown in FIG. 18, in response to an environmental content attention event triggered for environmental content in any one video picture, environmental content matched with the environmental content attention event is selected as selected content, including:
[0150] Step 1802: In response to an environmental content attention event triggered for environmental content in any one video picture, a target video picture indicated by the environmental content attention event is determined.
[0151] The target video picture refers to a video picture determined in response to an interest triggering operation triggered for one of the video pictures in a case where the interactive objects participating in video interaction are at least three, and a subsequent environmental content attention event is triggered for environmental content in the picture.
[0152] Specifically, on the basis of displaying, in a video interaction page of a login terminal of any interactive object, a video picture collected in real time by a login terminal of the remaining at least one interactive object, if the login terminal of the interactive object detects an environmental content attention event triggered for environmental content in any one of the video pictures, the login terminal determines a target video picture indicated by the environmental content attention event in response to the environmental content attention event. For example, if video picture 1, video picture 2, and video picture 3 are displayed in the video interaction page, and the environmental content attention event is triggered for environmental content in video picture 1, video picture 1 is the target video picture. The manner of determining the target video picture is similar to that in the foregoing embodiment, which is not described herein again.
[0153] Step 1804: A target video frame is filtered out from video stream data containing the target video picture.
[0154] The target video frame refers to a video frame selected from the video stream data containing the target video picture. Since not all video frames in the video stream data are available frames, key frame extraction is needed for the video stream data, the continuous video frames in the video stream data are pre-processed to improve the image quality, and then the video frames with large changes are selected as target video frames through comparison. Key frame extraction is an important step for selecting target video frames from video stream data. Since not all video frames in the video stream data are available frames, the video stream data needs to be processed. The specific method is to first pre-process the continuous video frames in the video stream data, and then improve the image quality through image enhancement algorithm, and then select the video frames with large changes as target video frames through comparison. The comparison and selection methods include calculating the difference between two consecutive frames, analyzing the object motion trajectory, comparing with the static background, comparing the color histogram, etc. Through key frame extraction, the computational burden can be reduced and the efficiency of content recognition can be improved. Image enhancement algorithm is an algorithm for improving the image quality of video frames in video stream data, such as histogram equalization. When processing the video stream data, the image enhancement algorithm can improve the problem of uneven illumination that may exist in the video frame, make the gray scale distribution of the image more uniform, enhance the global contrast of the image, improve the clarity and recognizability of the image, and facilitate subsequent video frame selection, content recognition and other operations.
[0155] Specifically, the login terminal of the interactive object acquires the video stream data containing the target video picture, that is, the login terminal of the interactive object needs to send a data acquisition request to the server, and the data acquisition request is used to request the video stream data containing the target video picture collected by the login terminal in real time. Then, the login terminal selects the target video frame from the video stream data containing the target video picture.
[0156] Since the login terminal of the interactive object matched with the target video picture can call the camera API to collect the video stream data in real time, and the analysis on the video stream data can display the corresponding video picture on the login terminal of any interactive object, but not all video pictures (i.e. video frames in the video stream data) are available video frames, it is necessary to extract key frames from the video stream data, so as to obtain the target video frame through key frame screening. That is, the continuous video frames in the video stream data can be first image preprocessed, and the image quality of each video frame can be improved through an image enhancement algorithm. The histogram equalization is selected as the image enhancement algorithm, because it can effectively enhance the global contrast of the image, has a good improvement effect on the problem of uneven illumination that may exist in the video frame, can make the gray scale distribution of the image more uniform, thereby improving the clarity and recognizability of the image, and facilitating subsequent video frame selection and content recognition operations. Taking the histogram equalization as an example, the specific operation is as follows: first, the number of pixels of each gray scale level in the video frame image is counted to obtain a gray scale histogram. Then, the cumulative distribution function (CDF) of each gray scale level is calculated, and the CDF is normalized to make its value range between 0-255. Finally, the gray scale level of each pixel of the original image is mapped according to the normalized CDF to obtain the histogram equalized image. In this way, the image quality of each video frame is improved, and then the video frame selection logic is implemented to process only the video frames with large changes among the multiple video frames, thereby reducing the computational burden. Image preprocessing is a series of processing operations on the video frames before key frame extraction and content recognition of the video stream data, and the purpose is to improve the image quality for subsequent analysis and processing. Common image preprocessing methods include grayscale, filtering, histogram equalization, etc. For example, the histogram equalization can effectively enhance the global contrast of the image, improve the problem of uneven illumination, make the gray scale distribution of the image more uniform, and improve the clarity and recognizability of the image.
[0157] Based on this, the target video frame is screened from the video stream data containing the target video picture, specifically: the multiple video frames in the video stream data are image preprocessed to improve the image quality of each video frame, and then the image preprocessed video frames are compared in succession to screen the video frames with large changes in the image preprocessed video frames, so as to determine the video frames with large changes in the image preprocessed video frames as the target video frame.
[0158] The way to compare and screen the video frames with large changes in the image preprocessed video frames is specifically introduced as follows:
[0159] 1. Calculate the image difference between the two consecutive image preprocessed video frames, and determine that the latter one of the two consecutive image preprocessed video frames is the video frame with large changes when the image difference exceeds a preset difference threshold.
[0160] 2. Generate image sequences for the same object in the pre-processed video frames, and then analyze the motion trajectories of the same object in the image sequences to estimate the motion patterns between the pre-processed video frames. If the motion pattern represents a vector change of the optical flow of the object greater than a change threshold, it means that the change between frames is large, that is, the video frames with a vector change greater than the change threshold are screened as video frames with large changes. The specific analysis method is as follows: First, use a target following algorithm (such as KCF following algorithm) to follow the same object in consecutive video frames, and record the position information of the object in each frame. Then, according to the position information of the object in different frames, the motion trajectory of the object is generated. Next, the optical flow vector between adjacent points on the motion trajectory is calculated, which represents the motion direction and speed of the object between adjacent frames. If the motion pattern represents a vector change of the optical flow of the object greater than a change threshold, it means that the change between frames is large, that is, the video frames with a vector change greater than the change threshold are screened as video frames with large changes.
[0161] 3. In the case of a static background in the video picture, compare each pre-processed video frame with the static background to detect the change of the video frame, and determine that the video frame is a large change video frame if the difference between the video frame and the static background exceeds a preset difference threshold.
[0162] 4. Compare the color histograms of two consecutive pre-processed video frames, and determine that the video frame is a large change video frame if the difference between the color histograms exceeds a preset difference threshold.
[0163] In practical applications, all video frames in the video stream data can also be directly determined as target video frames, or the similarity between consecutive video frames in the video stream data is calculated, and in the case that the similarity of multiple consecutive video frames is greater than a preset similarity threshold, one of the multiple video frames with similarity greater than the preset similarity threshold is selected as the target video frame. As can be seen, the image quality of each video frame in the video stream data can be improved by the image enhancement algorithm, and then the target video frame is extracted and screened by the method introduced above (such as the color histogram-based method, the motion vector-based method, the Local Binary Pattern (LBP)-based method, etc.). After the target video frame is extracted, the target video frame can be further preprocessed to facilitate subsequent content recognition of the target video frame. At this time, the preprocessing of the target video frame can be: scaling the target video frame to the desired size of the pre-trained model, such as 224x224 pixels, and then performing grayscale on the target video frame after size scaling to reduce the computational complexity. In addition, the pixel values of the target video frame can also be normalized to 0-1 to improve the stability of data processing. The Local Binary Pattern is a method for image feature extraction, which can be used in the present application to extract and screen target video frames from video stream data. It generates a binary pattern by comparing the gray values of each pixel point and its neighborhood pixel points in the image, thereby describing the local texture features of the image. The method based on the Local Binary Pattern can effectively extract the feature information in the video frame, help to screen out the video frame with large changes as the target video frame, and improve the accuracy of content recognition.
[0164] It can be understood that the way of how to screen the target video frame from the video stream data needs to be flexibly determined based on actual situations and scene requirements, and should not be understood as a specific limitation of the present application.
[0165] Step 1806, performing environmental content recognition on the target video frame, and taking the environmental content represented by the content recognition result as the selected content.
[0166] Among them, the environmental content recognition includes target video frame recognition and object matching, that is, the environmental content recognition is used to determine the objects existing in the target video frame and the object type of each object, and the candidate objects corresponding to the object and the object type are matched. The selected content can be an object belonging to a preset object type, or an object with the largest proportion in the content detection frame of the target video frame, which is not limited here. The content detection frame is used to define the region of the object in the target video frame when the environmental content recognition is performed on the target video frame. The selected content can be the object with the largest proportion in the content detection frame, or the object belonging to the preset object type, which helps to determine the selected content by analyzing and recognizing the objects in the content detection frame.
[0167] Specifically, the login terminal of the interactive object can perform environment content recognition on the target video frame to obtain a first environment content recognition result for the target video frame, and then determine that the video picture includes the selected content based on the first environment content recognition result. That is, the target video frame is first subjected to environment content recognition. As introduced in the foregoing, after each login terminal collects video stream data in real time and performs corresponding transmission, in order to improve the image quality of each video frame in the video stream data, each video frame in the video stream data can be subjected to image enhancement processing through an image enhancement algorithm. Then, the target video frame is extracted and screened through a method such as the method based on a color histogram, the method based on a motion vector, the method based on a local binary pattern (LBP), and the like. The target video frame is subjected to environment content recognition by using an image recognition model. In consideration of the stability and efficiency of the model, the login terminal of the interactive object performs environment content recognition on the target video frame, including: scaling the extracted target video frame to a size supported by the image recognition model, such as 224x224 pixels, then performing grayscale on the target video frame after the size is scaled, to reduce the calculation complexity of the image recognition model, and the pixel values of the target video frame can also be normalized to 0-1, to improve the stability.
[0168] After the foregoing processing is completed, the target video frame after preprocessing is subjected to feature extraction on the target video frame by using an image recognition model such as a convolutional neural network. Specifically, the target video frame after preprocessing is input into the convolutional neural network, which is composed of multiple convolutional layers and pooling layers alternately. In the convolutional layer, each convolutional filter performs sliding convolution operation on the input image to capture specific patterns or features in the image, such as edges, textures, color blocks, corner points, etc., to generate a feature map. The pooling layer performs down-sampling on the feature map to reduce the spatial size of the image while retaining important feature information. After the processing of multiple convolutional layers and pooling layers, a high-level feature representation of the image is obtained. These features can be regarded as a high-level representation of the image, which captures the main information of the image and can be used for subsequent environment content recognition.
[0169] In order to determine the object type of the object in the target video frame, a fully connected layer can be added to the output layer of the convolutional neural network, and the number of neurons of the fully connected layer is the same as the number of preset object types. The probability distribution of each object type is obtained by performing a softmax activation function on the output of the fully connected layer. The object type with the maximum probability is selected as the object type of the object.
[0170] When performing object matching, a candidate object database of object types is constructed in advance, and each candidate object in the database has a corresponding feature vector. The object feature vector extracted from the target video frame is compared with the feature vectors in the candidate object database for similarity calculation, for example, using the cosine similarity formula wherein is the feature vector of the object in the target video frame, is the feature vector of the candidate object. The candidate object with the highest similarity is selected as the object of the detection box. Among them, object matching is an important link in environmental content recognition. After determining the object type of the object in the target video frame, the object feature vector extracted from the target video frame is compared with the feature vectors in the candidate object database of the object type constructed in advance for similarity calculation, for example, using the cosine similarity formula. By comparing the similarity, the candidate object with the highest similarity is selected as the object of the detection box, thereby completing the object matching and providing a basis for determining the selected content.
[0171] Then the object in the content detection box with the largest proportion in the target video frame is determined as the selected content. Alternatively, objects belonging to a preset object type are directly selected, and similarity screening is performed from the candidate object database of the preset object type to determine the selected content of the preset object type. Among them, is the feature vector of the object in the target video frame, is the feature vector of the candidate object, is the cosine similarity between the two vectors.
[0172] Since the similarity between objects is considered during screening, the similarity can be obtained by comparing the image features and the candidate object features in the candidate object database, thereby querying the most similar object. At this time, some efficient search algorithms can be used, such as approximate nearest neighbor search, and some optimized data structures such as K-dimension tree (KD) tree, ball tree, etc. In addition, hash technology (such as Locality-Sensitive Hashing (LSH)) can also be used to further improve the query screening speed, which is not limited here.
[0173] It can be understood that the corresponding examples in the embodiments of the present application are used to understand the present scheme, but should not be understood as a specific limitation of the present scheme.
[0174] In the above embodiments, by screening the target video frame, the full video frame can be avoided for screening and content detection, thereby ensuring the reliability and efficiency of content detection, thereby more accurately and efficiently determining the selected content through content recognition, and further improving the reliability and efficiency of displaying the content picture associated with the selected content.
[0175] The following describes a manner of displaying a content screen associated with selected content in a case where the selected content is multimedia content played by a terminal device. In one embodiment, as shown in FIG. 19, a terminal device playing multimedia content is displayed in a video screen, and the selected content is the multimedia content played by the terminal device.
[0176] That is, the terminal device exists in the video screen, and the terminal device plays multimedia content, and the selected content can be the multimedia content played by the terminal device. The foregoing selected content can be a video object, a text object, a graphic-text object, etc. For example, taking the multimedia content as a video object as an example, the multimedia content can be “A drama, B episode, C minute”. Taking the multimedia content as a text object as an example, the multimedia content can be “A article, B page”.
[0177] Based on this, the content screen associated with the selected content is displayed in the video interactive page, including:
[0178] Step 1902, playing multimedia content in the video interactive page.
[0179] Specifically, the multimedia content is played in the video interactive page of the login terminal of the interactive object. In actual application, the video screen and the multimedia content can be simultaneously displayed in the video interactive page of the login terminal of the interactive object. Taking the login terminal of the first interactive object as a display reference, the video screen collected in real time by the login terminal of the second interactive object can be displayed in the video interactive page of the first interactive object. Since the terminal device playing multimedia content exists in the video screen collected in real time by the login terminal of the second interactive object, the multimedia content can be played in the video interactive page of the first interactive object. For ease of understanding, please refer to FIG. 20. As shown in (A) of FIG. 20, the video screen 2002 collected in real time by the login terminal of the second interactive object is displayed in the login terminal of the first interactive object. The terminal device 2006 playing the multimedia content 2004 exists in the video screen 2002, and the multimedia content 2004 is determined as the selected content. As shown in (B) of FIG. 20, the video screen 2002 and the multimedia content 2004 are displayed in the login terminal of the first interactive object.
[0180] In one specific embodiment, the terminal device plays the multimedia content through the target application, that is, the multimedia content is played by the terminal device through the target application. How to identify the manner in which the terminal device plays the multimedia content through the target application includes: determining a candidate application matching the category of the multimedia content; extracting a content feature of the multimedia content; performing similarity analysis on the content feature and an interface feature corresponding to each candidate application respectively to obtain a feature similarity; and taking the candidate application whose feature similarity meets a matching requirement as the target application. The matching requirement is a standard for judging whether the candidate application can be the target application, and is determined based on the feature similarity. After the similarity analysis on the content feature of the multimedia content and the interface feature of the candidate application to obtain the feature similarity, when the feature similarity meets a certain condition (such as the distance being less than a preset threshold), it is considered that the candidate application meets the matching requirement, and can be taken as the target application.
[0181] The content feature refers to a feature obtained by extracting the feature of the multimedia content through the trained image recognition model. The image recognition model usually alternately comprises a plurality of convolution layers and pooling layers. Each convolution layer applies a set of convolution filters to an input image to capture a specific pattern or feature in the image, such as an edge, a texture, a color block, a corner point, etc. The pooling layer reduces the spatial size of the image while retaining important feature information, thereby extracting the content feature of the multimedia content.
[0182] The extraction of the content feature of the multimedia content and the similarity analysis need to be based on the image recognition model. First, the image recognition model is introduced. The image recognition model in the present application is a CNN, which is mainly used for feature extraction. The CNN model alternately comprises a plurality of convolution layers and pooling layers. Each convolution layer applies a set of convolution filters to an input image, and each filter captures a specific pattern or feature in the image, such as an edge, a texture, a color block, a corner point, etc. The pooling layer reduces the spatial size of the image while retaining important feature information. These features can be regarded as a high-level representation of the image, which captures the main information of the image, and thus the content feature of the multimedia content can be extracted.
[0183] Then, for each content sample and the application sample to which the content sample belongs, the sample features corresponding to the content sample can include the color scheme of the application sample, the shape and position of the buttons in the application sample, the font and size of the text in the application sample, the UI component layout style in the application sample, the application logo of the application sample, and interface elements specific to the application sample, etc. Then, by performing data augmentation on each content sample, more content samples can be generated, thereby improving the generalization ability of the model. In the training process, the method of data augmentation on the content sample includes but is not limited to rotation, scaling, cropping, flipping, color transformation, etc. For example, the screenshot of the application sample can be taken as the content sample, and the content sample can be rotated and scaled to obtain a new content sample.
[0184] Therefore, based on the foregoing features, the image recognition model (mainly a classifier) is trained to identify the application to which the image recognition model belongs. The specific training steps are as follows: First, a large number of content samples and the application samples to which the content samples belong are collected. The sample features corresponding to the content sample can include the color scheme of the application sample, the shape and position of the buttons in the application sample, the font and size of the text in the application sample, the UI component layout style in the application sample, the application logo of the application sample, and interface elements specific to the application sample, etc. Then, each content sample is subjected to data augmentation such as rotation, scaling, cropping, flipping, color transformation, etc. to generate more content samples, thereby improving the generalization ability of the model. Next, a pre-trained convolutional neural network (such as ResNet, VGG, etc.) is used as a base model, and the content sample is input into the pre-trained model to extract features. In the training process, some layers of the pre-trained model are frozen, and only the classifier layer is trained. This way, the knowledge learned by the pre-trained model can be utilized to improve the training efficiency and recognition accuracy. At the same time, since the pre-trained model is trained on a large amount of image data, it can recognize various image features, which can improve the generalization ability of the model and enable it to better handle various types of images. Finally, the trained model is used to classify new images and identify the application to which they belong.
[0185] Based on this, because there are content samples and application samples to which the content samples belong, an application feature library can be constructed for different application types to which the application samples belong, and the application feature library contains interface features under different application types. The specific process of constructing the application feature library is as follows: first, a large number of interface samples of different applications are collected, which should cover various application types. For each application sample, extract its interface features, including the color scheme of the application, the shape and position of the button, the font and size of the text, the UI component layout style, the application logo, and specific interface elements. Quantify and encode these features to form a feature vector. Then, according to the application type, these feature vectors are classified and stored to form an application feature library. In practical applications, by comparing the content features with the interface features of different application types in the application feature library, the target application to which the selected content belongs is found. The interface features are the features extracted from the application feature library, each corresponding to a candidate application, including the color scheme of the application, the shape and position of the button, the font and size of the text, the UI component layout style, the application logo, and specific interface elements. These features are quantified and encoded to form a feature vector for similarity analysis with the content features of the multimedia content.
[0186] The specific training process is as follows: first, a large number of content samples and application samples to which the content samples belong are collected, and the sample features corresponding to the content samples can include the color scheme of the application sample, the shape and position of the button in the application sample, the font and size of the text in the application sample, the UI component layout style in the application sample, the application logo of the application sample, and specific interface elements of the application sample. Then, data augmentation is performed on each content sample, such as rotation, scaling, cropping, flipping, color transformation, etc., to generate more content samples and improve the generalization ability of the model. Next, use a pre-trained convolutional neural network (such as ResNet, VGG, etc.) as a base model, input the content sample into the pre-trained model, and extract the features. During the training process, freeze some layers of the pre-trained model, and only train the classifier layer, which can utilize the knowledge learned by the pre-trained model to improve the training efficiency and recognition accuracy. At the same time, since the pre-trained model is trained on a large amount of image data, it can recognize various image features, which can improve the generalization ability of the model and make it better handle various types of images. Finally, use the trained model to classify new images and identify their belonging application.
[0187] The content features are compared with interface features of different application program types in the application program feature library. The specific comparison method is as follows: first, the content features and the interface features in the application program feature library are represented by feature vectors. Then, an efficient search algorithm (such as approximate nearest neighbor search) and an optimized data structure (such as KD tree, ball tree, etc.) are used to calculate the distance (such as Euclidean distance, cosine distance, etc.) between the content feature vector and each interface feature vector in the application program feature library. The smaller the distance, the higher the similarity. Hash technology (such as local sensitive hashing (LSH)) can also be used to further improve the query screening speed. Finally, the application program corresponding to the interface feature with the smallest distance (the highest similarity) is selected as the target application program to which the selected content belongs.
[0188] Specifically, the server first determines the candidate application program matching the category of the multimedia content. Specifically, an application program database is constructed in advance, which records the multimedia content categories that each application program can support. When the category of the multimedia content is a text category, the application programs supporting the display of text are filtered out from the application program database as candidate application programs; similarly, when the category of the multimedia content is a video category, the application programs supporting the playing of videos are filtered out from the application program database as candidate application programs. The application program database is a pre-constructed database, which records the multimedia content categories that each application program can support. When it is necessary to determine the candidate application program matching the category of the multimedia content, the database can be filtered to determine the candidate application program. For example, when the category of the multimedia content is a text category, the application programs supporting the display of text are filtered out from the database as candidate application programs; when the category of the multimedia content is a video category, the application programs supporting the playing of videos are filtered out as candidate application programs.
[0189] For example, an application program database is constructed in advance, which records the multimedia content categories that each application program can support. Let the category of the multimedia content be C, the application program database be D, and for each application program A in the database D i , the supported multimedia content category set is S i . When C is a text category, the application program A i satisfying C∈S i is filtered out as a candidate application program; similarly, when C is a video category, the application program A i satisfying C∈S i is also filtered out as a candidate application program.
[0190] Based on this, the server extracts the content features of the multimedia content through the trained image recognition model. Specifically, the image of the multimedia content is input into the image recognition model, which is composed of multiple convolution layers and pooling layers alternately. Each convolution layer applies a set of convolution filters to the input image, and each filter captures a specific pattern or feature in the image, such as edges, textures, color blocks, corner points, etc. The pooling layer reduces the spatial size of the image while retaining important feature information. Through these convolution and pooling operations, the image recognition model can extract the content features of the multimedia content. Then the interface features corresponding to each candidate application are extracted from the application feature library, and the content features and the interface features corresponding to each candidate application are analyzed for similarity, and the feature similarity is obtained. The candidate application corresponding to the interface feature with the largest feature similarity value is selected as the target application to which the selected content belongs.
[0191] For example, there are candidate application A1, candidate application A2, and candidate application A3, interface feature B1 corresponding to candidate application A1 is extracted, interface feature B2 corresponding to candidate application A2 is extracted, and interface feature B3 corresponding to candidate application A3 is extracted. The feature similarity between the content features and the interface features B1, B2, and B3 is calculated. If the feature similarity between the content features and the interface feature B1 is 40%, the feature similarity between the content features and the interface feature B2 is 80%, and the feature similarity between the content features and the interface feature B3 is 60%, it can be determined that the feature similarity between the interface feature B2 is the largest (i.e. 80%), and therefore, the candidate application A2 corresponding to the interface feature B2 is determined as the target application to which the selected content belongs.
[0192] Based on this, in the video interaction page, the multimedia content is played, including: in the video interaction page, the program page of the target application is invoked; in the program page, the multimedia content is played.
[0193] Specifically, the login terminal of the interactive object invokes the program page of the target application in the video interaction page, and plays the multimedia content in the program page. That is, the program page of the target application is invoked in the video interaction page of the login terminal of the interactive object, and then the multimedia content is played in the program page. For ease of understanding, as shown in FIG. 21, the multimedia content 2104 is played in the program page 2102.
[0194] It can be understood that the terminal device plays multimedia content through the target application program, and needs to call the program page of the target application program for multimedia content playing, therefore, it needs to consider whether the target application program exists in the login terminal of the interactive object, and the following two cases of the target application program existing in the login terminal of the interactive object and the target application program not existing in the login terminal of the interactive object are introduced:
[0195] In one specific embodiment, in the video interactive page, the program page of the target application program is called, including: searching for the target application program in the existing application program; in the case of searching for the target application program, calling the program page of the target application program in the video interactive page.
[0196] Specifically, the terminal searches for the target application program in the downloaded existing application program, and the searching method of the target application program can be based on the program identifier of the target application program or based on the program name of the target application program, which is not limited here. Based on this, in the case of searching for the target application program in the existing application program of the terminal, the program page of the target application program is called in the video interactive page. The program page of the target application program can be displayed in a full-screen mode in the video interactive page, or can be displayed in a floating window mode in the video interactive page, which is not limited here. For ease of understanding, as shown in FIG. 22, the program page 2202 of the target application program is called in the video interactive page.
[0197] In one optional embodiment, in the case that the target application program does not exist in the existing application program, a target web page related to the target application program is called in the video interactive page, and the multimedia content is played in the target web page. Wherein, the target web page is a web page having the same or similar function as the target application program, and can be used to play the selected content (i.e. the multimedia content played by the terminal device). Specifically, in the case that the target application program does not exist in the existing application program, it means that the target application program cannot be called for content playing, at this time, a target web page related to the target application program can be called, and then the multimedia content is played in the target web page. The target web page can be displayed in a full-screen mode in the video interactive page, or can be displayed in a floating window mode in the video interactive page, which is not limited here.
[0198] For ease of understanding, as shown in FIG. 23, (A) of FIG. 23 illustrates that the target web page 2302 is called in the video interactive page, based on which, the multimedia content can be further played in the target web page 2302, that is, as shown in (B) of FIG. 23, the multimedia content 2304 is played in the target web page 2302.
[0199] In one optional embodiment, the method for content display further includes: in the case that the target application program does not exist in the existing application program, displaying a download prompt information of the target application program in the video interactive page.
[0200] Specifically, in the case that the target application does not exist in the existing application, the download prompt information of the target application is displayed on the video interaction page of the login terminal of the interactive object. Since the target application does not exist in the existing application, the video screen and the selected content displayed based on the webpage can be displayed simultaneously. Considering that the target webpage display exists webpage delay, and the related functions of the target application can not be fully supported, in order to ensure the display reliability of the selected content, the download prompt information of the target application can be displayed at this time, and the interactive object is instructed to download the target application through the download prompt information of the target application. That is, in the case that the target application does not exist in the login terminal of the interactive object, the download prompt information of the target application is displayed. For ease of understanding, as shown in FIG. 24, the download prompt information 2402 of the target application is displayed on the video interaction page.
[0201] Based on this, on the video interaction page, the program page of the target application is invoked, including: in response to the program download confirmation operation triggered for the download prompt information, in the case that the target application is downloaded successfully, the program page of the target application is displayed. Specifically, in the case that the interactive object determines to download the target application, that is, the interactive object triggers the program download confirmation operation for the download prompt information, at this time, the login terminal of the interactive object responds to the program download confirmation operation, and in the case that the target application is downloaded successfully, the program page of the target application is displayed.
[0202] Since in actual application, in the conference scene or part of the video interaction scene, it is essentially unnecessary to automatically push and display the content, the content can be displayed through the play confirmation information, and in one specific embodiment, on the video interaction page, the program page of the target application is invoked, including: on the video interaction page, the play confirmation information of playing multimedia content through the target application is displayed; in response to the confirmation operation triggered for the play confirmation information, the program page of the target application is invoked on the video interaction page.
[0203] The playing confirmation information is used to confirm whether the multimedia content is played through the target application, that is, the playing confirmation information is used to make the interactive object determine whether the multimedia content is played through the target application. The playing confirmation information can be displayed in a notification prompt area, which is used to display notification information for an application or for a system, and supports the minimization operation of the application. In this application, the notification prompt area is specifically used to display the playing confirmation information. Specifically, in the video interaction page of the login terminal of the interactive object, the playing confirmation information of playing the multimedia content through the target application is displayed, that is, during the video interaction process of the interactive object, the target application is not directly invoked, and the multimedia content is not played through the target application. Instead, the playing confirmation information of playing the multimedia content through the target application is displayed first, that is, the corresponding text of the playing confirmation information is displayed in the notification prompt area. The corresponding text of the playing confirmation information can be "multimedia content will be played for you" and the like, which is not limited here. For ease of understanding, as shown in FIG. 25, the playing confirmation information 2502 of playing the multimedia content through the target application is displayed on the login terminal of the interactive object, and the playing confirmation information 2502 is in the notification prompt area 2504.
[0204] Based on this, after the interactive object determines to play the multimedia content through the application, the interactive object can perform a confirmation operation on the playing confirmation information displayed on the login terminal. The confirmation operation refers to the operation performed by the interactive object on the playing confirmation information after the playing confirmation information of playing the multimedia content through the target application is displayed in the video interaction page, such as a long press operation, a click operation, a sliding operation and the like. The login terminal of the interactive object responds to the confirmation operation triggered on the playing confirmation information, invokes the program page of the target application in the video interaction page, and then plays the multimedia content through the program page of the target application. The way of invoking the program page and playing the multimedia content is similar to that introduced in the foregoing embodiments, which will not be described here.
[0205] Next, how to determine the playing confirmation information will be introduced. In one specific embodiment, the playing confirmation information of playing the multimedia content through the target application is displayed in the video interaction page, including: generating a deep link for indicating the target application and the multimedia content, and displaying the playing confirmation information of the multimedia content based on the deep link.
[0206] The playing confirmation information is used to confirm whether the multimedia content is played through the target application program. Since the selected content is the multimedia content played by the terminal device, that is, the multimedia content is necessarily displayed in the target application program in the terminal device, it is known from the foregoing embodiment that the target application program playing the multimedia content in the terminal device is determined when the multimedia content is determined. Specifically, a deep link indicating the target application program and the multimedia content needs to be generated by the server at this time. After the target application program playing the multimedia content in the terminal device is determined in the foregoing manner, the dynamic deep link is generated.
[0207] Specifically, the URI scheme is defined for each supported application program to realize in-application navigation, so that when the target application program is invoked through the deep link, the specific page where the multimedia content is located can be accurately located, the user experience and the accuracy of content display are improved, and the generation step is as follows: first, the URI scheme is defined for each supported application program to realize in-application navigation, and it is ensured that the target application program is an application program supporting the defined URI scheme. It is assumed that the target application program is A, the supported URI scheme is U, and the pre-defined URL structure is S=U+"?content_id={content_id}&content_name={content_name}". Then, a deep link generation request is initiated to the target application program, and the deep link generation request contains content parameters indicating the multimedia content, such as the ID of the multimedia content being id and the name being name. Then, the content parameters indicating the multimedia content are extracted from the deep link generation request, the content parameters id and name are filled into the URL structure S, that is, {content i d} is replaced by id, and {content n ame} is replaced by name, to construct a deep link L=U+"?content_id="+id+"&content_name="+name. Finally, the generated deep link L is returned to the login terminal of the interactive object. The deep link refers to a link service provided by a chaining website, so that the object can obtain the content on the linked website without leaving the chaining website page, and at this time, the page address bar displays the website address of the chaining website rather than the website address of the linked website. In the present application, the deep link is used to indicate the multimedia content in the target application program.
[0208] Therefore, after the login terminal of the interactive object obtains the deep link, the login terminal of the interactive object can display the play confirmation information for the multimedia content based on the deep link. That is, the login terminal of the interactive object displays a notification prompt area including the play confirmation information based on the deep link. At this time, the play confirmation information in the notification prompt area is triggered in the video interaction page. The deep link indicating the target application and the multimedia content is included in the play confirmation information. Therefore, the play confirmation information is triggered, the link indication of the deep link is performed, the program page of the target application is invoked, and the multimedia content is played in the program page.
[0209] Based on this, if the interactive object initiates a trigger operation for the play confirmation information, the program page of the target application is invoked in the video interaction page of the login terminal of the interactive object, and the multimedia content is played in the program page. Initiating the trigger operation for the play confirmation information is at least one of the following: the object line of sight of the interactive object focuses on the play confirmation information, and the duration of focusing on the play confirmation information reaches a preset focus duration; or, the object selection voice of the interactive object is for the play confirmation information; or, the information display duration of the play confirmation information reaches a preset display duration; or, the interactive object interacts with the play confirmation information. Therefore, how to determine the trigger operation for the play confirmation information will be introduced below. In an optional embodiment, the manner of determining the trigger operation for the play confirmation information includes: performing line-of-sight focus detection on the object line of sight of the interactive object to obtain a line-of-sight focus result; and triggering the trigger operation for the play confirmation information in a case where the line-of-sight focus result indicates that the object line of sight of the interactive object focuses on the play confirmation information and the duration of focusing on the play confirmation information reaches a preset focus duration.
[0210] Secondly, the line-of-sight focus detection is used to identify the focus area of the object line-of-sight. Therefore, the line-of-sight focus result is used to at least characterize the focus area of the object line-of-sight of the interactive object. Based on this, the preset focus duration can be 3 seconds, 5 seconds, 7 seconds, etc., and the specific preset focus duration needs to be flexibly determined based on the actual scene requirements. Specifically, in the case that the video picture collected by the login terminal of the interactive object in real time includes the interactive object, it indicates that the interactive object is in the video picture collected by the login terminal of the interactive object in real time. At this time, the line-of-sight focus detection can be further performed on the object line-of-sight of the interactive object to obtain the line-of-sight focus result, that is, to determine the focus area of the object line-of-sight of the interactive object. At this time, it is necessary to judge whether the focus area of the object line-of-sight is the playing confirmation information, that is, to judge whether the object line-of-sight of the interactive object is focused on the playing confirmation information. If yes, that is, the line-of-sight focus result characterizes that the object line-of-sight of the interactive object is focused on the playing confirmation information, then the duration of the interactive object to the playing confirmation information is timed, and in the case that the duration of focusing on the playing confirmation information reaches the preset focus duration, the trigger operation for the playing confirmation information is triggered, so as to arouse the program page of the target application program, and the multimedia content is played in the program page.
[0211] In an optional embodiment, the manner of determining the trigger operation for the playing confirmation information comprises: in the case that the object selection voice of the interactive object to the playing confirmation information is detected, triggering the trigger operation for the playing confirmation information.
[0212] Wherein, the object selection voice is the voice of the interactive object, which is used to trigger the playing confirmation information. The object selection voice can be "agree to display" and "can display content", etc., which is not limited here. Specifically, the login terminal of the interactive object can perform real-time voice detection. After the login terminal of the interactive object displays the playing confirmation information for indicating the selected content, if the object selection voice of the preset information (such as the aforementioned "agree to display" and "can display content", etc.) is detected, it is determined to trigger the trigger operation for the playing confirmation information, so as to arouse the program page of the target application program, and the multimedia content is played in the program page.
[0213] In an optional embodiment, the manner of determining the trigger operation for the playing confirmation information comprises: recording the information display duration of the playing confirmation information; in the case that the information display duration reaches the preset display duration, triggering the trigger operation for the playing confirmation information.
[0214] The preset display duration can be 3 seconds, 5 seconds, 7 seconds, etc. The specific preset display duration needs to be determined based on actual scene requirements. Specifically, after the login terminal of the interactive object displays the play confirmation information for indicating the selected content, the information display duration of the play confirmation information is recorded, and in the case that the information display duration of the play confirmation information reaches the preset display duration, the trigger operation for the play confirmation information is triggered, so as to arouse the program page of the target application program, and the multimedia content is played in the program page. In this case, when the login terminal of the interactive object displays the play confirmation information for indicating the selected content, the related notification of the preset display duration can be displayed at the same time, for example, when the preset display duration is 3 seconds, the related notification of the preset display duration is: "The selected content will be displayed for you in 3 seconds".
[0215] In an optional embodiment, the manner of determining the trigger operation for the play confirmation information includes: in the case that the interactive object interacts with the play confirmation information, triggering the trigger operation for the play confirmation information.
[0216] The interaction with the play confirmation information can include any of the following: a long press operation on the play confirmation information, a click operation on the play confirmation information, or a sliding operation on the play confirmation information, etc. Specifically, in the case that the login terminal of the interactive object detects that the interactive object interacts with the play confirmation information, the trigger operation for the play confirmation information can also be triggered, so as to arouse the program page of the target application program, and the multimedia content is played in the program page.
[0217] Further, in actual application, the multimedia content can also be displayed by a delivery device. Therefore, the delivery device that can deliver the content needs to be selected to display the selected content. The following describes this. In an optional embodiment, the method for displaying the content further includes: displaying, on the video interaction page, device identifiers of candidate delivery devices that can display the multimedia content; in response to a selection operation on the device identifier, selecting the candidate delivery device represented by the device identifier as a target delivery device; and delivering the multimedia content to the target delivery device.
[0218] The device identifier uniquely identifies the projection device, which is a terminal capable of projecting content, such as a projector or television. The device identifier can be the device name or device number, etc., without limitation. Specifically, the projection selection area displays the device identifiers of candidate projection devices capable of projecting multimedia content, and indicates that the candidate projection devices are in a communication connection with the login terminal of the interactive object. If no candidate projection devices capable of projecting multimedia content are available near the display terminal on the video interaction page, the multimedia content can be displayed directly on the login terminal's video interaction page. A candidate projection device refers to a terminal capable of projecting content, such as a projector or television, that is in a communication connection with the login terminal of the interactive object, and that supports multimedia content formats. If projection devices are available near the login terminal of the interactive object, a list of all projection devices is obtained through network scanning, etc., and their functions, status, and network connection stability with the login terminal are checked. Devices capable of projecting multimedia content are selected as candidate projection devices, and their device identifiers are displayed in the projection selection area of the video interaction page. The projection selection area is the area on the video interaction page used to display the device identifiers of candidate projection devices capable of projecting multimedia content. When there are candidate devices for displaying multimedia content near the login terminal of the interactive object, the device identifiers of these devices will be displayed in this area, making it convenient for the interactive object to select which device to display the multimedia content on.
[0219] Specifically, considering practical application scenarios, when there are delivery devices near the login terminal of the interactive object, candidate delivery devices for displaying multimedia content are first selected. The specific selection method is as follows: First, obtain a list L of all delivery devices near the login terminal of the interactive object through methods such as network scanning. Let the format of the multimedia content be F, and for each delivery device D in list L... i Check its functions and status to determine if it supports casting multimedia content in format F. If it does, mark it as S. i1 =1, otherwise S i1 =0. Simultaneously, the delivery device D is detected by sending test data packets. i The stability of the network connection between the login terminal and the interactive object is marked as S if the connection is normal. i2 =1, otherwise S i2 =0. Filter out S i1 =1 and S i2 Devices with a value of 1 are considered as candidate devices for displaying multimedia content.
[0220] Specifically, considering the actual application scenario, in the case where there is a delivery device near the login terminal of the interactive object, first, the candidate delivery device that can deliver the display multimedia content is screened out. The specific screening method is as follows:
[0221] First, network scanning is performed to obtain a list of all delivery devices near the login terminal of the interactive object. A network discovery protocol such as SSDP (Simple Service Discovery Protocol) or mDNS (Multicast Domain Name System) can be used. Taking SSDP as an example, the login terminal broadcasts an SSDP discovery message to the local network, and the delivery device supporting SSDP will respond to the message. The login terminal collects the information of the delivery device, such as the IP address of the device, the device name, etc., according to the response message, thereby forming a list of delivery devices.
[0222] Suppose the format of the multimedia content is F, for each delivery device D i in the list, check its function and state, determine whether it supports the delivery of multimedia content with format F, if it supports, mark S i =1, otherwise S i =0.
[0223] At the same time, the network connection stability between the delivery device D i and the login terminal of the interactive object is detected by sending test data packets. The specific method is that the login terminal sends a certain number (such as n) of test data packets to the delivery device, records the sending time t send and the time t recv when the response data packet is received, calculates the round-trip time RTT=t recv -t send for each data packet. The average round-trip time RTT of all test data packets is calculated, and the standard deviation σ of the round-trip time is calculated. If the average round-trip time RTT avg is less than a preset time threshold T th , and the standard deviation σ RTT is less than a preset fluctuation threshold σ th , it is considered that the connection is normal, marked as C i =1, otherwise C i =0.
[0224] The delivery device D i with S i =1 and C i =1 is screened out as a candidate delivery device that can deliver the display multimedia content. Wherein, F represents the format of the multimedia content, D i represents the i-th delivery device in the list, S i represents the i-th delivery device whether it supports the delivery of multimedia content with this format, C ia marker representing the network connection stability between the ith delivery device and the interactive object login terminal, n represents the number of test data packets sent, RTT i a marker representing the round-trip time of the ith test data packet, RTT avg a marker representing the average round-trip time, σ RTT a marker representing the standard deviation of the round-trip time, T th a marker representing the preset time threshold, σ th a marker representing the preset fluctuation threshold.
[0225] Then, the device identifiers of the candidate delivery devices that can deliver the display multimedia content are displayed in the delivery selection area in the video interaction page, that is, at this time, the multimedia content is not directly displayed, but the device identifiers of the candidate delivery devices that can deliver the display multimedia content are displayed, so that the interactive object can choose to play the multimedia content through the delivery device or in the video interaction page of the login terminal. For ease of understanding, as shown in FIG. 26, the delivery selection area 2602 is displayed in the video interaction page of the login terminal of the interactive object, and the device identifiers 2604 of the candidate delivery devices that can deliver the display multimedia content are displayed in the delivery selection area 2602.
[0226] Based on this, if the interactive object wants to play the multimedia content through the delivery device, the interactive object can perform a selection operation on the displayed device identifier, and the aforementioned selection operation can be a long press operation on the device identifier, a click operation on the device identifier, etc., which is not limited here. Thus, the login terminal of the interactive object can respond to the selection operation on the device identifier, and take the candidate delivery device represented by the selected device identifier as the target delivery device, so as to deliver the multimedia content to the target delivery device, so that the target delivery device plays the multimedia content. For example, the device identifier C1, the device identifier C2 and the device identifier C3 are displayed, the device identifier C1 represents the delivery device D1, the device identifier C2 represents the delivery device D2, and the device identifier C3 represents the delivery device D3, if the interactive object performs a selection operation on the device identifier C2, then the delivery device D2 represented by the device identifier C2 is the target delivery device, at this time, the multimedia content is delivered to the delivery device D2, and the multimedia content is played through the delivery device D2. The target delivery device refers to that after the device identifiers of the candidate delivery devices that can deliver the display multimedia content are displayed in the video interaction page, the interactive object performs a selection operation on the device identifier, and the candidate delivery device represented by the selected device identifier is the target delivery device, and the multimedia content will be delivered to the device for display.
[0227] The following will be described in detail how to screen the candidate delivery device and the way of displaying the terminal identifier: in a specific embodiment, on the video interaction page, the device identifier of the candidate delivery device capable of delivering the display multimedia content is displayed, comprising: searching for the available delivery device in connection with the display terminal of the video interaction page; taking the searched available delivery device as the candidate delivery device for delivering the multimedia content; and displaying the device identifier of the candidate delivery device on the video interaction page.
[0228] Specifically, after determining the multimedia content and the target application program, the login terminal of the interactive object, i.e., the display terminal of the video interaction page, can detect the available devices in the vicinity of the login terminal of the interactive object, and the available devices in the vicinity of the display terminal of the video interaction page can be the devices with a signal strength greater than a signal strength threshold between the display terminal of the video interaction page. Based on this, to ensure normal communication connection, the display terminal of the video interaction page also needs to detect the connection state between each available device and the display terminal of the video interaction page, and determine the available devices in connection with the display terminal of the video interaction page as the available delivery devices, which can be single or multiple.
[0229] Based on this, the searched available delivery device is taken as the candidate delivery device for delivering the multimedia content, and then the device identifier of the candidate delivery device is displayed on the video interaction page. The specific manner is similar to the foregoing embodiment, which will not be described here. It can be understood that if there is no available device in the vicinity of the display terminal of the video interaction page, or the connection state between the available devices and the display terminal of the video interaction page is all not connected, the display manner introduced in the foregoing embodiment is used for content display at this time.
[0230] Secondly, since the present application considers the target application program, if the interactive object wants to directly jump to the target application program for playing the multimedia content, the interactive object can directly interact with the multimedia content, directly jump to the target application program, and display and play the multimedia content through the target application program, at this time the video picture can be displayed through the floating window. Alternatively, the interactive object can interact with the target application program, at this time the target application program is displayed and jumped, and the video picture can also be displayed on the page for displaying and playing the multimedia content through the target application program in the form of a floating window. The foregoing interaction operation can be single-click operation, double-click operation, sliding operation, etc., which is not limited here.
[0231] For ease of understanding, taking displaying a video picture in the login terminal of the first interactive object as an example, as shown in FIG. 27, (A) of FIG. 27 shows that the video picture 2702 and the multimedia content 2704 are displayed in the login terminal of the first interactive object. If the login terminal of the first interactive object performs an interactive operation on the multimedia content 2704, (B) of FIG. 27 shows that the display mode of displaying the multimedia content 2704 through a target application program, and the video picture 2702 is displayed on the multimedia content 2704 through the target application program in the form of a floating window.
[0232] In addition, in the present embodiment, the selected content is mainly multimedia content. If the selected content is entity content, the content picture associated with the selected content is different, but the display mode of displaying in the form of full screen / floating window is similar to the display mode described in the foregoing embodiments, which will not be described herein again.
[0233] In the foregoing embodiments, in the case that the terminal device playing the multimedia content is displayed in the video picture, the multimedia content can be identified, and the target application program playing the multimedia content can be determined, so as to consider the multimedia content and the target application program, display the multimedia content in the login terminal of the interactive object, and ensure reliable content display in various actual scenarios, thereby improving the reliability and flexibility of content display.
[0234] Based on the detailed description of the foregoing embodiments, the complete flow of the method for content display in the present embodiment will be described. In one embodiment, as shown in FIG. 28, a method for content display is provided. Taking the case that the method is applied to the server 104 in FIG. 1 as an example, it can be understood that the method can also be applied to the terminal 102, and can also be applied to a system including the terminal 102 and the server 104, and is realized through the interaction of the terminal 102 and the server 104. In the present embodiment, the method includes the following steps:
[0235] In step 2801, in a video interaction scene in which at least two interactive objects participate, a video picture collected in real time by a login terminal of at least one interactive object is displayed in a video interaction page of another interactive object. The interactive objects participating in the video interaction include at least a first interactive object and a second interactive object. The video interaction page is a display page of the first interactive object. The terminal device playing the multimedia content is displayed in the video picture.
[0236] In step 2802, in response to an environmental content attention operation triggered on one of the video pictures, the video picture triggered by the environmental content attention operation is taken as a target video picture. The target video picture is a video picture corresponding to the second interactive object.
[0237] Step 2803, in a case where the voice interaction content of the first interaction object and the second interaction object contains a preset keyword, triggering the environmental content attention event.
[0238] Step 2804, in a case where the preset gesture operation is triggered for the target video screen, triggering the environmental content attention event.
[0239] Step 2805, in a case where the interaction operation is triggered for the target video screen, triggering the environmental content attention event.
[0240] Step 2806, screening the target video frame from the video stream data containing the target video screen.
[0241] Step 2807, performing environmental content recognition on the target video frame, taking the environmental content represented by the content recognition result as the selected content; the selected content is the multimedia content played by the terminal device.
[0242] Step 2808, determining the candidate application program matching the category of the multimedia content; extracting the content feature of the multimedia content; performing similarity analysis on the content feature and the interface feature corresponding to each candidate application program respectively to obtain the feature similarity; taking the candidate application program whose feature similarity meets the matching requirement as the target application program.
[0243] Specifically, let the content feature vector be V c , the interface feature vector corresponding to the candidate application program A i be V i , and the Euclidean distance be used to calculate the distance between them, where n is the dimension of the feature vector, and are the jth components of the content feature vector and the interface feature vector respectively. The smaller the distance, the higher the similarity. Take the candidate application program whose distance is less than a preset threshold T as the target application program.
[0244] Step 2809, searching for the target application program in the existing application programs.
[0245] Step 2810, in a case where the target application program is found, displaying, on the video interaction page, the play confirmation information of playing the multimedia content through the target application program.
[0246] Step 2811, in response to the confirmation operation triggered for the play confirmation information, invoking the program page of the target application program on the video interaction page; playing the multimedia content on the program page.
[0247] Step 2812, in the case that the target application does not exist in the existing application program, displaying a download prompt information of the target application on the video interaction page.
[0248] Step 2813, in response to a program download confirmation operation triggered for the download prompt information, displaying a program page of the target application in the case that the target application is downloaded successfully; playing the multimedia content on the program page.
[0249] Step 2814, displaying device identifiers of candidate display devices on which the multimedia content can be displayed on the video interaction page; in response to a selection operation for the device identifiers, taking a candidate display device represented by a selected device identifier as a target display device; and displaying the multimedia content to the target display device.
[0250] It should be understood that the specific implementation of steps 2801 to 2814 is similar to the foregoing embodiments, and will not be described here.
[0251] It should be understood that, although the steps in the flowcharts involved in the above-described embodiments are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least some of other steps or stages or steps in other steps.
[0252] Based on the same inventive concept, the embodiments of the present application also provide a content display device for implementing the above-described content display method. The implementation scheme of the problem solving provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more content display device embodiments provided below can refer to the limitations of the content display method in the foregoing, which will not be described here.
[0253] In one embodiment, as shown in FIG. 29, a content display device is provided, comprising a video picture display module 2902, an event triggering module 2904 and a content picture display module 2906, wherein:
[0254] The video picture display module 2902 is configured to display a video picture collected in real time by a login terminal of at least one other interactive object in a video interaction page of one interactive object in a video interaction scene in which the one interactive object and the at least one other interactive object participate.
[0255] The event triggering module 2904 is configured to, in response to an environmental content attention event triggered for environmental content in any one video picture, take the environmental content matched with the environmental content attention event as selected content;
[0256] The content picture display module 2906 is configured to display a content picture associated with the selected content on a video interaction page.
[0257] In an embodiment, the terminal device displays a multimedia content playing terminal device in the video picture, and the selected content is the multimedia content played by the terminal device.
[0258] The content picture display module is specifically configured to play the multimedia content on the video interaction page.
[0259] In an embodiment, the terminal device plays the multimedia content through a target application program.
[0260] The content picture display module is specifically configured to arouse a program page of the target application program on the video interaction page, and play the multimedia content on the program page.
[0261] In an embodiment, the content picture display module is specifically configured to search for the target application program in an existing application program, and arouse a program page of the target application program on the video interaction page in a case where the target application program is searched.
[0262] In an embodiment, the content display device further includes an information display module.
[0263] The information display module is configured to display download prompt information of the target application program on the video interaction page in a case where the target application program does not exist in the existing application program.
[0264] The content picture display module is specifically configured to, in response to a program download confirmation operation triggered for the download prompt information, display a program page of the target application program in a case where the target application program is successfully downloaded.
[0265] In an embodiment, the content picture display module is specifically configured to display playing confirmation information of the multimedia content played through the target application program on the video interaction page, and arouse a program page of the target application program on the video interaction page in response to a confirmation operation triggered for the playing confirmation information.
[0266] In an embodiment, the content picture display module is specifically configured to generate a deep link for indicating the target application program and the multimedia content, and display playing confirmation information of the multimedia content based on the deep link, the playing confirmation information being used to confirm whether to play the multimedia content through the target application program.
[0267] In an embodiment, the content presentation device further comprises an application determining module;
[0268] The application determining module is configured to determine candidate applications matching the category of the multimedia content; extract content features of the multimedia content; perform similarity analysis on the content features and interface features corresponding to each candidate application respectively to obtain feature similarity; and select a candidate application whose feature similarity meets a matching requirement as a target application.
[0269] In an embodiment, the content presentation device further comprises a content delivery module;
[0270] The content delivery module is configured to display device identifiers of candidate delivery devices capable of delivering the presentation multimedia content on a video interaction page; in response to a selection operation on a device identifier, select a candidate delivery device represented by the selected device identifier as a target delivery device; and deliver the multimedia content to the target delivery device.
[0271] In an embodiment, the content delivery module is specifically configured to search for available delivery devices in a connected state with a display terminal of the video interaction page; select the searched available delivery devices as candidate delivery devices for delivering the multimedia content; and display device identifiers of the candidate delivery devices on the video interaction page.
[0272] In an embodiment, the event triggering module is specifically configured to, in response to an environmental content attention event triggered by an environmental content in any one video frame, determine a target video frame indicated by the environmental content attention event; filter out a target video frame from video stream data containing the target video frame; and perform environmental content recognition on the target video frame, and select environmental content represented by a content recognition result as selected content.
[0273] In an embodiment, the event triggering module is specifically configured to, in a case where the interactive object participating in the video interaction is at least three, in response to an environmental content attention operation triggered on one of the video frames, select a video frame triggered by the environmental content attention operation as a target video frame; and trigger an environmental content attention event for environmental content in the target video frame.
[0274] In an embodiment, the interactive object participating in the video interaction includes at least a first interactive object and a second interactive object.
[0275] The event triggering module is specifically configured to, in a case where the video interaction page is a display page of the first interactive object, and a video frame triggered by the environmental content attention event is a target video frame corresponding to the second interactive object, trigger the environmental content attention event in a case where voice interaction content of the first interactive object and the second interactive object contains a preset keyword.
[0276] In an embodiment, the interactive objects participating in the video interaction include at least a first interactive object and a second interactive object.
[0277] The event triggering module is specifically configured to, in a case where the video interaction page is a display page of the first interactive object, a video picture triggered with the environmental content attention event is a target video picture corresponding to the second interactive object, and a preset gesture operation is triggered for the target video picture, trigger the environmental content attention event.
[0278] In an embodiment, the interactive objects participating in the video interaction include at least a first interactive object and a second interactive object.
[0279] The event triggering module is specifically configured to, in a case where the video interaction page is a display page of the first interactive object, a video picture triggered with the environmental content attention event is a target video picture corresponding to the second interactive object, and an interactive operation is triggered for the target video picture, trigger the environmental content attention event.
[0280] In an embodiment, the content picture display module is specifically configured to display the selected content in a preset angle on the video interaction page.
[0281] In an embodiment, the content picture display module is specifically configured to display a global image of the selected content on the video interaction page.
[0282] In an embodiment, the content picture display module is specifically configured to display a content promotion page of the selected content on the video interaction page.
[0283] In an embodiment, the content picture display module is specifically configured to display a content detail page of the selected content on the video interaction page.
[0284] In an embodiment, the content picture display module is specifically configured to display a content picture associated with the selected content in a full-screen mode, and display the video picture on the content picture associated with the selected content in a floating window mode.
[0285] In an embodiment, the content picture display module is specifically configured to display the video picture in a full-screen mode, and display the content picture associated with the selected content on the video picture in a floating window mode.
[0286] In an embodiment, the content picture display module is specifically configured to display the video picture through a first floating window, and display the content picture associated with the selected content through a second floating window.
[0287] In an embodiment, the computer device can be a server or a terminal. In the embodiment, the computer device is taken as a terminal as an example for introduction, and its internal structure diagram can be shown in FIG. 30. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to perform wired or wireless communication with external terminals. The wireless communication can be achieved through WIFI, mobile cellular network, NFC (Near Field Communication) or other technologies. The computer program is executed by the processor to implement a content display method. The display unit of the computer device is configured to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, a trackball or a touchpad arranged on the shell of the computer device, or an external keyboard, a touchpad or a mouse, etc.
[0288] Those skilled in the art can understand that the structure shown in FIG. 30 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. Specifically, the computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0289] In an embodiment, a computer device is also provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0290] In an embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0291] In an embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0292] To sum up, the present application provides a content display method, device, computer equipment, computer readable storage medium and computer program product. In a video interaction scene involving at least two interaction objects, the computer equipment displays the video screen collected by the terminal of the rest of the at least one interaction object on the video interaction page of one of the interaction objects. From a technical point of view, this process uses network communication technology and video stream analysis technology to ensure the real-time and smoothness of the video screen. The video stream data collected by each interaction object terminal is transmitted to the designated terminal through the network, and then the video stream data is analyzed to obtain the video screen. Real-time display of the video screen ensures the continuity and authenticity of the video interaction, allowing the interaction objects to communicate as if they were in the same space, improving the interaction experience and verification convenience. When an environmental content attention event triggered by the environmental content in any one video screen is detected, the environmental content matched with the event is selected as the selected content, and the content screen associated with the selected content is displayed on the video interaction page. This process involves image recognition and content matching technology. The computer equipment identifies the environmental content in the video screen through an image recognition model, determines the selected content, and then performs content matching according to the type and characteristics of the selected content to find the associated content screen. This automatic identification and display method avoids the tedious operation of manual content sharing by the interaction objects, reduces the operation time and errors, improves the efficiency of content display, and also ensures the continuity of video interaction, which will not be interrupted due to content sharing operation.
[0293] Further, when a terminal device playing multimedia content is displayed in the video screen, the selected content is the multimedia content played by the terminal device, and the computer equipment plays the multimedia content on the video interaction page. Technically, this requires format recognition and playback adaptation of multimedia content. The computer equipment analyzes the format of the multimedia content, calls the corresponding playback module for playback, and ensures that the multimedia content can be normally displayed on the video interaction page. This targeted content playback method meets the user's demand for instant viewing of multimedia content during video interaction, improves the practicality and user experience of content display.
[0294] Further, if the terminal device plays multimedia content through a target application program, the computer equipment invokes the program page of the target application program on the video interaction page and plays the multimedia content on the program page. This process involves application program invocation and content delivery technology. The computer equipment invokes the target application program through a system interface and delivers the relevant information of the multimedia content to the application program, so that it can accurately play the multimedia content. By using the professional playback function of the target application program, more stable and high-quality playback effect can be provided, and other functions of the application program, such as playback control and content recommendation, can also be used to improve the quality and user experience of content playback.
[0295] Further, when the program page of the target application is invoked on the video interaction page, the computer device first searches for the target application in the existing applications, and if found, invokes the program page thereof. This process utilizes application management technology, in which the computer device searches for the target application according to its identifier or name by traversing the list of applications in the device. This approach quickly locates the target application, avoids blind search among a large number of applications, improves the efficiency of application invocation, and reduces user waiting time.
[0296] Further, when the target application does not exist in the existing applications, the computer device displays download prompt information of the target application on the video interaction page. If the user triggers a program download confirmation operation in response to the download prompt information, the program page of the target application is displayed when the target application is successfully downloaded. This involves application download and installation management technology, in which the computer device connects to an application download source through a network, downloads and installs the target application. Displaying the download prompt information provides the user with a way to obtain the target application, expanding the scope of content display. Meanwhile, the automatic download and installation function reduces the steps of manual operation by the user, improving the convenience of application acquisition.
[0297] Further, when the program page of the target application is invoked on the video interaction page, the computer device first displays play confirmation information of playing the multimedia content through the target application, and then invokes the program page of the target application on the video interaction page in response to a confirmation operation triggered in response to the play confirmation information. This process utilizes user interaction and information prompting technology, which allows the user to decide whether to play the multimedia content by displaying the play confirmation information, increasing the user's control over content playback. In some scenarios that are not suitable for automatic playback, such as meetings, quiet environments, etc., this approach can avoid unnecessary interference, improving the comfort and flexibility of the user experience.
[0298] Further, when the play confirmation information of playing the multimedia content through the target application is displayed, the computer generates a deep link for indicating the target application and the multimedia content, and displays the play confirmation information of the multimedia content based on the deep link. Deep link technology can accurately locate the multimedia content in the target application, and by generating a link containing multimedia content information, the user can directly jump to the corresponding page of the target application to play the multimedia content when confirming playback. This improves the accuracy and efficiency of content display, reduces the time for the user to search for content in the application, and enhances the user experience.
[0299] Further, the computer device determines a candidate application matching the category of the multimedia content, extracts a content feature of the multimedia content, performs similarity analysis on the content feature and an interface feature corresponding to each candidate application respectively, obtains a feature similarity, and takes the candidate application satisfying a matching requirement as a target application. This process involves feature extraction, similarity calculation and application screening techniques. The computer device extracts the feature of the multimedia content through an image recognition model, and then compares the feature with interface features in an application feature library to calculate a similarity. This application screening method based on feature matching can accurately find a target application suitable for playing the multimedia content, improves the accuracy and reliability of application recognition, and ensures that the multimedia content can be well displayed in a suitable application.
[0300] Further, the computer device displays device identifiers of candidate display devices that can display the multimedia content on the video interaction page, and in response to a selection operation on a device identifier, takes the candidate display device represented by the selected device identifier as a target display device, and displays the multimedia content on the target display device. This involves device discovery, selection and content display techniques. The computer device discovers candidate display devices that can display the multimedia content through network scanning, etc., and displays their device identifiers for user selection. After the user selects the target display device, the computer device transmits the multimedia content to the target display device through the network for display. This multi-device display method provides users with more content display options, expands the range and effect of content display, and meets the display requirements in different scenarios.
[0301] Further, when displaying the device identifiers of the candidate display devices that can display the multimedia content on the video interaction page, the computer searches for available display devices connected to the display terminal of the video interaction page, takes them as candidate display devices for displaying the multimedia content, and displays their device identifiers. This process uses device connection detection and screening techniques. The computer device detects devices connected to the display terminal through network communication, and screens out devices that can display the multimedia content. Displaying the identifiers of available devices can help users clearly understand the available display devices, improve the accuracy and reliability of device selection, and avoid content display failures caused by device connection problems.
[0302] Further, in response to an environmental content attention event triggered for environmental content in any one video picture, the computer device first determines a target video picture indicated by the event, screens a target video frame from the video stream data containing the target video picture, performs environmental content identification on the target video frame, and takes the environmental content represented by the content identification result as the selected content. In terms of technology, the screening of the target video frame utilizes key frame extraction and image preprocessing technologies, which improves image quality by performing image preprocessing on consecutive video frames in the video stream data, and then screens a video frame with large changes as the target video frame, thereby reducing the computational burden. The environmental content identification on the target video frame utilizes an image recognition model, such as a convolutional neural network, to extract features of the video frame and determine the environmental content. This way avoids screening and content detection on full-quantity video frames, improves the efficiency and accuracy of content detection, and thus more accurately and efficiently determines the selected content, thereby improving the reliability and efficiency of displaying the content picture associated with the selected content.
[0303] Further, in the case where the interactive object participating in the video interaction includes at least three interactive objects, if an environmental content attention operation is triggered for one of the video pictures, the computer device takes the video picture triggered by the operation as a target video picture, and triggers an environmental content attention event for the environmental content in the target video picture. This process utilizes user interaction detection and video picture positioning technologies, and the computer device determines the target video picture by detecting user operations such as long pressing, gesture operations, etc. This way can accurately determine the target video picture in multiple video pictures, improve the accuracy and pertinence of the triggering of the environmental content attention event, and ensure content detection and display for the video picture that the user is interested in.
[0304] Further, when the interactive object participating in the video interaction includes at least a first interactive object and a second interactive object, and the video interaction page is a display page of the first interactive object, and the video picture triggered by the environmental content attention event is a target video picture corresponding to the second interactive object, the triggering mode of the environmental content attention event includes: triggering the event in the case where the voice interaction content of the first interactive object and the second interactive object contains a preset keyword; triggering the event in the case where a preset gesture operation is triggered for the target video picture; and triggering the event in the case where an interactive operation is triggered for the target video picture. These triggering modes utilize voice recognition, gesture recognition, and interaction detection technologies. The voice recognition technology can recognize the preset keyword in the voice interaction content, the gesture recognition technology can recognize the preset gesture operation, and the interaction detection technology can detect the user's interactive operation. The diversified triggering modes provide the user with more interactive options, increase the flexibility and convenience of the interaction, and enable the user to trigger the environmental content attention event according to his own habits and needs.
[0305] Further, when displaying the content screen associated with the selected content on the video interaction page, at least one of the following modes is included: displaying the selected content at a preset angle, displaying a global image of the selected content, displaying a content promotion page of the selected content, and displaying a content detail page of the selected content. These display modes utilize image display and page rendering technologies, and select appropriate display modes according to the type and characteristics of the selected content. Displaying the selected content at a preset angle can highlight the local features of the content, displaying a global image can allow the user to understand the content comprehensively, and displaying a content promotion page and a content detail page can provide the user with more content information. The rich display modes meet the user's demand for display of different types of content, and improve the diversity and attractiveness of content display.
[0306] Further, when displaying the content screen associated with the selected content on the video interaction page, at least one of the following modes is included: displaying the selected content at a preset angle, displaying a global image of the selected content, displaying a content promotion page of the selected content, and displaying a content detail page of the selected content. These display modes utilize image display and page rendering technologies, and select appropriate display modes according to the type and characteristics of the selected content. Displaying the selected content at a preset angle can highlight the local features of the content, displaying a global image can allow the user to understand the content comprehensively, and displaying a content promotion page and a content detail page can provide the user with more content information. The rich display modes meet the user's demand for display of different types of content, and improve the diversity and attractiveness of content display.
[0307] Further, in the case that the video screen of the first interaction object is collected in real time by the login terminal of the first interaction object, the computer device performs line-of-sight focus detection on the object line-of-sight of the first interaction object to obtain a line-of-sight focus result. If the line-of-sight focus result indicates that the object line-of-sight of the first interaction object is focused on the environmental content in the video screen collected in real time by the login terminal of the second interaction object, and the focusing time of the environmental content reaches a preset focusing time, an environmental content attention event is triggered. This process utilizes line-of-sight focus detection technology, captures the eye image of the first interaction object through a camera, analyzes the rotation angle and direction of the eyeball, and determines the line-of-sight focus area. The triggering mode based on line-of-sight focus detection can automatically identify the user's interest point without additional user operation, improving the automation degree and accuracy of the triggering of the environmental content attention event, and providing the user with a more intelligent and convenient interactive experience.
[0308] In practical applications, when key frame extraction is performed on video stream data, image preprocessing is first performed on continuous video frames in the video stream data, and the image quality of each video frame is improved through an image enhancement algorithm such as histogram equalization. Histogram equalization can improve the uneven lighting problem that may exist in the video frame, make the gray scale distribution of the image more uniform, enhance the global contrast of the image, and improve the clarity and recognizability of the image. This helps subsequent video frame selection and content recognition operations, reduces recognition errors caused by image quality problems, and improves the accuracy and efficiency of content recognition.
[0309] When the target video frame is subjected to environmental content recognition, the extracted target video frame is scaled to a size supported by the image recognition model, such as 224x224 pixels, and then the size-scaled target video frame is subjected to grayscale processing to reduce the computational complexity of the image recognition model. The pixel values of the target video frame are also normalized to between 0 and 1 to improve stability. These preprocessing operations can optimize the input format of the target video frame, allowing the image recognition model to more efficiently perform feature extraction and environmental content recognition, reducing the consumption of computing resources and improving the running efficiency and recognition accuracy of the model. When building an application feature library, interface samples of a large number of different applications are collected, and interface features thereof are extracted, including the color scheme of the application, the shape and position of the button, the font and size of the text, the UI component layout style, the application logo, and specific interface elements, etc. These features are quantized and coded to form feature vectors, which are stored in classified application types. The construction of the application feature library provides rich reference information for determining the target application. By comparing the content features with the interface features of different application types in the application feature library, the selected content can be more accurately determined to belong to the target application, improving the accuracy and reliability of application recognition.
[0310] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the object or fully authorized by all parties, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0311] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, etc., without being limited thereto.
[0312] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0313] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the scope of protection of the patent of the present application should be subject to the appended claims.
Claims
1. A method of content presentation, performed by a computer device, the method comprising: in a video interaction scenario involving at least two interaction objects, displaying, in a video interaction page of one of the interaction objects, a video screen captured in real time by a login terminal of the rest of the at least one interaction object; in response to an environmental content focus event triggered by environmental content in any one of the video screens, taking the environmental content matched with the environmental content focus event as selected content; and displaying, in the video interaction page, a content screen associated with the selected content. 2.The method of claim 1, wherein, in a case where a terminal device playing multimedia content is displayed in the video screen, the selected content is the multimedia content played by the terminal device; and the displaying, in the video interaction page, the content screen associated with the selected content comprises: playing the multimedia content in the video interaction page. 3.The method of claim 2, wherein the terminal device plays the multimedia content through a target application; and the playing the multimedia content in the video interaction page comprises: invoking a program page of the target application in the video interaction page; and playing the multimedia content in the program page. 4.The method of claim 3, wherein the invoking the program page of the target application in the video interaction page comprises: searching for the target application in existing applications; and in a case where the target application is found, invoking the program page of the target application in the video interaction page. 5.The method of claim 3 or 4, further comprising: in a case where the target application does not exist in the existing applications, displaying, in the video interaction page, download prompt information of the target application; and the invoking the program page of the target application in the video interaction page comprises: in response to a program download confirmation operation triggered by the download prompt information, displaying the program page of the target application in a case where the target application is successfully downloaded. 6.The method of any one of claims 3 to 5, wherein the invoking the program page of the target application in the video interaction page comprises: displaying, in the video interaction page, play confirmation information of playing the multimedia content through the target application; and in response to a confirmation operation triggered by the play confirmation information, invoking the program page of the target application in the video interaction page. 7.The method of claim 6, wherein the displaying, in the video interaction page, the play confirmation information of playing the multimedia content through the target application comprises: generating a deep link for indicating the target application and the multimedia content, and displaying the play confirmation information of the multimedia content based on the deep link, the play confirmation information being used to confirm whether to play the multimedia content through the target application. 8.The method of any one of claims 3 to 7, further comprising: determining a candidate application matching the category of the multimedia content; extracting a content feature of the multimedia content; performing similarity analysis on the content feature and an interface feature corresponding to each of the candidate applications, to obtain a feature similarity; selecting a candidate application with a feature similarity satisfying a matching requirement as a target application.
9. The method of any one of claims 2-8, further comprising: displaying, on the video interaction page, device identifiers of candidate display devices on which the multimedia content can be displayed; in response to a selection operation on a device identifier, selecting a candidate display device represented by the selected device identifier as a target display device; and displaying the multimedia content on the target display device.
10. The method of claim 9, wherein the displaying, on the video interaction page, device identifiers of candidate display devices on which the multimedia content can be displayed comprises: searching for available display devices connected to a display terminal of the video interaction page; selecting the searched available display devices as candidate display devices on which the multimedia content can be displayed; and displaying, on the video interaction page, device identifiers of the candidate display devices.
11. The method of any one of claims 1-10, wherein the selecting, in response to an environmental content attention event triggered by environmental content in any one of the video frames, environmental content corresponding to the environmental content attention event as selected content comprises: in response to an environmental content attention event triggered by environmental content in any one of the video frames, determining a target video frame indicated by the environmental content attention event; screening a target video frame from video stream data including the target video frame; and performing environmental content recognition on the target video frame, and selecting environmental content represented by a content recognition result as selected content.
12. The method of any one of claims 1-11, wherein, when there are at least three interactive objects participating in the video interaction, the triggering manner of the environmental content attention event triggered by environmental content in any one of the video frames comprises: in response to an environmental content attention operation triggered by a video frame, selecting the video frame as a target video frame; and triggering an environmental content attention event for environmental content in the target video frame.
13. The method of any one of claims 1-12, wherein the interactive objects participating in the video interaction include at least a first interactive object and a second interactive object; when the video interaction page is a display page of the first interactive object, and a video frame in which the environmental content attention event is triggered is a target video frame corresponding to the second interactive object, the triggering manner of the environmental content attention event comprises any one of the following manners: triggering the environmental content attention event when voice interaction content of the first interactive object and the second interactive object includes a preset keyword; and triggering the environmental content attention event when a preset gesture operation is triggered for the target video frame. In a case that the interaction operation is triggered for the target video picture, an environmental content attention event is triggered.
14. The method of any one of claims 1-13, wherein the displaying, on the video interaction page, the content picture associated with the selected content comprises at least one of the following: displaying the selected content in a preset angle on the video interaction page; displaying a global image of the selected content on the video interaction page; displaying a content promotion page of the selected content on the video interaction page; and displaying a content detail page of the selected content on the video interaction page.
15. The method of any one of claims 1-14, wherein the displaying, on the video interaction page, the content picture associated with the selected content comprises any one of the following: displaying the content picture associated with the selected content in a full screen mode and displaying the video picture on the content picture associated with the selected content in a floating window mode; displaying the video picture in a full screen mode and displaying the content picture associated with the selected content on the video picture in a floating window mode; and displaying the video picture in a first floating window and displaying the content picture associated with the selected content in a second floating window.
16. The method of any one of claims 1-15, further comprising: performing a line-of-sight focus detection on a line-of-sight of the first interaction object to obtain a line-of-sight focus result in a case that the first interaction object is included in the video picture collected by the login terminal of the first interaction object in real time; and triggering an environmental content attention event in a case that the line-of-sight focus result indicates that the line-of-sight of the first interaction object is focused on the environmental content in the video picture collected by the login terminal of the second interaction object in real time and a time length of focusing on the environmental content reaches a preset focus time length.
17. A content display apparatus, comprising: a video picture display module configured to display, in a video interaction page of one of at least two interaction objects, a video picture collected by a login terminal of at least one of the rest of the at least two interaction objects in a video interaction scenario in which the at least two interaction objects participate; an event trigger module configured to, in response to an environmental content attention event triggered for an environmental content in any one of the video pictures, match the environmental content attention event to the environmental content as selected content; and a content picture display module configured to display, on the video interaction page, a content picture associated with the selected content.
18. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements steps of the method in any one of claims 1-16 when executing the computer program.
19. A computer readable storage medium, having a computer program stored thereon, wherein the computer program, when executed by a processor, implements steps of the method in any one of claims 1-16.
20. A computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements steps of the method in any one of claims 1-16.
Citation Information
Patent Citations
Method, device and system for sharing video picture in video call
CN104780336A
Video playing method and device, electronic equipment and storage medium
CN111274449A
Information acquisition method and device and electronic equipment
CN111601066A
Video sharing method and device, equipment and medium
CN113489937A
Displaying additional external video content along with users participating in a video exchange session
US11082464B1