Information display method and device, computer readable storage medium, and electronic device
By acquiring the first time point of video content, extracting target video content, processing target objects and their actions, and generating mixed keywords to match and display target comment content, the problem of not being able to personalize comments in existing technologies is solved, thus improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NETEASE (HANGZHOU) NETWORK CO LTD
- Filing Date
- 2022-10-09
- Publication Date
- 2026-07-24
AI Technical Summary
The existing comment display method cannot personalize the display based on the user's interests, forcing users to browse through a large number of irrelevant comments to find content of interest.
By acquiring the first time point of the video content, the target video content is extracted and processed to obtain the target object and its actions, generating mixed keywords, matching and displaying the target comment content corresponding to the keywords.
It enables personalized comment display based on user interests, improving the accuracy of comment content and user experience.
Smart Images

Figure CN115618864B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of human-computer interaction technology, and more specifically, to an information display method, an information display device, a computer-readable storage medium, and an electronic device. Background Technology
[0002] Existing comment display methods rely on the order or popularity of comments, failing to personalize the display based on user interests.
[0003] It should be noted that the information in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0004] The purpose of this disclosure is to provide an information display method, information display device, computer-readable storage medium, and electronic device, thereby overcoming, at least to some extent, the problem of the inability to personalize comment content due to limitations and defects in related technologies.
[0005] According to one aspect of this disclosure, an information display method is provided, comprising:
[0006] In response to a first interactive operation event acting on a preset interactive control, the first time node of the currently playing video content on the display interface at the time the first interactive operation event occurs is obtained;
[0007] Based on the first time point, target video content is extracted from the current video content, and the target video content is processed to obtain the target object and the action image of the target object;
[0008] Generate mixed keywords based on the target object, the action image, and the current attribute information of the current video content;
[0009] Match the target comment content corresponding to the hybrid keyword among the current comment content of the current video content, and display the target comment content.
[0010] In one exemplary embodiment of this disclosure, obtaining the first time point of the currently playing video content on the display interface includes:
[0011] In response to the first interactive operation event, obtain the first time node of the currently playing video content on the display interface when the first interactive operation event occurs;
[0012] The first interactive operation event includes interactive operation events and / or voice triggering events acting on a first preset interactive control; the preset interactive control is a video comment information interactive control.
[0013] In one exemplary embodiment of this disclosure, extracting target video content from the current video content based on the first time node includes:
[0014] Based on the user identification information of the event producer of the first interactive operation event, obtain the reference time range corresponding to the event producer;
[0015] Based on the first time node and the reference time range, calculate the start time node and end time node of the target video content in the current video content;
[0016] Based on the start time node and the end time node, extract the target video content from the current video content.
[0017] In one exemplary embodiment of this disclosure, the information display method further includes:
[0018] Obtain the historical browsing records of the event producer, and calculate the reference time range based on the historical browsing records.
[0019] In one exemplary embodiment of this disclosure, obtaining the historical browsing records of the event producer and calculating the reference time range based on the historical browsing records includes:
[0020] The target historical comments that the event producer focused on during the browsing of historical video content within the historical time period are obtained from the historical browsing records, and the target historical comments are segmented to obtain historical word groups.
[0021] Based on the frequency of occurrence of the historical phrases, target phrases are selected from the historical phrases, and the historical attribute information of the historical video content is obtained;
[0022] The reference time range is determined based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content.
[0023] In one exemplary embodiment of this disclosure, the information display method further includes:
[0024] Based on the image acquisition device included in the terminal device, the event producer's eye image is acquired during the process of browsing the original historical comment content corresponding to the historical video content;
[0025] Based on the human eye image, the eye gaze area of the event producer on the display interface of the terminal device is determined, and based on the eye gaze area, the target historical comment content is determined.
[0026] In one exemplary embodiment of this disclosure, determining the reference time range based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on a preset interactive control corresponding to the historical video content includes:
[0027] Obtain historical reference images including the historical attribute information and the target phrase, and determine the reference time node where the historical reference image is located in the historical video content;
[0028] Using the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content as the base point, calculate the time difference between the reference time node and the second time node;
[0029] The reference time range is obtained based on the time difference between the reference time node and the second time node.
[0030] In one exemplary embodiment of this disclosure, processing the target video content to obtain a target object and the motion image possessed by the target object includes:
[0031] The target video content is processed based on a preset image processing model to obtain keyframe images contained in the target video content;
[0032] Based on a preset target object extraction model, the target objects and their motion images included in the keyframe images are extracted.
[0033] In one exemplary embodiment of this disclosure, the preset target object extraction model includes a backbone feature extraction network, a neck feature fusion network, and a head feature detection network;
[0034] Specifically, based on a preset target object extraction model, the target objects and their motion characteristics included in the keyframe images are extracted, including:
[0035] The keyframe image is downsampled using the backbone feature extraction network to obtain the first local features;
[0036] The first local features are fused bidirectionally from deep to shallow and then from shallow to deep using the neck feature fusion network to obtain the first global features;
[0037] The head feature detection network is used to detect the category information and location information included in the first global feature to obtain the target object and the action image of the target object.
[0038] In one exemplary embodiment of this disclosure, a set of mixed keywords is generated based on the target object, the action image, and the current attribute information of the current video content, including:
[0039] Based on the target object and the action image, a first keyword is constructed, and based on the target object and the current attribute information of the current video content, a second keyword is constructed.
[0040] A third keyword is constructed based on the action image and the current attribute information of the current video content, and a fourth keyword is constructed based on the target object, the action image, and the current attribute information of the current video content;
[0041] The mixed keywords are constructed based on the first keyword and / or the second keyword and / or the third keyword and / or the fourth keyword.
[0042] According to one aspect of this disclosure, an information display device is provided, comprising:
[0043] The first time node acquisition module is used to obtain the first time node of the currently playing video content on the display interface.
[0044] The target video content processing module is used to extract target video content from the current video content according to the first time node, and process the target video content to obtain the target object and the action image of the target object;
[0045] The hybrid keyword generation module is used to generate hybrid keywords based on the target object, action image, and current attribute information of the current video content;
[0046] The target comment content display module is used to match the target comment content corresponding to the hybrid keyword in the current comment content of the current video content, and to display the target comment content.
[0047] According to one aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the information display method described in any of the preceding claims.
[0048] According to one aspect of this disclosure, an electronic device is provided, comprising:
[0049] Processor; and
[0050] Memory for storing the executable instructions of the processor;
[0051] The processor is configured to execute any of the above-described information display methods by executing the executable instructions.
[0052] This disclosure provides an information display method that, on the one hand, extracts target video content from the current video content based on a first time point, processes the target video content to obtain a target object and its associated actions and appearance; then, based on the target object, actions and appearance, and the current attribute information of the current video content, generates hybrid keywords, and matches the target comment content corresponding to the hybrid keywords with the current comment content in the current video content, thereby displaying the target comment content. This achieves the display of comment content based on hybrid keywords, thus solving the problem in the prior art that it is impossible to personalize the display of comment content according to the user's interests. This approach enables personalized display of comment content, thereby enhancing the user experience. Furthermore, it allows for the acquisition of the current time frame of the video content being played on the display interface. Based on this time frame, target video content is extracted from the current video content and processed to obtain the target object and its associated actions. Finally, based on the target object, actions, and current attribute information of the video content, a set of keywords is generated and matched with the target comment content corresponding to these keywords within the current comment content of the video content. This improves the accuracy of the target comment content and further enhances the user's viewing experience.
[0053] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0054] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0055] Figure 1 A flowchart illustrating an information display method according to an example embodiment of the present disclosure is shown schematically.
[0056] Figure 2 An example diagram schematically illustrates an interactive interface according to an exemplary embodiment of the present disclosure.
[0057] Figure 3An example diagram schematically illustrates a display interface for current comment content according to an exemplary embodiment of this disclosure.
[0058] Figure 4 The flowchart illustrates a method for extracting target video content from the current video content based on a first time point, according to an example embodiment of the present disclosure.
[0059] Figure 5 The flowchart illustrates a method for calculating the reference time range based on the historical browsing records of the event producer, according to an example embodiment of the present disclosure.
[0060] Figure 6 The illustration schematically shows an example diagram of a target object and its motion image obtained by processing target video content according to an example embodiment of the present disclosure.
[0061] Figure 7 An example diagram schematically illustrates a display interface for target comment content according to an exemplary embodiment of the present disclosure.
[0062] Figure 8 The flowchart illustrates a method for extracting target objects and their motion images from keyframe images based on a preset target object extraction model, according to an example embodiment of the present disclosure.
[0063] Figure 9 The diagram schematically illustrates a structural example of a visual feature extraction model according to an exemplary embodiment of the present disclosure.
[0064] Figure 10 The diagram schematically illustrates an example structure of a backbone feature extraction network according to an exemplary embodiment of the present disclosure.
[0065] Figure 11 The diagram schematically illustrates an example structure of a neck feature fusion network (Neck) according to an exemplary embodiment of the present disclosure.
[0066] Figure 12 The diagram schematically illustrates an example structure of a head feature detection network (Head) according to an exemplary embodiment of the present disclosure.
[0067] Figure 13 The illustration shows a scene example of a target object extraction model according to an example embodiment of the present disclosure, which extracts a target object included in a keyframe image and the motion image of the target object.
[0068] Figure 14 A block diagram schematically illustrates an information display device according to an exemplary embodiment of the present disclosure.
[0069] Figure 15 An electronic device for implementing the above-described information display method is illustrated according to an example embodiment of the present disclosure. Detailed Implementation
[0070] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0071] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0072] With the development of video shooting and editing technologies, more and more users are participating in the production and viewing of short videos. In practical applications, when a user watches a short video and finds a part of it that interests them, they will view the comments section.
[0073] In some short video comment displays, comments are presented based on their order or popularity. However, when there are many comments, the first few displayed may be irrelevant to the user's interests, requiring the user to scroll through numerous comments to find those that interest them. Therefore, how to display relevant comments based on the user's interests has become a pressing issue.
[0074] Based on this, this exemplary embodiment first provides an information display method, which can run on a terminal device (such as a mobile phone, tablet computer, PC, or smart display terminal, etc.); of course, those skilled in the art can also run the method disclosed herein on other platforms as needed, and this exemplary embodiment does not impose any special limitations on this. Specifically, refer to Figure 1 As shown, the information display method may include the following steps:
[0075] Step S110. Obtain the first time point of the currently playing video content on the display interface;
[0076] Step S120. Extract target video content from the current video content according to the first time node, and process the target video content to obtain the target object and the action image of the target object;
[0077] Step S130. Generate mixed keywords based on the target object, action image, and current attribute information of the current video content;
[0078] Step S140. Match the target comment content corresponding to the hybrid keyword in the current comment content of the current video content, and display the target comment content.
[0079] In the aforementioned information display method, on the one hand, since the target video content can be extracted from the current video content based on the first time node, and the target video content is processed to obtain the target object and the action image of the target object; then, based on the target object, action image, and the current attribute information of the current video content, a hybrid keyword is generated, and the target comment content corresponding to the hybrid keyword is matched in the current comment content of the current video content, and then the target comment content is displayed, the method realizes the display of comment content based on hybrid keywords, thereby solving the problem in the prior art that it is impossible to personalize the display of comment content according to the user's interests, and realizing the display of comment content based on hybrid keywords. Personalized content display enhances user experience. Furthermore, by capturing the current time frame of the video content on the display interface, target video content can be extracted and processed to obtain the target object and its actions. Finally, based on the target object, actions, and current video content attributes, mixed keywords are generated and matched with the corresponding target comments in the current video content. This improves the accuracy of the target comments and further enhances the user's viewing experience.
[0080] The steps included in the information display method of the present disclosure in the following will be explained and described in detail with reference to the accompanying drawings.
[0081] Specifically, in an information display method shown in the example embodiments of this disclosure:
[0082] In step S110, the first time point of the currently playing video content on the display interface is obtained.
[0083] In this example embodiment, obtaining the first time point of the currently playing video content on the display interface can be achieved as follows: In response to a first interactive operation event, obtain the first time point of the currently playing video content on the display interface at the time the first interactive operation event occurs; wherein, the first interactive operation event includes an interactive operation event acting on a first preset interactive control and / or a voice trigger event; the preset interactive control is a video comment information interactive control. That is, the first interactive operation event described here can be triggered manually by the current user to the video comment information interactive control, or it can be an event triggered by the terminal device receiving a specific voice message input by the current user; wherein, the voice message can be, for example, "open the comment section" or "open comment information," etc., and this example does not impose any special restrictions on this.
[0084] For details, please refer to Figure 2 As shown, during the current user's (i.e., the event producer of the first interactive operation event) viewing of a short video, if they need to view the current comments on the currently viewed video content, they can touch (i.e., the first interactive operation event applied to the preset interactive control) the preset interactive control (i.e., the video comment information interactive control) included in the current display interface where the current video content is located. After the display interface receives the first interactive operation event, it can display the current comments on the current video content. For details, please refer to... Figure 3 As shown. Simultaneously, when the first interactive operation event is received, it is also necessary to determine the first time point of the current video content at the time the first interactive operation event occurs; where the first time point refers to the specific playback progress of the current video content, such as 00:05:30, which means the current playback progress is five minutes and thirty seconds.
[0085] In step S120, target video content is extracted from the current video content according to the first time node, and the target video content is processed to obtain the target object and the action image of the target object.
[0086] In this example embodiment, after obtaining the first time node, the first step is to extract the target video content from the current video content based on the first time node. Specifically, refer to... Figure 4 As shown, the following steps may be included:
[0087] Step S410: Based on the user identification information of the event producer of the first interactive operation event, obtain the reference time range corresponding to the event producer.
[0088] In this example embodiment, the reference time range is first explained and described. Specifically, the reference time range described here refers to a video that moves forward a certain duration and backward a certain duration, based on a first time node. This method allows for the acquisition of target videos within a certain time period, and then inference of the comment content that the user needs to focus on based on these target videos. This reference time range can also be referred to as a time interval or a reference time range; this example does not impose any special restrictions on this. Furthermore, the specific calculation process for this reference time range can be obtained as follows: the historical browsing records of the event producer are obtained, and the reference time range is calculated based on these historical browsing records.
[0089] In one example embodiment, reference is made to... Figure 5 As shown, obtaining the historical browsing records of the event producer and calculating the reference time range based on the historical browsing records may include the following steps:
[0090] Step S510: Obtain the target historical comment content that the event producer paid attention to during the browsing of historical video content within the historical time period, as included in the historical browsing record, and perform word segmentation on the target historical comment content to obtain historical word groups.
[0091] In this example embodiment, the specific process for determining the target historical comment content is first explained and described. Specifically, the specific process for determining the target historical comment content can be implemented as follows: First, based on the image acquisition device included in the terminal device, the event producer's eye image is acquired during the browsing of the original historical comment content corresponding to the historical video content; second, the eye gaze area of the event producer on the display interface of the terminal device is determined based on the eye image, and the target historical comment content is determined based on the eye gaze area. That is, in a specific application process, when a user starts to view the comment content of a short video through the display interface of the terminal device, the image acquisition device connected to the terminal device can be activated to acquire the user's eye image, and then the eye gaze area can be determined based on the eye image, and the historical comment content appearing in the eye gaze area can be determined as the target historical comment content.
[0092] It should be noted that the image acquisition device described here can be an image acquisition device owned by the terminal device itself (such as the front or rear camera of a mobile phone), or an external image acquisition device connected to the terminal device. This example does not impose any special restrictions on this. At the same time, after obtaining the human eye image, the corresponding human eye gaze area can be determined based on the preset human eye image recognition algorithm.
[0093] Furthermore, once the target historical comment content is obtained, it can be segmented into words to obtain historical word groups. The specific segmentation process can be implemented using a word segmentation model, such as Word2Vec or other models. This example does not impose any special restrictions on this.
[0094] Step S520: Select a target phrase from the historical phrases based on the number of times they appear, and obtain the historical attribute information of the historical video content.
[0095] Specifically, once historical phrases are obtained, the frequency of each historical phrase in all target historical comments can be counted, and then target phrases can be determined based on this frequency (that is, extracting high-frequency phrases from historical phrases as target phrases). Then, the historical attribute information of the historical video content is obtained. This historical attribute information may include the video category, video producer, video source, video popularity, main cast and crew members included in the video, etc. For example, taking "Princess Pearl" as an example, the video category of the historical video content includes TV series, the video producer may include the director and producer of "Princess Pearl", etc., the main cast and crew members may include Xiao Yanzi, Ziwei, Erkang, Wu Age, etc., and the specific actors who played Xiao Yanzi, Ziwei, Erkang, and Wu Age are Zhao XX, Lin XX, Zhou XX, and Su XX, etc.
[0096] Step S530: Determine the reference time range based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content.
[0097] In this example embodiment, the reference time range is determined based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content. Specifically, this can be achieved as follows: First, a historical reference image including the historical attribute information and the target phrase is acquired, and the reference time node of the historical reference image within the historical video content is determined. Second, using the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content as a base point, the time difference between the reference time node and the second time node is calculated. Finally, the reference time range is obtained based on the time difference between the reference time node and the second time node. In other words, in specific applications, the comments that users are interested in can be obtained based on the user's eye movement and browsing gaze. High-frequency words in the comments can be extracted, and the high-frequency words can be matched with the content background of the video (such as the cast list for a TV series) to obtain a relevant image library. The search is performed before and after the second time point, and the frame video is broken down and compared with the image library to obtain the content corresponding to the high frequency. Then, based on the reference time point where the content corresponding to the high frequency appears in the historical video content, the corresponding reference time range is determined.
[0098] For example, if the target phrase is "eating," and the historical attribute information includes characters like Xiao Yanzi, Ziwei, Erkang, and Wu Age, then the relevant image library could include images of Xiao Yanzi eating, images of Ziwei eating, images of everyone eating together, etc. Based on these images, an image library is constructed. Then, a search is performed in the direction before and after the historical video content from the second time point to determine the position of these images within the historical video content, thus obtaining the reference time point. Finally, the reference time range is calculated. It should be noted that as long as any image from the image library appears in the historical video content, the corresponding reference time point can be determined. Of course, if multiple images from the image library appear in the historical video content, any of the following methods can be selected to determine the reference time point: the earliest appearing time point, the latest appearing time point, the intermediate appearing time point, or the average of all time points, etc. This example does not impose any special restrictions on this.
[0099] Step S420: Calculate the start and end time points of the target video content within the current video content based on the first time point and the reference time range.
[0100] Specifically, once the reference time range is obtained, the start and end time points can be determined based on the first time point and the reference time range. The start time point is the reference time range forward from the first time point, and the end time point is the reference time range backward from the first time point. At the same time, if the current video content is already at the start point of the video before the reference time range has been reached, or if the current video content has already ended before the reference time range has been reached, then the start point can be directly used as the start time point, and the end point can be used as the end time point.
[0101] Step S430: Extract the target video content from the current video content based on the start time node and the end time node.
[0102] Specifically, once the start and end times are obtained, the target video content, including the segment from the start to the end time, can be extracted from the current video content. The specific method for extracting video content can be implemented through video editing or through a video extraction model; this example does not impose any special restrictions on either approach.
[0103] Furthermore, once the target video content is obtained, it can be processed to obtain the target object and its motion image. Specifically, this can be achieved as follows: First, the target video content is processed based on a preset image processing model to obtain keyframe images within the target video content; second, based on a preset target object extraction model, the target object and its motion image included in the keyframe images are extracted. The image processing model described here can be a convolutional neural network model, a recurrent neural network model, or a deep neural network model, etc., and this example does not impose any special limitations on it. Further, once the keyframe images are obtained, the target object and its motion image can be extracted based on the preset target object extraction model. Specific example images of the extracted target object and its motion image can be found in [reference needed]. Figure 6 As shown.
[0104] In step S130, a set of mixed keywords is generated based on the target object, the action image, and the current attribute information of the current video content.
[0105] In this example embodiment, generating hybrid keywords based on the target object, the action image, and the current attribute information of the current video content can be achieved as follows: First, construct a first keyword based on the target object and the action image, and construct a second keyword based on the target object and the current attribute information of the current video content; second, construct a third keyword based on the action image and the current attribute information of the current video content, and construct a fourth keyword based on the target object, the action image, and the current attribute information of the current video content; finally, construct the hybrid keywords based on the first keyword and / or the second keyword and / or the third keyword and / or the fourth keyword.
[0106] Specifically, with Figure 6 Taking the target objects and actions shown as examples, specific target objects can include Zhen Huan and the Emperor, actions can include embracing, and current attribute information can include all information related to the Legend of Zhen Huan, such as the director, cast list, producer, theme song singer, etc. This example does not impose any special restrictions on this. Meanwhile, the first keyword can be "Zhen Huan and the Emperor embrace," or "Zhen Huan and the Fourth Prince embrace," the second keyword can include "Sun XX and Zhen Huan," "Sun XX and Chen XX," "Sun XX and the Emperor," etc.; the third keyword can include, for example, "Sun XX and Chen XX embrace," "Chen XX embrace," "Sun XX embrace," etc.; the fourth keyword can include "Zhen Huan (Sun XX) and the Emperor (Chen XX) embrace," etc. In generating mixed keywords, the first, second, third, and fourth keywords can be used as mixed keywords separately, or they can be combined in any pair or multiple combinations. This example does not impose any special restrictions on this.
[0107] In step S140, the target comment content corresponding to the hybrid keyword is matched among the current comment content of the current video content, and the target comment content is displayed.
[0108] In this example embodiment, after obtaining the mixed keywords, the target comment content, including all or part of the mixed keywords, can be matched within the current comment content of the current video content. Once the target comment content is obtained, the relevance between the target comment content and the mixed keywords can be calculated. The target comment content is then sorted according to the degree of relevance, and finally, the sorted target comment content is displayed. The specific display result of the obtained target comment content can be referenced... Figure 7 As shown.
[0109] The following will explain and illustrate the specific process of extracting the target object and the action image it possesses.
[0110] For details, please refer to Figure 8 As shown, based on a preset target object extraction model, the target objects and their motion characteristics included in the keyframe images are extracted. Specifically, this may include the following steps:
[0111] Step S810: Use the backbone feature extraction network to downsample the keyframe image to obtain the first local feature;
[0112] Step S820: The neck feature fusion network is used to perform bidirectional fusion of the first local features from deep to shallow and then from shallow to deep to obtain the first global features;
[0113] Step S830: The head feature detection network is used to detect the category information and location information included in the first global feature to obtain the target object and the action image of the target object.
[0114] The following will explain and describe steps S810-S830.
[0115] First, the pre-defined target object extraction model is explained and described. (Referencing...) Figure 9 As shown, the preset target object extraction model may include an input layer 901, a backbone feature extraction network 902, a neck feature fusion network 903, a head feature detection network 904, and an output layer 905. The input layer 901, backbone feature extraction network 902, neck feature fusion network 903, head feature detection network 904, and output layer 905 are connected sequentially.
[0116] In one example embodiment, reference is made to... Figure 10As shown, the backbone feature extraction network 902 may include a CBM module 1001 and multiple CSP modules, namely a first CSP module 1002, a second CSP module 1003, a third CSP module 1004, a fourth CSP module 1005, and a fifth CSP module 1006. That is, the backbone used in this exemplary embodiment can be implemented based on CSPDarknet53, which includes one CBM module and five CSP modules. The CBM module can be composed of Conv+Bn+Mish activation functions; Conv is convolutional convolution, Bn is batch normalization, and Mish is the activation function. The CSP module can be composed of the CBM module and one or more Res Unit modules concatenated. CSP can be used to represent Cross-Stage Partial, that is, it is a cross-stage local network with enhanced learning capabilities.
[0117] In one example embodiment, the first CSP module 1002 may consist of a CBM module and a Resunint module (Concatenated). In specific applications, the feature map 608 processed by the CBM module can be... 608 The 32-bit F_conv2 is passed to the first CSP module for processing (which contains only one residual unit).
[0118] In one example embodiment, the second CSP module 1003 can be composed of a CBM module and two Resunint modules concatenated. In specific applications, the 304 processed by the first CSP module can be... 304 The 64 feature maps are fed into the second CSP module for processing. Specifically, after downsampling, they become 152. 152 The feature map is 128, and then it is processed by two 1-bit systems respectively. 1 After a 64-fold convolution with s=1, two branches are obtained. The feature blocks of one branch are processed by the residual module and then concatenated with the other branch. The final output of the second CSP module is 152. 152 128 (Convolutional layers in the residual module: 1) 1 64 and 3 3 64).
[0119] In one example embodiment, the third CSP module 1004 can be composed of eight Res Unint modules and a CBM module concatenated. In specific applications, the 152 processed by the second CSP module can be... 152 The feature map of 128 is fed into the third CSP module for processing, and the final output of the third CSP module is 76. 76 256 (Convolutional layers in the residual module: 1) 1 128 and 3 3 (128) This module then splits into two branches: one branch continues to process the fourth CSP module, and the other branch goes directly to the Neck processing.
[0120] In one example embodiment, the fourth CSP module 1005 consists of eight Res Unint modules and a CBM module concatenated. In specific applications, a branch 76 of the third CSP module can be used... 76 The feature map of 256 is fed into the fourth CSP module for processing, and the final output of the fourth CSP module is 38. 38 512 (Convolutional layers in the residual module: 1) 1 256 and 3 3 (256) This module then splits into two branches: one branch continues processing the fifth CSP module, and the other branch goes directly into Neck processing.
[0121] In one example embodiment, the fifth CSP module 1006 consists of four Res Unint modules and a CBM module concatenated. In specific applications, a branch 38 of the fourth CSP module can be used... 38 The feature map of 512 is fed into the fifth CSP module for processing, and the final output of the fifth CSP module is 19. 19 1024 (Convolutional layers in the residual module: 1) 1 512 and 3 3 (512), the output of this module goes directly into the Neck processing.
[0122] In one example embodiment, reference is made to... Figure 11As shown, the Neck Feature Fusion Network (Neck) 903 may include an SPP module 1101, multiple CBL modules, multiple upsampling modules, and multiple stitching modules; wherein, the multiple CBL modules include a first CBL module 1102, a second CBL module 1103, a third CBL module 1104, a fourth CBL module 1105, a fifth CBL module 1106, a sixth CBL module 1107, a seventh CBL module 1108, an eighth CBL module 1109, a ninth CBL module 1110, a tenth CBL module 1111, an eleventh CBL module 1112, and a twelfth CBL module 1113; the multiple upsampling modules include a first upsampling module 1114 and a second upsampling module 1115; and the multiple stitching modules include a first stitching module 1116, a second stitching module 1117, a third stitching module 1118, and a fourth stitching module 1119.
[0123] The CBL module can be composed of three activation functions: Conv, Bn, and Leaky_relu. The three CBL modules preceding and following the SPP (Spatial Pyramid Pooling) module (the first and second CBL modules each containing three modules) are symmetrical, and their convolutions are 1... 1 512,3 3 1024 and 1 1 512, with a step size of 1 for all values; the SPP module can use max pooling methods of 1×1, 5×5, 9×9, and 13×13 for multi-scale fusion; specifically, the SPP module uses max pooling methods of 1×1, 5×5 padding=5 / / 2, 9×9 padding=9 / / 2, and 13×13 padding=13 / / 2 for multi-scale fusion. The output from the previous three CBL modules is: 19 19 The feature map of 512 is fed into the SPP module, and the final result is 19. 19 2048, after convolution by three CBL modules, yields 19. 19 Feature map of 512.
[0124] Furthermore, refer to Figure 12As shown, the Head feature detection network 904 can include multiple CBL modules and multiple convolutional modules; the multiple CBL modules can include a thirteenth CBL module 1201, a fourteenth CBL module 1202, and a fifteenth CBL module 1203; the multiple convolutional modules can include a first convolutional module 1204, a second convolutional module 1205, and a third convolutional module 1206. In specific applications, the Yolo Head uses the obtained features for prediction, which is a decoding process. Specifically, in the feature utilization part, YoloV4 extracts multi-scale features for object detection, extracting three feature layers: the middle layer, the lower-middle layer, and the bottom layer, with shapes of (19, 19, 255), (38, 38, 255), and (76, 76, 255), respectively. These three feature maps represent the entire detection result output by Yolo, including the bounding box position (4-dimensional), detection confidence (1-dimensional), and category (80-dimensional), totaling 85 dimensions. The final dimension of the feature map, 85, represents this information. The other dimensions of the feature map are N×N×3. N×N represents the reference position information of the detection box, and 3 represents the three prior boxes of different scales.
[0125] Secondly, in the process of extracting the target object and its associated action image, firstly, the backbone feature extraction network can be used to downsample the original image to obtain the first local features. Specifically, this can be achieved as follows: First, the CBM module is used to perform convolutional normalization and activation processing on the keyframe image to obtain the first convolutional feature map; secondly, the first CSP module is used to perform a first downsampling process on the first convolutional feature map to obtain the first downsampling result, and the second CSP module is used to downsample the first downsampling result to obtain the second downsampling result; then, the third CSP module is used... The sampling steps are repeated in the first, fourth, and fifth CSP modules to obtain the third, fourth, and fifth downsampling results in sequence, and the fifth downsampling result is used as the first local feature. Next, the neck feature fusion network is used to perform bidirectional fusion of the first local feature from deep to shallow and then from shallow to deep to obtain the first global feature. Specifically, this can be achieved as follows: First, the first CBL module is used to perform convolutional normalization and activation processing on the first local feature to obtain the second local feature, and the SPP module is used to perform multi-scale fusion processing on the second local feature to obtain the first global feature. The context features of the two local features are analyzed. First, the upper and lower features are convolved, normalized, and activated using a second CBL module to obtain a third local feature. Then, the third local feature is convolved, normalized, and activated using a third CBL module to obtain a fourth local feature. Next, the fourth local feature is upsampled using a first upsampling module to obtain a first upsampling result. Then, the fourth downsampling result is convolved, normalized, and activated using a fourth CBL module. Finally, the first concatenation module concatenates the first upsampling result and the convolved, normalized, and activated fourth downsampling result to obtain... The first stitching result is processed by convolution normalization and activation using the fifth CBL module to obtain the second stitching result. Further, the second stitching result is processed by convolution normalization and activation using the sixth CBL module to obtain the third stitching result. The third stitching result is then processed by upsampling using the second upsampling module to obtain the second upsampling result. Even further, the third downsampling result is processed by convolution normalization and activation using the seventh CBL module. The second stitching module then stitches the second upsampling result and the third downsampling result after convolution normalization and activation to obtain the fourth stitching result.Furthermore, the eighth CBL module performs convolutional normalization and activation processing on the fourth concatenation result to obtain a first global feature with a first preset scale. The ninth CBL module then performs convolutional normalization and activation processing on the first global feature with the first preset scale. The third concatenation module then concatenates the first global feature with the first preset scale (after convolutional normalization and activation processing) with the third concatenation result to obtain a fifth concatenation result. The tenth CBL module then performs convolutional normalization and activation processing on the fifth concatenation result to obtain a first global feature with a second preset scale. Finally, the eleventh CBL module... The CBL module performs convolutional normalization and planning processing on the first global feature with a second preset scale, and uses the fourth concatenation module to concatenate the third local feature and the first global feature with the second preset scale after convolutional normalization and activation processing to obtain the sixth concatenation result; then, the twelfth CBL module performs convolutional normalization and activation processing on the sixth concatenation result to obtain the first global feature with a third preset scale; finally, the head feature detection network is used to detect the category information and position information included in the first global feature to obtain the target object and the action image of the target object. Specifically, this can be achieved as follows: First, the thirteenth CBL module and the first convolutional module can be used to detect the category and location information of the target object included in the first global feature with a first preset scale, thereby obtaining image features with the first preset scale. Second, the fourteenth CBL module and the second convolutional module can be used to detect the category and location information of the target object included in the first global feature with a second preset scale, thereby obtaining image features with the second preset scale. Finally, the fifteenth CBL module and the third convolutional module can be used to detect the category and location information of the target object included in the first global feature with a third preset scale, thereby obtaining the target object with the third preset scale and the action image of the target object.
[0126] It should be noted that the target object extraction model used in the example embodiments of this disclosure is simpler and more convenient than the traditional Faster R-CNN (Region-CNN) model. Furthermore, this target object extraction model can use lightweight convolutions, resulting in fewer parameters, faster detection speed, and easier deployment in mobile terminal settings. Simultaneously, to fuse information from multiple receptive fields, this target object extraction model incorporates an SPP module in the Neck feature fusion network, thereby separating significant contextual features. Additionally, it achieves bidirectional fusion of feature information from deep to shallow layers and then from shallow to deep layers in the feature extraction part. For a specific example of feature extraction from keyframe images by the target feature extraction model, please refer to the relevant documentation. Figure 13 As shown.
[0127] This disclosure also provides an information display device through exemplary embodiments. Specifically, refer to... Figure 14 As shown, the information display device may include a first time node acquisition module 1410, a target video content processing module 1420, a hybrid keyword generation module 1430, and a target comment content display module 1440. Wherein:
[0128] The first time node acquisition module 1410 can be used to acquire the first time node of the currently playing video content on the display interface.
[0129] The target video content processing module 1420 can be used to extract target video content from the current video content according to the first time node, and process the target video content to obtain the target object and the action image of the target object;
[0130] The mixed keyword generation module 1430 can be used to generate mixed keywords based on the target object, action image and current attribute information of the current video content;
[0131] The target comment content display module 1440 can be used to match the target comment content corresponding to the hybrid keyword in the current comment content of the current video content, and display the target comment content.
[0132] In one exemplary embodiment of this disclosure, obtaining the first time point of the currently playing video content on the display interface includes:
[0133] In response to the first interactive operation event, obtain the first time node of the currently playing video content on the display interface when the first interactive operation event occurs;
[0134] The first interactive operation event includes interactive operation events and / or voice triggering events acting on a first preset interactive control; the preset interactive control is a video comment information interactive control.
[0135] In one exemplary embodiment of this disclosure, extracting target video content from the current video content based on the first time node includes:
[0136] Based on the user identification information of the event producer of the first interactive operation event, obtain the reference time range corresponding to the event producer;
[0137] Based on the first time node and the reference time range, calculate the start time node and end time node of the target video content in the current video content;
[0138] Based on the start time node and the end time node, extract the target video content from the current video content.
[0139] In one exemplary embodiment of this disclosure, the information display device further includes:
[0140] The reference time range calculation module can be used to obtain the historical browsing records of the event producer and calculate the reference time range based on the historical browsing records.
[0141] In one exemplary embodiment of this disclosure, obtaining the historical browsing records of the event producer and calculating the reference time range based on the historical browsing records includes:
[0142] The target historical comments that the event producer focused on during the browsing of historical video content within the historical time period are obtained from the historical browsing records, and the target historical comments are segmented to obtain historical word groups.
[0143] Based on the frequency of occurrence of the historical phrases, target phrases are selected from the historical phrases, and the historical attribute information of the historical video content is obtained;
[0144] The reference time range is determined based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content.
[0145] In one exemplary embodiment of this disclosure, the information display device further includes:
[0146] The human eye image acquisition module can be used to acquire human eye images of the event producer during the process of browsing the original historical comment content corresponding to the historical video content, based on the image acquisition device included in the terminal device;
[0147] The target historical comment content determination module can be used to determine the area of human eye gaze of the event producer on the display interface of the terminal device based on the human eye image, and determine the target historical comment content based on the area of human eye gaze.
[0148] In one exemplary embodiment of this disclosure, determining the reference time range based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on a preset interactive control corresponding to the historical video content includes:
[0149] Obtain historical reference images including the historical attribute information and the target phrase, and determine the reference time node where the historical reference image is located in the historical video content;
[0150] Using the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content as the base point, calculate the time difference between the reference time node and the second time node;
[0151] The reference time range is obtained based on the time difference between the reference time node and the second time node.
[0152] In one exemplary embodiment of this disclosure, processing the target video content to obtain a target object and the motion image possessed by the target object includes:
[0153] The target video content is processed based on a preset image processing model to obtain keyframe images contained in the target video content;
[0154] Based on a preset target object extraction model, the target objects and their motion images included in the keyframe images are extracted.
[0155] In one exemplary embodiment of this disclosure, the preset target object extraction model includes a backbone feature extraction network, a neck feature fusion network, and a head feature detection network;
[0156] Specifically, based on a preset target object extraction model, the target objects and their motion characteristics included in the keyframe images are extracted, including:
[0157] The keyframe image is downsampled using the backbone feature extraction network to obtain the first local features;
[0158] The neck feature fusion network is used to perform bidirectional fusion of the first local features from deep to shallow and then from shallow to deep to obtain the first global features;
[0159] The head feature detection network is used to detect the category information and location information included in the first global feature to obtain the target object and the action image of the target object.
[0160] In one exemplary embodiment of this disclosure, a set of mixed keywords is generated based on the target object, the action image, and the current attribute information of the current video content, including:
[0161] Based on the target object and the action image, a first keyword is constructed, and based on the target object and the current attribute information of the current video content, a second keyword is constructed.
[0162] A third keyword is constructed based on the action image and the current attribute information of the current video content, and a fourth keyword is constructed based on the target object, the action image, and the current attribute information of the current video content;
[0163] The mixed keywords are constructed based on the first keyword and / or the second keyword and / or the third keyword and / or the fourth keyword.
[0164] The specific details of each module in the aforementioned information display device have been described in detail in the corresponding information display methods, so they will not be repeated here.
[0165] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0166] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0167] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0168] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0169] The following reference Figure 15 To describe an electronic device 1500 according to such an embodiment of the present disclosure. Figure 15 The electronic device 1500 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0170] like Figure 15 As shown, the electronic device 1500 is manifested in the form of a general-purpose computing device. The components of the electronic device 1500 may include, but are not limited to: at least one processing unit 1510, at least one storage unit 1520, a bus 1530 connecting different system components (including storage unit 1520 and processing unit 1510), and a display unit 1540.
[0171] The storage unit stores program code that can be executed by the processing unit 1510, causing the processing unit 1510 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1510 can perform actions such as... Figure 1 The steps shown are as follows: Step S110: Obtain the first time node of the currently playing video content on the display interface; Step S120: Extract target video content from the current video content according to the first time node, and process the target video content to obtain the target object and the action image of the target object; Step S130: Generate hybrid keywords according to the target object, action image and the current attribute information of the current video content; Step S140: Match the target comment content corresponding to the hybrid keywords in the current comment content of the current video content, and display the target comment content.
[0172] Storage unit 1520 may include readable media in the form of volatile storage units, such as random access memory (RAM) 15201 and / or cache memory 15202, and may further include read-only memory (ROM) 15203.
[0173] Storage unit 1520 may also include a program / utility 15204 having a set (at least one) of program modules 15205, such program modules 15205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0174] Bus 1530 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0175] Electronic device 1500 can also communicate with one or more external devices 1600 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1500, and / or any device that enables electronic device 1500 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1550. Furthermore, electronic device 1500 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1560. As shown, network adapter 1560 communicates with other modules of electronic device 1500 via bus 1530. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0176] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0177] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible implementations, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this disclosure described in the "Exemplary Methods" section above.
[0178] The program product for implementing the above-described method according to embodiments of the present disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0179] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0180] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0181] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0182] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0183] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0184] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention described herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not invented by this disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
Claims
1. An information display method, characterized in that, include: In response to the first interactive operation event, obtain the first time node of the currently playing video content on the display interface when the first interactive operation event occurs; The first interactive operation event includes interactive operation events acting on the first preset interactive control and / or voice trigger events; The first preset interactive control is a video comment information interactive control; Based on the user identification information of the event producer of the first interactive operation event, obtain the reference time range corresponding to the event producer; based on the first time node and the reference time range, calculate the start time node and end time node of the target video content in the current video content; based on the start time node and end time node, extract the target video content from the current video content, and process the target video content to obtain the target object and the action image of the target object. The reference time range is determined based on the historical browsing records of the event producer; Generate mixed keywords based on the target object, the action image, and the current attribute information of the current video content; Match the target comment content corresponding to the hybrid keyword among the current comment content of the current video content, and display the target comment content.
2. The information display method according to claim 1, characterized in that, The information display method further includes: Obtain the historical browsing records of the event producer, and calculate the reference time range based on the historical browsing records.
3. The information display method according to claim 2, characterized in that, Obtain the historical browsing records of the event producer, and calculate the reference time range based on the historical browsing records, including: The target historical comments that the event producer focused on during the browsing of historical video content within the historical time period are obtained from the historical browsing records, and the target historical comments are segmented to obtain historical word groups. Based on the frequency of occurrence of the historical phrases, target phrases are selected from the historical phrases, and the historical attribute information of the historical video content is obtained; The reference time range is determined based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content.
4. The information display method according to claim 3, characterized in that, The information display method further includes: Based on the image acquisition device included in the terminal device, the event producer's eye image is acquired during the process of browsing the original historical comment content corresponding to the historical video content; Based on the human eye image, the eye gaze area of the event producer on the display interface of the terminal device is determined, and based on the eye gaze area, the target historical comment content is determined.
5. The information display method according to claim 3, characterized in that, The reference time range is determined based on the historical attribute information, the target phrase, and the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content, including: Obtain historical reference images including the historical attribute information and the target phrase, and determine the reference time node where the historical reference image is located in the historical video content; Using the second time node of the second interactive operation event of the event producer acting on the preset interactive control corresponding to the historical video content as the base point, calculate the time difference between the reference time node and the second time node; The reference time range is obtained based on the time difference between the reference time node and the second time node.
6. The information display method according to claim 1, characterized in that, The target video content is processed to obtain the target object and the action image of the target object, including: The target video content is processed based on a preset image processing model to obtain keyframe images contained in the target video content; Based on a preset target object extraction model, the target objects and their motion images included in the keyframe images are extracted.
7. The information display method according to claim 6, characterized in that, The preset target object extraction model includes a backbone feature extraction network, a neck feature fusion network, and a head feature detection network. Specifically, based on a preset target object extraction model, the target objects and their motion characteristics included in the keyframe images are extracted, including: The keyframe image is downsampled using the backbone feature extraction network to obtain the first local features; The neck feature fusion network is used to perform bidirectional fusion of the first local features from deep to shallow and then from shallow to deep to obtain the first global features; The head feature detection network is used to detect the category information and location information included in the first global feature to obtain the target object and the action image of the target object.
8. The information display method according to claim 1, characterized in that, Based on the target object, the action image, and the current attribute information of the current video content, generate mixed keywords, including: Based on the target object and the action image, a first keyword is constructed, and based on the target object and the current attribute information of the current video content, a second keyword is constructed. A third keyword is constructed based on the action image and the current attribute information of the current video content, and a fourth keyword is constructed based on the target object, the action image, and the current attribute information of the current video content; The mixed keywords are constructed based on the first keyword and / or the second keyword and / or the third keyword and / or the fourth keyword.
9. An information display device, characterized in that, include: The first time node acquisition module is used to respond to the first interactive operation event and acquire the first time node of the currently playing video content on the display interface when the first interactive operation event occurs. The first interactive operation event includes interactive operation events acting on the first preset interactive control and / or voice trigger events; The first preset interactive control is a video comment information interactive control; The target video content processing module is used to obtain a reference time range corresponding to the event producer based on the user identification information of the event producer of the first interactive operation event; calculate the start time node and end time node of the target video content in the current video content based on the first time node and the reference time range; extract the target video content from the current video content based on the start time node and end time node, and process the target video content to obtain the target object and the action image of the target object; the reference time range is determined based on the historical browsing records of the event producer; The hybrid keyword generation module is used to generate hybrid keywords based on the target object, action image, and current attribute information of the current video content; The target comment content display module is used to match the target comment content corresponding to the hybrid keyword in the current comment content of the current video content, and display the target comment content.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the information display method according to any one of claims 1-8.
11. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the information display method according to any one of claims 1-8 by executing the executable instructions.