Content display method and device, electronic equipment and readable storage medium
By generating query information related to the subtitle text and utilizing the correspondence between cached key information and subtitle text, the problem of users being unable to obtain detailed competition information was solved, enabling accurate review of competition video content and improving the live streaming experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MIGU CO LTD
- Filing Date
- 2023-10-18
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies cannot meet users' needs for obtaining the video content they require, especially in match videos where detailed match information cannot be accurately obtained through subtitle text.
By receiving user input on subtitle text, query information related to the subtitle text is generated. Using the correspondence between pre-cached key information and subtitle text, the subtitle text corresponding to the target key information is determined and displayed on the playback interface, enabling a review of the video content.
Users can accurately obtain detailed information about the match based on the subtitle operation, review the video content, improve the user viewing experience, and not affect the live broadcast of the match.
Smart Images

Figure CN117312543B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, specifically relating to a content display method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] In related technologies, to facilitate understanding of the competition content, the commentary is typically displayed as subtitles to users. When a user wants to understand a particular participant's performance in the current match, the subtitles corresponding to the commentary can be integrated and displayed. Methods for integrating subtitles mainly include storing existing subtitle texts in a list arranged chronologically. Users can retrieve subtitle texts within a specific time range by selecting a time point or time period, thus understanding the detailed content of the match video. However, in this case, since the match time does not represent the specific video content, it cannot meet the user's need to obtain the desired video content. Summary of the Invention
[0003] The purpose of this application is to provide a content display method, apparatus, electronic device, and readable storage medium to solve the problem that related technologies cannot meet users' needs for obtaining required video content.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows:
[0005] Firstly, it provides a content display method, including:
[0006] Receive user input for a first subtitle text, which is the subtitle text displayed on the video playback interface;
[0007] In response to the input operation, query information related to the first subtitle text is generated, and the query information includes at least one target key information;
[0008] Based on the correspondence between pre-cached key information and subtitle text, determine the second subtitle text corresponding to the at least one target key information;
[0009] The second subtitle text is displayed on the playback interface.
[0010] Secondly, a content display device is provided, including:
[0011] The receiving module is used to receive user input for the first subtitle text, which is the subtitle text displayed on the video playback interface.
[0012] The generation module is configured to generate query information related to the first subtitle text in response to the input operation, wherein the query information includes at least one target key information;
[0013] The first determining module is used to determine the second subtitle text corresponding to the at least one target key information based on the correspondence between the pre-cached key information and the subtitle text;
[0014] A display module is used to display the second subtitle text on the playback interface.
[0015] Thirdly, an electronic device is provided, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0016] Fourthly, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0017] In this embodiment, by receiving user input on a first subtitle text (the subtitle text displayed on the video playback interface), query information related to the first subtitle text is generated. The query information includes at least one target key information. Based on the correspondence between pre-cached key information and subtitle text, a second subtitle text corresponding to the at least one target key information is determined and displayed on the playback interface. This allows for the determination and display of previous subtitle text based on user operations on subtitles and related target key information, thereby enabling video content review and satisfying the user's need to obtain the required video content. Attached Figure Description
[0018] Figure 1 This is a flowchart of a content display method provided in an embodiment of this application;
[0019] Figure 2 This is a schematic diagram of subtitle caching in an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of the system architecture in an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of the event content review process in an embodiment of this application;
[0022] Figure 5 This is a schematic diagram of the subtitle display method in the embodiments of this application;
[0023] Figure 6This is a schematic diagram of the structure of a content display device provided in an embodiment of this application;
[0024] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] The following description, in conjunction with the accompanying drawings, details the content display method, apparatus, electronic device, and readable storage medium provided in the embodiments of this application through specific examples and application scenarios.
[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating a content display method provided in an embodiment of this application. The method is applied to an electronic device, specifically a client terminal, such as a mobile phone, tablet computer, or laptop computer. Figure 1 As shown, the method includes the following steps:
[0029] Step 11: Receive user input for the first subtitle text, which is the subtitle text displayed on the video playback interface;
[0030] Step 12: In response to the input operation, generate query information related to the first subtitle text, the query information including at least one target key information;
[0031] Step 13: Determine the second subtitle text corresponding to the at least one target key information based on the correspondence between the pre-cached key information and the subtitle text;
[0032] Step 14: Display the second subtitle text on the playback interface.
[0033] Here, the video can be a match video, meaning that the video playback interface displays a match. The match can include, but is not limited to, football matches, basketball matches, etc. The playback interface can be a live streaming interface, meaning that video / event content is displayed during a live match broadcast. The subtitle text is specifically derived from the commentary on the match.
[0034] The input operation can be selected as, but is not limited to, the user pressing or sliding the first subtitle text, or the user pressing or sliding the playback interface where the first subtitle text is located. There is no limitation on this, and it can be set according to actual needs.
[0035] The target key information can include one or more key information pieces; that is, the same subtitle text can contain different key information pieces simultaneously. The key information can be divided into different categories, and key information within the same category can be divided into different key information pieces. For different categories of key information, separate caches can be set up, that is, the subtitle text corresponding to different categories of key information is stored in different caches, thereby distinguishing the subtitle text and speeding up the retrieval of the required subtitle text. Each cache contains multiple cache queues, and each cache queue corresponds to one key information piece, thus forming a correspondence between key information and subtitle text, allowing for more accurate distinction of subtitle text. For example, taking a football match as an example, the key information corresponding to the subtitle text can be divided into three categories: athletes (i.e., players), teams (i.e., teams), and match events, such as goals, fouls, and substitutions. In this case, if... Figure 2 As shown, three caches can be used to store the corresponding subtitle texts: 1) Subtitles containing different key information are stored in different caches. For example, subtitle texts corresponding to athletes are stored in the athlete-associated cache, subtitle texts corresponding to sports teams are stored in the team-associated cache, and subtitle texts corresponding to match events are stored in the match event-associated cache; 2) Each cache contains multiple cache queues, such as Redis queues, and each cache queue corresponds to one piece of key information. For example, for the key information "athlete," the key of each cache queue is the athlete's identifier, which can be the athlete's name, athlete number, athlete nickname, etc.; for the key information "sports team," the key of each cache queue is the team's identifier, which can be the team name, team nickname, etc.; for the key information "match event," the key of each cache queue is the match event's identifier; 3) Each cache queue is sorted in chronological order (first-in, first-out), meaning the relevant subtitle texts are sorted from earliest to latest. This allows for more precise differentiation of subtitle texts, thus meeting the need for fast and accurate retrieval of relevant subtitle texts.
[0036] Optionally, the correspondence between the key information and the subtitle text can be obtained from the server. This correspondence is obtained by the server based on the analysis of real-time acquired subtitle data and corresponding video frame groups. For example, by combining subtitle text with video understanding technology, a correspondence can be established between the subtitle text and key information (such as athletes, sports teams, and / or competition events).
[0037] Optionally, after generating the query information related to the first subtitle text, the corresponding second subtitle text can be retrieved using the server. The specific process includes: first, sending the query information to the server; then, the server, based on the pre-cached correspondence between key information and subtitle text, determines at least one second subtitle text corresponding to a target key piece of information and sends it to the client terminal, which then displays the second subtitle text.
[0038] Optionally, the query information may also include the time point of the first subtitle text. Then, when determining the second subtitle text, the preset bar (e.g., 20 subtitles) closest to that time point is selected for display. Regarding the display method, the selected second subtitle text can be displayed in a paginated format on the side of the live stream interface, allowing users to scroll through the content; alternatively, it can be displayed as a pop-up message, with no limitation on this method.
[0039] The content display method of this application embodiment receives user input on first subtitle text, generates query information related to the first subtitle text, the query information includes at least one target key information, determines the second subtitle text corresponding to the at least one target key information based on the correspondence between the pre-cached key information and the subtitle text, and displays the second subtitle text on the playback interface. It can determine and display previous subtitle text based on user operation on subtitles and related target key information, thereby realizing the review of video / event content and satisfying the user's need to obtain the required video / event content.
[0040] Furthermore, with the solution in this application, when reviewing video / event content, there is no need to pause the game, which will not affect the user's viewing of the live game and improve the user's viewing experience; for users who watch the game in the middle, the accurate commentary subtitles can help them review the event content; for people with hearing impairments, the intelligent subtitle application can not only help them understand the current game content, but also help them review the previous game content.
[0041] In this embodiment, query information can be obtained based on the analysis of the subtitle text, or it can be obtained by combining an understanding of the relevant video frame groups. The generation of query information related to the first subtitle text can include at least one of the following:
[0042] 1) Analyze the first subtitle text to obtain the target key information contained in the first subtitle text;
[0043] 2) Obtain the video frame group corresponding to the first subtitle text, and analyze each video frame in the video frame group to obtain the target key information contained in the video frame group.
[0044] Here, for the video frame group corresponding to the first subtitle text, the corresponding video frame group can be determined based on the time point of the first subtitle text. For example, using the time point of the first subtitle text as a reference, determine the video frame at that time point, and then shift forward and backward by multiple frames (such as 8 frames) to obtain the corresponding video frame group.
[0045] In some embodiments, the target key information can be obtained by analyzing the video frame group corresponding to the first subtitle text, even if the first subtitle text does not contain the target key information. Alternatively, the target key information can be obtained by analyzing the first subtitle text and its corresponding video frame group.
[0046] In this embodiment, different methods can be used to analyze video frame groups based on different target key information. If the target key information is an athlete, the correlation between the video frame group and the athlete can be determined by identifying the pixel proportion of each athlete in the video frame. The process of analyzing each video frame in the video frame group can include:
[0047] Identify the pixel proportion of each athlete in each video frame related to the competition video; for example, a pre-trained SSD (Single Shot MultiBox Detector) model can be used to identify athletes in each video frame frame by frame and calculate the pixel proportion of each athlete in the video frame in which they appear.
[0048] The appearance ratio of each athlete is determined based on the pixel percentage of each athlete in each video frame; the appearance ratio of each athlete is equal to the ratio of the sum of the pixel percentages of each athlete in the video frame group to the sum of the pixel percentages of all athletes in the video frame group.
[0049] The occurrence ratio of each athlete is normalized, and when the normalized occurrence ratio of each athlete follows a normal distribution, the preset athletes with the largest normalized occurrence ratio are confirmed as the target key information.
[0050] Here, when the normalized occurrence ratio of each athlete follows a normal distribution, it means that this set of video frames focuses on several athletes, and the preset athletes (such as 1 or 2 athletes) with the highest normalized occurrence ratio can be identified as the target key information. If the normalized occurrence ratio of each athlete does not follow a normal distribution, it can be considered that this set of video frames and the corresponding subtitle text do not involve the relevant athletes.
[0051] Optionally, if the target key information is a sports team, the correlation between the video frame group and the sports team can be determined based on the changes in the number of athletes from each sports team within the video frame group. The process of analyzing each video frame in the video frame group described above may include:
[0052] Identify the number of first athletes for each team in a video frame group and the number of second athletes in the last video frame group; for example, a pre-trained ByteTracker model can be used to track athletes within a video frame group.
[0053] Based on the first number of athletes and the second number of athletes, the reduction rate of each sports team is determined. This reduction rate actually reflects the change in the camera shots of each sports team. If the reduction rate is large, it means that the camera is moving away from the main sports team in the current frame, which means that the relevance between the narration and the main sports team in the current frame will also decrease.
[0054] Based on the number of athletes and the reduction rate of each sports team, the relevance weight of each sports team is determined, and sports teams with a relevance weight greater than a preset threshold are indeed key target information.
[0055] Here, the reduction rate for each sports team can be determined based on the difference between the number of first-place athletes and the number of second-place athletes for each team. For example, if the number of first-place athletes and the number of second-place athletes for sports team i are Num0 and Num, respectively... n The reduction rate Lost of team i can then be calculated using the following formula. i :
[0056] Lost i -(Num0-Num n ) / Num0
[0057] Optionally, considering that the number of tracked athletes may increase or decrease over time, and the differences between the initial number of athletes, the relevance weight of the sports teams can be used to determine whether the relevant video frame group contains target key information. That is, the sports team is determined to be target key information only if the relevance weight of the sports team is greater than a preset threshold; otherwise, it can be considered that the video frame group and the corresponding subtitle text do not involve the relevant sports team.
[0058] For example, the correlation weight W of each sports team can be calculated using the following formula. i :
[0059]
[0060] Optionally, if the target key information is a match event, since match events such as goals, fouls, and substitutions all require referee intervention, the correlation between the video frame group and the match event can be determined based on the frequency of referee appearances within the video frame group. The process of analyzing each video frame in the video frame group described above can include:
[0061] Identify the number of times a referee appears in a group of video frames related to a match video; for example, a pre-trained SSD model can be used to identify the number of times a referee appears in a group of video frames.
[0062] When the number of occurrences meets a preset condition, the competition event in the video frame group is confirmed as the target key information. For example, the preset condition may be that the number of occurrences is greater than or equal to a preset threshold, or it may be that the ratio of the number of occurrences to the number of frames in the video frame group is greater than or equal to a preset threshold; there is no limitation on this.
[0063] It should be noted that the same subtitle text may contain different key information and may belong to different cache queues. Therefore, the above analysis process for the corresponding video frame group can be performed to fully determine the relevant key information.
[0064] In this embodiment, since the understanding of video frames may be erroneous, the degree of correlation between the video frame group and the subtitle text can be further determined based on the above analysis results of the video frame group. After analyzing each video frame in the video frame group, this embodiment may further include:
[0065] Based on the target key information, the search starts from the head of the cache queue corresponding to the target key information to obtain the first subtitle text containing the target key information, and selects multiple subtitle texts that are adjacent to the first subtitle text containing the target key information; the head of the cache queue stores the latest related subtitle texts;
[0066] Based on the time corresponding to each subtitle text in the plurality of subtitle texts, a first time interval is determined, wherein the first time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the plurality of subtitle texts;
[0067] Compare the first time interval with the second time interval, where the second time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the cache queue corresponding to the target key information;
[0068] When the first time interval is less than or equal to the second time interval, the target key information is determined to be related to the first subtitle text, and the target key information is used as the query information; or, when the first time interval is greater than the second time interval, the target key information is determined to be unrelated to the first subtitle text, and the target key information is deleted from the query information.
[0069] Here, since commentators will inevitably mention athletes' names or nicknames, teams' names or nicknames, and match events during their commentary, the content expressed in the subtitles will focus on these key information when they appear. In effect, these subtitles will appear very close in time. Therefore, the size of the time interval can be used to determine whether new subtitles are related to key information.
[0070] The present application will now be described in detail with reference to specific embodiments.
[0071] In the embodiments of this application, such as Figure 3 As shown, the process is mainly divided into two parts: one is the subtitle processing and querying on the backend server, and the other is the user interaction and result viewing on the client terminal. The overall execution flow can include: 1) The backend server receives real-time subtitle data of the event, performs text processing on the subtitle data, and performs corresponding analysis to extract key information such as athletes, teams, and event information; 2) The extracted key information and corresponding subtitle text are associated, and the subtitle text and associated information are stored, that is, the correspondence between key information and subtitle text is established and cached; 3) For subtitle text that cannot be associated, video understanding technology is used to analyze and process the corresponding video frames, establish the correspondence between key information and subtitle text, and cache it; 4) Users can press the subtitles of interest or swipe the playback screen on the viewing interface to query the corresponding subtitle text they need / are interested in, and the queried subtitle text is displayed on the viewing interface without affecting the user's viewing of the event.
[0072] The specific implementation process of the solution in this application can be divided into server-side and front-end, as described below.
[0073] For the server-side, the key information corresponding to the subtitle text is divided into three categories: athletes, teams, and match events, including goals, fouls, and substitutions; in this case, such as Figure 2 As shown, three caches are used to store the corresponding subtitle text: 1) Subtitles containing different key information are stored in different caches. For example, the subtitle text corresponding to athletes is stored in the subtitle cache associated with athletes, the subtitle text corresponding to sports teams is stored in the subtitle cache associated with sports teams, and the subtitle text corresponding to competition events is stored in the subtitle cache associated with competition events; 2) Each cache contains multiple cache queues, such as Redis queues, and each cache queue corresponds to a key piece of information. For example, for the key information "athlete," the key of each cache queue is the athlete's identifier, which can be used to identify athletes. The key information includes the athlete's name, athlete number, and athlete nickname. For teams, the key for each cache queue is the team's identifier, which can be the team name, nickname, etc. For match events, the key for each cache queue is the match event's identifier. 3) Each cache queue is sorted in chronological order (first-in, first-out), meaning the relevant subtitle texts are sorted from earliest to latest, which facilitates front-end display. 4) As the match continues, new subtitle data is continuously generated. This new subtitle data is continuously input into the server for processing, and the correspondence between key information and subtitle text is stored.
[0074] Optionally, the storage structure of the cache queue in this embodiment can be as follows:
[0075] {Key information, subtitle text, subtitle display time}
[0076] Key information includes, for example, the athlete's name, team name, or the name of the match event. The subtitle text is the specific content of the subtitle. The time point at which the subtitle appears in the match video is the exact moment it appears.
[0077] The specific processing procedure on the server side is as follows:
[0078] Step 1: The server performs initialization, which mainly includes the following parts:
[0079] 1) Initialize the cache to store the processed subtitle text;
[0080] 2) Preload information related to sports teams and athletes;
[0081] ① For sports teams, since commentators or fans often refer to them by nicknames, and these nicknames are relatively fixed, the nicknames of the participating sports teams can be pre-loaded to improve the accuracy of identifying key information in the subtitles. For example, the data format for pre-loaded information can be as follows:
[0082] {Team Name, Team Nickname 1, Team Nickname 2, ...}
[0083] ② For athletes, similar to the function of team nicknames, since commentators or fans often refer to athletes by their numbers or nicknames, the numbers or nicknames of the athletes participating in the competition can be pre-loaded to improve the accuracy of identifying key information in the subtitles; for example, the data format for pre-loaded information can be as follows:
[0084] {Athlete Name, Athlete Number, Nickname 1, Nickname 2, Nickname 3, ...}
[0085] 3) Initialize a cache to store match event information, including goals, fouls, and substitutions, in the following format:
[0086] {Event ID, Event Content, Event Time}
[0087] The event ID indicates a goal, a foul, or a substitution. The event content is determined in real-time by match data during the live broadcast, such as who committed the foul. The event time is the time when the event occurred, for example, calculated based on the time elapsed in the match, such as 49:50 in the entire match.
[0088] Step 2: The server receives subtitle-related data, which includes: 1) subtitle text, denoted as sub_text, and the corresponding display time point sub_time; 2) video frame group, denoted as frames_i, with each video frame group numbered as i, and each video frame group consisting of multiple video frames.
[0089] After receiving the above two types of data, the subtitle text is aligned with the video frame group. For example, the time point when the subtitle text sub_text_i is displayed is sub_time_i; the time of the video frame group can be calculated from the start time of the live broadcast and the video frame rate. Video frames from time point sub_time_i-1 to sub_time_i can be collected and used as the video frames related to the subtitles to form a new video frame group n_frame_i. Optionally, for video frame group n_frame_i, 8 frames from the previous time period and 8 frames from the next time period can be concatenated to the beginning and end of video frame group n_frame_i to solve the problem of the subtitle's correlation with the preceding and following video frames.
[0090] Step 3: For the subtitle text, perform preliminary key information extraction, that is, use the information pre-loaded in Step 1, such as using athlete names or nicknames, sports team names or nicknames, and competition event names for keyword matching; if the match is successful, store the subtitle text in the relevant cache queue according to the extracted key information; if the match is unsuccessful, proceed to Step 4.
[0091] For example, if a player from Team 1 scores 10 goals, and the commentary reads "Player scores 10 goals, Team 1 takes the lead," then the following cached result can be generated:
[0092] Matched with athlete: {Athlete 10, "Athlete 10 scores a goal, Team 1 is leading", time point};
[0093] Matched with a sports team: {Sports Team 1, "Athlete has scored 10 goals, Sports Team 1 is in the lead", time point};
[0094] Matching event: {Goal, "Athlete scores 10 goals, team leads by 1", time point}.
[0095] Based on the above caching results, it can be seen that a single subtitle text can simultaneously identify multiple key pieces of information.
[0096] Step 4: Based on video understanding, determine the relationship between the subtitle text and the video frames. This step mainly addresses subtitles without subjects, such as "He should have passed the ball through" or "The formation is too far to the right, resulting in too much space on the left," where key information cannot be directly extracted. Taking the subtitle text sub_text_i and the corresponding video frame group n_frame_i as an example, the specific determination process includes:
[0097] (1) For athletes:
[0098] First, a pre-trained SSD model is used to identify athletes appearing in each video frame, and the pixel proportion of each athlete in the video frame in which they appear is calculated. For example, for athlete i, the following results can be obtained:
[0099] Pi = {Pix1, Pix2, ..., Pix} n}
[0100] Where Pi is a list of pixel proportions for athlete i, Pix is the pixel proportion of athlete i in a video frame, and n is the number of video frames in the video frame group. If athlete i does not appear in a video frame, its pixel proportion is recorded as 0.
[0101] Based on the above, the following results can be obtained for the identified athletes:
[0102] CP={P1, P2,…,Pi,…,P m}
[0103] Where m is the number of athletes identified. Pi represents the percentage of pixels of athlete i in each video frame.
[0104] Then, based on the pixel percentage of each athlete in each video frame, the appearance ratio of each athlete is determined. The appearance ratio of each athlete is equal to the ratio of the sum of the pixel percentages of each athlete in the video frame group to the sum of the pixel percentages of all athletes in the video frame group. That is, the appearance ratio of each athlete is obtained by summing up the pixel percentages of all identified athletes in each video frame as the denominator, and summing up the pixel percentages of each individual athlete in each video frame as the numerator, and dividing by the sum, as shown in the following formula:
[0105]
[0106] Finally, the occurrence ratio of each athlete is normalized. When the normalized occurrence ratio of each athlete follows a normal distribution, the preset number of athletes (e.g., 2 athletes) with the largest normalized occurrence ratio are selected as key information, and the following content is output:
[0107] {Athlete Name, Subtitle Text, Time Point}
[0108] If the normalized occurrence ratio of each athlete follows a normal distribution, it means that this set of video frames focuses on several athletes, and these can be used to extract key information about the athletes. However, if the normalized occurrence ratio of each athlete does not follow a normal distribution, it is considered that this subtitle does not involve the relevant athletes.
[0109] (2) While extracting key information about athletes, key information about the sports team can also be extracted. The specific process includes:
[0110] First, a pre-trained Byte Tracker model is used to track athletes within video frame groups. Simultaneously, a pre-trained ResNet model is used to distinguish teams based on jersey color; for example, team 1 is blue, and team 2 is white. The reduction rate for team i is then calculated using the following formula:
[0111] Lost i =(Num0-Num n ) / Num0
[0112] Where Num0 represents the number of tracked objects identified for the first time; Num n This indicates the number of tracked objects at the end of a video frame group. For two sports teams, this can be denoted as Lost_a and Lost_b.
[0113] Considering that the number of tracked objects may increase or decrease, and the differences in the initial number of tracked objects, the relevance weight of each sports team can be calculated, and key information can be determined based on the relevance weight.
[0114] It should be noted that using the sigmoid function to normalize the reduction rate of each team, and multiplying it by the initial number of athletes, can illustrate the camera movement of each team. This reduction rate actually reflects the camera movement of each team; a large reduction rate indicates that the camera is moving away from the main team in the current shot, meaning the correlation between the narration and the main team in the current shot will decrease. Even if the reduction rate is negative, meaning the number of tracked targets has increased, the monotonically increasing property of the sigmoid function can maintain the accuracy of the correlation weight calculation.
[0115] For example, if the initial number of followers for one sports team is significantly greater than that for another, and the reduction rate is very small, then the sports team with the larger number of followers can be identified as the sports team mentioned in the subtitles. If the reduction rate for the sports team with the larger number of followers is also very large, then the determination can be made based on the relevance weight. That is, sports teams with a relevance weight greater than a preset threshold are indeed key information. The input result is as follows:
[0116] {Sports team name, subtitle text, time point}
[0117] (3) While performing (1) and (2) above, key information about the competition event can also be extracted. The specific process includes:
[0118] First, a pre-trained SSD model is used to identify the number of times the referee appears, K, in a group of video frames; this is because match events such as goals, fouls, and substitutions all require the intervention of the referee.
[0119] Then, when the number K is greater than [0.75*r], where r is the number of frames in the video frame group, that is, when 3 / 4 of the frames in the video frame group contain the referee, it is considered that the referee's participation is high, i.e., a match event has occurred. After obtaining the relevant match data, the output result is as follows:
[0120] {Match events, subtitles, time points}
[0121] Step 5: Based on Step 4 above, further analysis is conducted to determine the association information of the subtitles. The specific processing procedure is as follows:
[0122] S51: For the three different types of cache, search the corresponding cache queue according to the result of step 4, that is, search according to athlete name, team name or competition event; search from the head of the queue backwards, find the first subtitle text containing key information in reverse chronological order, that is, the latest subtitle text containing key information; for example, if the number of subtitles in the cache queue is m and numbered from 1 to m, and the subtitle text containing key information is found to be m-3, then extract the time points corresponding to the four subtitle texts m-3, m-2, m-1, and m, and subtract them from each other to obtain the average value to get the time interval n_space_i;
[0123] S52: Subtract the time points of the subtitle texts in the same cache queue from each other in chronological order to obtain the time interval between the subtitle texts. Then, calculate the average of these time intervals to obtain the average time interval, denoted as r_space_i.
[0124] S53: For the same type of cache queue, such as the caption cache queue for athlete 5, if the time interval n_space_i is less than or equal to the average time interval r_space_i, then the association result of step 4 is considered to be correct, that is, the current caption is related to athlete 5; otherwise, the association result of step 4 is considered to be incorrect, that is, the caption does not belong to the current category, and the result of step 4 is discarded.
[0125] The reason for executing step 5 is that commentators will inevitably mention athletes' names or nicknames, teams' names or nicknames, and match events during their commentary. When this key information appears, the content of the subtitle text will focus on this key information. In practice, this means that these subtitles will appear very close in time. Therefore, the time interval can be used to determine whether new subtitle text is related to key information, thereby improving the accuracy of subtitle association. It is important to note that the same subtitle text can belong to three different cache queues simultaneously, meaning that this subtitle contains three types of key information.
[0126] Step 6: Store the determined results into the corresponding cache queue based on the associated key information, including athletes, teams, and / or match events. This allows for the integration of video understanding to establish the association between subtitle text and athletes, teams, and / or match events, thereby achieving intelligent classification of subtitle text.
[0127] The front-end in this embodiment primarily handles user interaction and the viewing / display of query results. For example, when a user is watching a match and the commentator says, "He already wasted a chance," if the user doesn't understand the commentary, they can directly press the commentary subtitle on the interface to display previous subtitles related to the athlete or team, thus helping the user quickly review previous match content while watching the live broadcast. Figure 4 As shown, the specific steps are as follows:
[0128] Step 01: The user turns on the event replay switch on the event live broadcast interface; this switch can be set at the bottom of the live broadcast screen, and there is no limitation on this.
[0129] Step 02: When users watch the live broadcast of the game, commentary subtitles are generated in real time; when users are interested in the current commentary content, or want to review related events / video content, they can press the subtitles of interest on the screen to obtain the relevant event / video content.
[0130] Step 03: Obtain query information based on video content and subtitle content. Refer to the above embodiment for details. The following query information can be obtained:
[0131] {Athlete's name, caption text, time point};
[0132] {Sports team name, subtitle text, time point};
[0133] {Match events, subtitle text, time points}.
[0134] Step 04: Based on the query information in Step 03, query different cache queues and retrieve the data in the cache queues in a paginated manner. For example, each time, retrieve the latest 20 subtitle texts from the cache queue. The retrieved content includes the subtitle text and the subtitle display time.
[0135] Step 05: Display the retrieved data on the live stream interface, i.e., display past subtitles on the live stream interface; the display method can be as follows: Figure 5 As shown, the subtitles that are about to return are displayed on the left side of the live stream screen, forming a curve.
[0136] Step 06: Users swipe to view the results, such as by swiping up and down to browse the displayed content, including the time points and detailed information of the subtitles. Displaying subtitles on the live stream interface not only allows users to easily review past matches by flipping through previous subtitles, but also doesn't interfere with watching the ongoing match. For example, the font of the subtitles can be set to semi-transparent to further enhance the user's viewing experience.
[0137] For example, the final effect would be: if a user starts watching midway through the match, pressing on a subtitle of interest, such as "Athlete 3 scores twice," the latest commentary subtitles about Athlete 3 would be displayed on the left side of the live stream interface, such as "49:34 His speed is his biggest advantage" or "49:02 Athlete 3 gets the ball, Team 2 starts attacking," etc. Users can also use the up and down swipe operation to view earlier commentary subtitles, thus enabling them to review the match content.
[0138] It should be noted that the content display method provided in this application embodiment can be executed by a content display device or a control module within that content display device for executing the content display method. This application embodiment uses the execution of the content display method by a content display device as an example to illustrate the content display device provided in this application embodiment.
[0139] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a content display device provided in an embodiment of this application. This device is applied to electronic devices, such as... Figure 6 As shown, the content display device 60 includes:
[0140] The receiving module 61 is used to receive the user's input operation on the first subtitle text, which is the subtitle text displayed on the video playback interface.
[0141] The generation module 62 is configured to generate query information related to the first subtitle text in response to the input operation, wherein the query information includes at least one target key information.
[0142] The first determining module 63 is used to determine the second subtitle text corresponding to the at least one target key information based on the correspondence between the pre-cached key information and the subtitle text;
[0143] Display module 64 is used to display the second subtitle text on the playback interface.
[0144] Therefore, based on the user's operation on the subtitles and related key information, the previous subtitle text can be determined and displayed, thereby enabling the review of video content and satisfying the user's need to obtain the required video content.
[0145] Optionally, the generation module 62 is specifically used for:
[0146] Analyze the first subtitle text to obtain the target key information contained in the first subtitle text; or,
[0147] Obtain the video frame group corresponding to the first subtitle text, and analyze each video frame in the video frame group to obtain the target key information contained in the video frame group.
[0148] Optionally, the video is a competition video, and the generation module 62 is specifically used to perform the following steps:
[0149] Identify the pixel percentage of each athlete in each video frame related to the competition video;
[0150] The appearance ratio of each athlete is determined based on the pixel percentage of each athlete in each video frame; the appearance ratio of each athlete is equal to the ratio of the sum of the pixel percentages of each athlete in the video frame group to the sum of the pixel percentages of all athletes in the video frame group.
[0151] The occurrence ratio of each athlete is normalized, and when the normalized occurrence ratio of each athlete follows a normal distribution, the preset athletes with the largest normalized occurrence ratio are confirmed as the target key information.
[0152] Optionally, the video is a competition video, and the generation module 62 is specifically used to perform the following steps:
[0153] Identify the number of first athletes and the number of second athletes for each team in the video frame group related to the competition video in the first video frame of the video frame group;
[0154] The reduction rate of each sports team is determined based on the first number of athletes and the second number of athletes;
[0155] Based on the number of the first athletes and the reduction rate of each sports team, the relevance weight of each sports team is determined, and sports teams with a relevance weight greater than a preset threshold are confirmed as the target key information.
[0156] Optionally, the video is a competition video, and the generation module 62 is specifically used to perform the following steps:
[0157] Identify the number of times the referee associated with the match video appears in the video frame group;
[0158] When the number of times meets the preset conditions, the competition events in the video frame group are confirmed as the target key information.
[0159] Optionally, the content display device 60 may also include:
[0160] The search module is used to, after obtaining the target key information contained in the video frame group, search from the head of the cache queue corresponding to the target key information to obtain the first subtitle text containing the target key information, and select multiple subtitle texts adjacent to the first subtitle text containing the target key information.
[0161] The second determining module is used to determine a first time interval based on the time corresponding to each subtitle text in the plurality of subtitle texts, wherein the first time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the plurality of subtitle texts;
[0162] The comparison module is used to compare the size of the first time interval and the second time interval, wherein the second time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the cache queue corresponding to the target key information;
[0163] The third determining module is used to determine that the target key information is related to the first subtitle text when the first time interval is less than or equal to the second time interval, and to use the target key information as the query information; or, when the first time interval is greater than the second time interval, determine that the target key information is unrelated to the first subtitle text, and to delete the target key information from the query information.
[0164] Optionally, the content display device 60 may also include:
[0165] The acquisition module is used to obtain the correspondence between the key information and the subtitle text from the server; the correspondence between the key information and the subtitle text is obtained by the server based on the analysis of the subtitle data acquired in real time and the corresponding video frame groups.
[0166] The content display device 60 of this application embodiment can implement each process of the above-described content display method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0167] Optional, such as Figure 7 As shown, this application embodiment also provides an electronic device 70, including a processor 71, a memory 72, and a program or instructions stored in the memory 72 and executable on the processor 71. When the program or instructions are executed by the processor 71, they implement the various processes of the above-described method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0168] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they can implement the various processes of the above-described method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0169] Computer-readable media include both permanent and non-permanent, removable and non-removable media, which can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0170] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0171] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0172] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a service classification device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0173] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A content display method, characterized in that, include: Receive user input for a first subtitle text, which is the subtitle text displayed on the video playback interface; In response to the input operation, query information related to the first subtitle text is generated, and the query information includes at least one target key information; Based on the correspondence between pre-cached key information and subtitle text, determine the second subtitle text corresponding to the at least one target key information; The second subtitle text is displayed on the playback interface; The generation of query information related to the first subtitle text includes: Obtain the video frame group corresponding to the first subtitle text, and analyze each video frame in the video frame group to obtain the target key information contained in the video frame group; After analyzing each video frame in the video frame group to obtain the target key information contained in the video frame group, the method further includes: Based on the target key information, start searching from the head of the cache queue corresponding to the target key information to obtain the first subtitle text containing the target key information, and select multiple subtitle texts that are adjacent to the first subtitle text containing the target key information; A first time interval is determined based on the time corresponding to each subtitle text in the plurality of subtitle texts, wherein the first time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the plurality of subtitle texts; Compare the first time interval with the second time interval, wherein the second time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the cache queue corresponding to the target key information; When the first time interval is less than or equal to the second time interval, the target key information is determined to be related to the first subtitle text, and the target key information is used as the query information; or, when the first time interval is greater than the second time interval, the target key information is determined to be unrelated to the first subtitle text, and the target key information is deleted from the query information.
2. The method according to claim 1, characterized in that, The generation of query information related to the first subtitle text also includes: The first subtitle text is analyzed to obtain the target key information contained in the first subtitle text.
3. The method according to claim 1, characterized in that, The video is a competition video. The analysis of each video frame in the video frame group to obtain the target key information contained in the video frame group includes: Identify the pixel percentage of each athlete in each video frame related to the competition video; The appearance ratio of each athlete is determined based on the pixel percentage of each athlete in each video frame; wherein, the appearance ratio of each athlete is equal to the ratio of the sum of the pixel percentages of each athlete in the video frame group to the sum of the pixel percentages of all athletes in the video frame group. The occurrence ratio of each athlete is normalized, and when the normalized occurrence ratio of each athlete follows a normal distribution, the preset athletes with the largest normalized occurrence ratio are confirmed as the target key information.
4. The method according to claim 1, characterized in that, The video is a competition video. The analysis of each video frame in the video frame group to obtain the target key information contained in the video frame group includes: Identify the number of first athletes and the number of second athletes for each team in the video frame group related to the competition video in the first video frame of the video frame group; The reduction rate of each sports team is determined based on the first number of athletes and the second number of athletes; Based on the number of the first athletes and the reduction rate of each sports team, the relevance weight of each sports team is determined, and sports teams with a relevance weight greater than a preset threshold are confirmed as the target key information.
5. The method according to claim 1, characterized in that, The video is a competition video. The analysis of each video frame in the video frame group to obtain the target key information contained in the video frame group includes: Identify the number of times the referee associated with the match video appears in the video frame group; When the number of times meets the preset conditions, the competition events in the video frame group are confirmed as the target key information.
6. The method according to claim 1, characterized in that, The method further includes: The correspondence between the key information and the subtitle text is obtained from the server; wherein the correspondence between the key information and the subtitle text is obtained by the server based on the analysis of the subtitle data acquired in real time and the corresponding video frame groups.
7. A content display device, characterized in that, include: The receiving module is used to receive user input for the first subtitle text, which is the subtitle text displayed on the video playback interface. The generation module is configured to generate query information related to the first subtitle text in response to the input operation, wherein the query information includes at least one target key information; The first determining module is used to determine the second subtitle text corresponding to the at least one target key information based on the correspondence between the pre-cached key information and the subtitle text; A display module is used to display the second subtitle text on the playback interface; Specifically, the generation module is used to: obtain the video frame group corresponding to the first subtitle text, and analyze each video frame in the video frame group to obtain the target key information contained in the video frame group; The content display device also includes: The search module is used to, after obtaining the target key information contained in the video frame group, search from the head of the cache queue corresponding to the target key information to obtain the first subtitle text containing the target key information, and select multiple subtitle texts adjacent to the first subtitle text containing the target key information. The second determining module is used to determine a first time interval based on the time corresponding to each subtitle text in the plurality of subtitle texts, wherein the first time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the plurality of subtitle texts; The comparison module is used to compare the size of the first time interval and the second time interval, wherein the second time interval is equal to the average of the time intervals between any two adjacent subtitle texts in the cache queue corresponding to the target key information; The third determining module is used to determine that the target key information is related to the first subtitle text when the first time interval is less than or equal to the second time interval, and to use the target key information as the query information; or, when the first time interval is greater than the second time interval, determine that the target key information is unrelated to the first subtitle text, and to delete the target key information from the query information.
8. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the content display method as described in any one of claims 1 to 6.
9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the content display method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Caption data pushing method, caption displaying method, device, equipment and medium
CN108600773A
Video playing method, apparatus and device, storage medium, and program product
US20230057963A1