Video definition level determination method, apparatus and device, storage medium and program product

The video duration and decision tree model are predicted through neural network models to perceive user preferences, solving the problem of inaccurate video position decisions and improving user experience and video fluency.

WO2025130602A1PCT designated stage expired Publication Date: 2025-06-26BEIJINGLUOTA INFORMATION TECHNOLOGYCO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136538
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-03
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, the server preferentially issues cached video stitches, resulting in poor accuracy of stitch decisions, and the video stitches issued do not match the actual needs of the audience, affecting the user experience.

Method used

By introducing a neural network model, the video output time index after the audience determines the target video, and effectively perceives the user's preferences and habits through the decision tree model, assists in the position decision making, and ensures that the video position matches the user's needs.

Benefits of technology

It improves the accuracy of gear decisions, improves the second output rate and reduces the lag rate, while ensuring user viewing clarity, and improving user experience and user retention rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136538_26062025_PF_FP_ABST
    Figure CN2024136538_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video definition level determination method, apparatus and device, a storage medium and a program product. The method comprises: acquiring attribute information, selectable video definition level information and historical watching information of the current viewer side; inputting the attribute information and the selectable video definition level information into a trained neural network model to obtain an output video duration index corresponding to each selectable video definition level; inputting the selectable video definition level information and the historical watching information into a pre-constructed decision tree model to obtain a preference weight of each selectable video definition level; and on the basis of the output video duration index and the preference weight that correspond to each selectable video definition level, determining a target video definition level corresponding to the current viewer side. The solution guarantees the watching definition of users, effectively senses preferences and habits of the users, and facilitates assistance in fitting a definition level decision to target requirements of the users, thereby improving the user experience and the user retention rate.
Need to check novelty before this filing date? Find Prior Art

Description

Video gear determination method, device, equipment, storage medium and program product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 20, 2023, with application number 202311764343.0, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a method, apparatus, device, storage medium, and program product for determining a video gear position. Background Art

[0003] When users watch live broadcasts, the host's source video is transcoded into a variety of different levels to accommodate different users' device performance and network conditions. Each level has different bitrates, resolutions, and encoding methods. During actual viewing, the video level is adjusted in real time based on the user's bandwidth and video cache to ensure smooth viewing and clarity.

[0004] However, in related technologies, the server tends to prioritize sending cached video levels to meet the real-time requirements of live broadcast scenarios, which in turn reduces the accuracy of level decisions. The sent video levels do not match the actual needs of the audience, affecting the user experience. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, device, storage medium, and program product for determining video gears. These methods address the issue of prioritizing the delivery of cached video gears, which results in poor gear decision accuracy and a mismatch between the delivered video gears and the actual needs of the viewer. By introducing a neural network model, the solution accurately predicts the duration of the video after the viewer determines the target video. This duration helps improve gear decision-making, increasing the playback rate within seconds and reducing the freeze rate while ensuring user viewing clarity. Furthermore, by introducing a decision tree model, users' preferences can be effectively perceived, helping gear decisions to align with the user's target needs, thereby improving user experience and user retention.

[0006] In a first aspect, an embodiment of the present application provides a method for determining a video gear position, the method comprising:

[0007] Obtain the current viewer's attribute information, available video slot information, and historical viewing information;

[0008] Inputting the attribute information and the optional video gear information into the trained neural network model to obtain a video output duration indicator corresponding to each optional video gear, wherein the video output duration indicator is used to represent the time from the current viewer end determining the target video to the target video being displayed on the screen;

[0009] Inputting the optional video gear information and the historical viewing information into a pre-built decision tree model to obtain a preference weight for each optional video gear;

[0010] The target video gear corresponding to the current viewer terminal is determined based on the video output duration index and the preference weight corresponding to each optional video gear.

[0011] In a second aspect, an embodiment of the present application further provides a device for determining a video gear position, the device comprising:

[0012] An acquisition module configured to acquire attribute information, selectable video gear information, and historical viewing information of the current viewer terminal;

[0013] a duration indicator determination module configured to input the attribute information and the optional video gear information into a trained neural network model to obtain a video output duration indicator corresponding to each optional video gear, wherein the video output duration indicator is used to represent the time from the current viewer end determining the target video to the target video being displayed;

[0014] a preference weight determination module configured to input the selectable video gear information and the historical viewing information into a pre-built decision tree model to obtain a preference weight for each selectable video gear;

[0015] The target gear determination module is configured to determine the target video gear corresponding to the current viewer terminal based on the video output duration indicator and preference weight corresponding to each optional video gear.

[0016] In a third aspect, an embodiment of the present application further provides a device for determining a video gear position, the device comprising:

[0017] one or more processors;

[0018] a storage device configured to store one or more programs,

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the video gear determination method described in the embodiment of the present application.

[0020] In a fourth aspect, an embodiment of the present application further provides a non-volatile storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are configured to execute the video gear determination method described in the embodiment of the present application.

[0021] In a fifth aspect, an embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads and executes the computer program from the computer-readable storage medium, so that the device executes the video gear determination method described in the embodiment of the present application.

[0022] In the embodiment of the present application, by obtaining the attribute information, optional video gear information and historical viewing information of the current viewer end, the attribute information and the optional video gear information are input into the trained neural network model to obtain the video output time index corresponding to each optional video gear, wherein the video output time index is used to represent the time from the current viewer end determining the target video to the target video screen display; the optional video gear information and the historical viewing information are input into a pre-built decision tree model to obtain the preference weight of each optional video gear, and the target video gear corresponding to the current viewer end is determined based on the video output time index and the preference weight corresponding to each optional video gear. In the above scheme, by introducing the neural network model, the video output time index after the viewer end determines the target video is accurately predicted. The video output time index is conducive to improving the gear decision while improving the second output rate and reducing the freeze rate, ensuring the user's viewing clarity. In addition, by introducing the decision tree model, the user's preference habits are effectively perceived, which is conducive to assisting the gear decision to meet the user's target needs, improving the user experience and user retention rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG1 is a flow chart of a method for determining a video gear position provided by an embodiment of the present application;

[0024] FIG2 is a flow chart of a method for determining a target video level corresponding to a current viewer terminal provided by an embodiment of the present application;

[0025] FIG3 is a flow chart of a method for determining a video gear position including a neural network model training process provided by an embodiment of the present application;

[0026] FIG4 is a flow chart of another method for determining a video gear position including a neural network model training process provided by an embodiment of the present application;

[0027] FIG5 is a flow chart of a method for determining a video gear position including a decision tree model building process provided by an embodiment of the present application;

[0028] FIG6 is a flowchart of a method for constructing a decision tree model provided in an embodiment of the present application;

[0029] FIG7 is a structural block diagram of a device for determining a video gear position provided by an embodiment of the present application;

[0030] FIG8 is a schematic structural diagram of a video gear determination device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the embodiments of the present application, and are not intended to limit the embodiments of the present application. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions of the embodiments of the present application, rather than all structures.

[0032] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0033] The video gear determination method provided in the embodiments of the present application can be used to adaptively determine the initial gear of a video in video output scenarios. Its application scenarios can be short videos, online live broadcasts, etc. Taking the application scenario of short videos as an example, based on the relevant information obtained from the viewer end, the appropriate initial gear can be matched for the current video stream, and video data can be sent based on the initial gear. The several application scenarios listed above are only exemplary and explanatory. In actual applications, the video gear determination method can also be used in video output in other scenarios, and the embodiments of the present application are not limited to this.

[0034] When users watch live video, the host's source video is transcoded into a variety of different levels to suit different users' device performance and network conditions. Each level has different bitrates, resolutions, encoding methods, and so on. Taking short video applications as an example, mobile phone users often use their fragmented entertainment time to browse videos, often spending only a few seconds in a live video room. If the user selects a live room through screen operation and the initial playback screen is stuck, blurry, or the video is too slow to play, it will greatly affect the user's viewing experience. Therefore, deciding on the appropriate target level is crucial.

[0035] In related technologies, ABR (Adaptive Bitrate Streaming) technology can determine the initial video level that is appropriate for the user's bandwidth. Due to the real-time requirements of live broadcast scenarios, when a user clicks or swipes to switch to a live broadcast room, the viewer will request the server to send the corresponding video stream data. The server will tend to send the level that has the cached stream, allowing direct playback without having to pull data. This can reduce the loading time caused by temporary data pulls and thus affect the user experience. Therefore, if the edge node server does not have the data stream corresponding to the initial video level cached, the server will tend to send the other video levels that are cached, resulting in lower latency for the viewer, but the level decision accuracy is poor. If the server sends a lower-definition video stream than the target video level, the viewer's video clarity will be affected, thus affecting the user's viewing experience. If the server sends a higher-definition video stream than the target video level, the viewer is more likely to experience slow video speeds or lag. In addition, different audience groups have different preferences. Some viewers have high requirements for video clarity, and can accept a few seconds of video loading time or video freezes when the viewer's device model or network conditions cannot cover the high-definition video bit rate; some viewers have high requirements for video smoothness and can accept low resolution, but cannot accept freezes or slow loading. Therefore, the decision on the target video gear needs to take into account the user's preferences to improve user retention. The embodiments of the present application provide a method, device, equipment, storage medium and program product for determining a video gear, which aims to solve the problem that the accuracy of the gear decision is reduced due to the priority of sending the cached video gear, and the sent video gear does not match the actual needs of the audience.

[0036] In the video gear determination method provided in the embodiment of the present application, the execution entity of each step can be a computer device, which refers to any electronic device with data calculation, processing and storage capabilities, such as mobile phones, PCs (Personal Computers), tablet computers and other terminal devices, or servers and other devices. The embodiment of the present application does not limit this.

[0037] FIG1 is a flow chart of a method for determining a video gear position provided by an embodiment of the present application. As shown in FIG1 , the method includes the following steps:

[0038] Step S101: Acquire the attribute information, selectable video level information, and historical viewing information of the current viewer terminal.

[0039] Among them, the attribute information of the current viewer terminal may be information related to the terminal device where the current viewer terminal is located and the network connection where it is located. The attribute information may include network information and device performance. The device performance may include device model, number of CPU cores and CPU main frequency, etc. The network information may include the region where the viewer terminal is located, network type, network bandwidth, etc. The optional video gear information of the current viewer terminal may be information related to multiple transcoding gears that can be provided to the current viewer terminal for selection. The relevant information corresponding to each transcoding gear may include bit rate information of the video stream, whether the viewer video server has cached data for the gear, whether the video proxy server has cached data for the gear, whether the host video server and the viewer video server belong to the same video server or are in the same computer room, and the round-trip delay between the host video server and the viewer video server when they are in different computer rooms, etc. The historical viewing information of the current viewer end may be the historical record information of the online video watched by the current viewer end. The historical viewing information may include the video viewing time, whether the viewer end has manually switched gears and the previous and next gears corresponding to the gear switching behavior, ABR gear switching behavior records, network bandwidth status (bandwidth size and whether the bandwidth suddenly increases or decreases), viewing video gear, viewing video bit rate, video output speed and freeze rate, etc. Of course, the above attribute information, optional video gear information and historical viewing information are for illustrative purposes only, and this application does not limit the adaptive adjustment of information selection for different application scenarios.

[0040] Step S102: Input the attribute information and the optional video gear information into the trained neural network model to obtain the video output duration index corresponding to each optional video gear, wherein the video output duration index is used to indicate the time from the current viewer end determining the target video to the target video screen display.

[0041] Among them, since users may frequently switch between live broadcast rooms when watching videos online, if the video output speed cannot adapt to the user's frequent switching, it will greatly affect the user's viewing experience. Therefore, this embodiment introduces a video output duration indicator, which is used to indicate the time from the current audience end determining the target video to the target video screen display. The higher the score corresponding to this indicator, the shorter the predicted video output duration.

[0042] In one embodiment, the video output duration can be the sum of the delay of connecting to the video server, the first packet duration, the frame duration and the playback delay. The delay of connecting to the video server can be obtained by subtracting the time when the viewer end successfully connects to the video server from the time when the user enters the room through interactive behaviors such as clicking or sliding. The first packet duration can be obtained by subtracting the time when the viewer end receives the first packet from the time when the viewer end successfully connects to the video server. The frame duration can be obtained by subtracting the time when the viewer end receives the first packet from the time when the viewer end receives the first packet. The playback delay can be obtained by subtracting the time when the first I frame is played from the time when the first I frame is composed. Taking into account the factors affecting the duration of the video, for example, the optional video levels are low-definition, high-definition and ultra-high-definition. The target video level is set to high-definition, and the audience video server does not have a cache of high-definition levels. If the audience video server sends a cached low-definition level, the audience's clarity and experience will be affected; if the audience video server sends a cached ultra-high-definition level, then since the bit rate, frame rate and resolution corresponding to the ultra-high-definition level will be higher, the framing time will be increased, thereby increasing the duration of the video; if the audience video server sends an uncached high-definition level, then temporary streaming is required. Compared with sending a cached video level, the first packet time will be increased, thereby increasing the duration of the video. Furthermore, the impact of temporary streaming on first-packet duration is related to the CDN (Content Delivery Network). If the viewer's video server doesn't have a corresponding video stream, pulling the stream from the proxy server will incur a delay of several milliseconds if the proxy server does. If the proxy server doesn't have a corresponding video stream, pulling the stream from the host's video server will require round-trip latency between the server room where the host and viewer servers are located. Therefore, a reasonable prediction of the video duration indicator requires combining multiple attribute information and information about available video slots.

[0043] In one embodiment, a neural network model can be used to predict the video output duration index corresponding to each optional video gear according to the input attribute information and the optional video gear information. The neural network model can be CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network) and DNN (Deep Neural Network), etc., which are not limited in this application. Taking the DNN model as an example, the DNN model can be composed of an input layer, multiple hidden layers and an output layer. Each hidden layer contains multiple neurons or nodes. The output features of the previous layer can be used as the input of the next layer for feature learning. Through the feature transformation of multiple nonlinear mappings, the nonlinear correspondence between the attribute information and the optional video gear information and the video output duration index can be effectively fitted. Optionally, the neural network model can use the viewer's bandwidth a1, the viewer's device performance b1, b2, b3..., the video stream bit rate c1 of the optional video gear, and the round-trip delay d1 between the host video server and the viewer video server to learn and fit the prediction of the video length indicator. The fitting relationship learned by the neural network model is as follows: Q = q1 × a1 + q 21 ×b1+q 22 ×b2+q 23 ×b3+…-q3×c1-q4×d1,

[0044] Among them, q1>0, q 21 >0,q 22 >0,q 23 >0, ..., q3>0, q4>0, where Q is the video output duration indicator. The higher the score of the video output duration indicator, the shorter the predicted video output duration. Of course, this fitting relationship is an exemplary description. The form, fitting coefficient, and fitting method of other fitting relationships can be adjusted based on actual scenario requirements and are not limited in this application.

[0045] Step S103: Input the optional video gear information and historical viewing information into a pre-built decision tree model to obtain a preference weight for each optional video gear.

[0046] Among them, in the online video scenario, some viewers prefer high-resolution videos and can accept slow video speed or video freeze, while some viewers prefer high-smoothness videos and can accept low definition but do not want to encounter problems such as freeze, long loading time or frame loss. Therefore, this embodiment introduces preference weights, which can be user preference coefficients corresponding to each optional video gear setting, used to reflect the audience's preference for each optional video gear.

[0047] In one embodiment, the decision tree model can be constructed based on algorithms such as the ID3 algorithm, the C4.5 algorithm, and the Cart algorithm, which are not limited in this application. Taking the ID3 algorithm as an example, the idea of ​​constructing a decision tree can be to use the user's viewing time as the first attribute feature, divide the data set into two parts based on the set viewing time threshold, and then calculate the information gain rate of the remaining attribute features in each part respectively, and continue to select the feature corresponding to the maximum information gain rate as the partition feature, and so on, until there is no feature partition or the labels of the partition features are all one category.

[0048] In one embodiment, after the decision tree model is fully constructed, the non-leaf nodes can be examined from the bottom up. If the subtree corresponding to the non-leaf node is replaced with a leaf node, it can bring about an improvement in generalization performance. In this way, the subtree is replaced with a leaf node, thereby reducing the problem of overfitting of the decision tree model and improving the generalization ability of the model.

[0049] Step S104: Determine the target video gear corresponding to the current viewer terminal based on the video output duration indicator and the preference weight corresponding to each optional video gear.

[0050] The target video gear can be a video gear determined by comprehensively analyzing various user attribute information and preference information. Optionally, FIG2 is a flow chart of a method for determining the target video gear corresponding to the current viewer terminal provided by an embodiment of the present application, as shown in FIG2, including the following steps:

[0051] Step S1041: Determine a candidate video level that meets the bandwidth requirement of the current viewer terminal from the optional video levels based on an adaptive bit rate algorithm.

[0052] Among them, the adaptive bit rate algorithm can be an algorithm that monitors the network status of the current audience end and determines the video gear that meets the bandwidth conditions of the current audience end from the optional video gears. The adaptive bit rate algorithm can be a DASH algorithm, HLS algorithm, etc., which is not limited in this application.

[0053] Step S1042: Calculate the corresponding position score of each candidate video position according to the video output duration indicator and the preference weight corresponding to the position.

[0054] Among them, the gear score corresponding to each candidate video gear can be a reflection of the degree of matching between the candidate video gear and the current audience end. The higher the gear score, the higher the degree of matching between the candidate video gear and the current audience end. Optionally, the video duration index and the preference weight can be normalized and then summed to obtain the gear score. Optionally, a weighted calculation coefficient can be set for the video duration index and the preference weight respectively, and the calculation coefficient can be obtained based on the statistical results of the online test. Of course, there can also be other ways to calculate the gear score based on the video duration index and the preference weight, which are not limited in this application.

[0055] Step S1043: Determine the video gear with the highest gear score from the candidate video gears as the target video gear.

[0056] Therefore, the adaptive bit rate algorithm is used to preliminarily screen the optional video gears in advance to meet the current network conditions of the audience, reducing the amount of data calculation for determining the target video gear in the subsequent process and improving the gear determination efficiency. By integrating the multi-factor effects of video duration indicators and preference weights, the appropriate video gear is accurately allocated to the current audience, providing users with the video gear with the best smoothness and clarity in scenarios with network fluctuations.

[0057] As can be seen from the above, by obtaining the attribute information, optional video gear information and historical viewing information of the current viewer end; inputting the attribute information and optional video gear information into the trained neural network model, the corresponding video output time index of each optional video gear is obtained, wherein the video output time index is used to represent the time from the current viewer end determining the target video to the target video screen display; inputting the optional video gear information and historical viewing information into the pre-built decision tree model, the preference weight of each optional video gear is obtained; based on the corresponding video output time index and preference weight of each optional video gear, the target video gear corresponding to the current viewer end is determined. In the above scheme, by introducing the neural network model, the video output time index after the viewer end determines the target video is accurately predicted. The video output time index is conducive to improving the gear decision while increasing the second output rate and reducing the freeze rate, ensuring the user's viewing clarity. In addition, by introducing the decision tree model, the user's preference habits are effectively perceived, which is conducive to assisting the gear decision to meet the user's target needs, improving the user experience and user retention rate.

[0058] FIG3 is a flow chart of a method for determining a video gear position including a neural network model training process provided by an embodiment of the present application. As shown in FIG3 , the method includes the following steps:

[0059] Step S201: Acquire the attribute information, selectable video level information, and historical viewing information of the current viewer terminal.

[0060] Step S202: Acquire model training data, wherein the model training data includes attribute information of the sample viewer terminal and transcoding gear information.

[0061] Among them, the attribute information of the sample viewer terminal may be information related to the terminal device where the sample viewer terminal is located and the network connection. The attribute information may include network information and device performance. The device performance may include device model, number of CPU cores and CPU main frequency, etc. The network information may include the region where the viewer terminal is located, network type, network bandwidth, etc. The transcoding gear information of the sample viewer terminal may be relevant information of multiple transcoding gears that can be provided to the sample viewer terminal for selection. The relevant information corresponding to each transcoding gear may include bit rate information of the video stream, whether the viewer video server has cached data for the gear, whether the video proxy server has cached data for the gear, whether the host video server and the viewer video server belong to the same video server or are in the same computer room, and the round-trip delay between the host video server and the viewer video server when they are in different computer rooms.

[0062] Step S203: training the neural network model based on the model training data.

[0063] Before model training, the model training data can be preprocessed, including but not limited to normalization, standardization, filtering, and denoising, to improve the model training effect. Then, based on the video duration prediction problem, an appropriate neural network structure can be selected and the weights and biases of the neural network can be initialized. Initialization methods can include random initialization, Xavier initialization, He initialization, etc. Next, an appropriate loss function and gradient descent optimization algorithm are selected. The loss function can include cross-entropy loss, mean error, etc., and the optimization algorithm can include stochastic gradient descent (SGD), Adam algorithm, RMSprop algorithm, etc. During model training using the model training data, each training iteration passes the input data to the model, calculates the loss, and then updates the model parameters through backpropagation, gradually reducing the loss until it stabilizes, completing model training. Furthermore, the trained model can be evaluated using a validation set or a test set. Evaluation metrics can include accuracy, precision, recall, F1 score, etc.

[0064] Step S204: input the attribute information and the optional video gear information into the trained neural network model to obtain the video output duration indicator corresponding to each optional video gear, wherein the video output duration indicator is used to indicate the time length from the current viewer end determining the target video to the target video screen display.

[0065] Step S205: input the optional video gear information and historical viewing information into a pre-built decision tree model to obtain a preference weight for each optional video gear;

[0066] Step S206: Determine the target video level corresponding to the current viewer terminal based on the video output duration indicator and the preference weight corresponding to each optional video level.

[0067] As mentioned above, by collecting attribute information of the sample audience end and transcoding gear information to train the neural network model, a reliable model for accurately predicting the video output duration indicator of the current audience end can be obtained, providing a strong reference for the decision of the target video gear.

[0068] FIG4 is a flow chart of another method for determining a video gear position including a neural network model training process provided by an embodiment of the present application. As shown in FIG4 , the method includes the following steps:

[0069] Step S301: Acquire the attribute information, selectable video level information, and historical viewing information of the current viewer terminal.

[0070] Step S302: Acquire model training data, wherein the model training data includes attribute information of the sample viewer terminal and transcoding gear information.

[0071] Step S303: Divide the model training data into a plurality of regional training data based on the set geographical partitions.

[0072] Because live video streaming tends to be regionally concentrated, user terminals, network equipment deployment, and viewing habits may vary significantly across regions. Therefore, model training data can be divided based on geographic regions to train neural network models tailored to the specific regions. For example, the geographic regions could be South America, North America, Europe, and Asia. This application does not limit the granularity of geographic regions, and adaptive adjustments can be made based on different application scenarios.

[0073] Step S304: training the corresponding neural network model based on the regional training data corresponding to each geographical partition.

[0074] Step S305: input the current attribute information and the optional video gear information into the trained neural network model corresponding to the geographical partition where the current viewer terminal is located, and obtain the video output duration indicator corresponding to each optional video gear, wherein the video output duration indicator is used to indicate the time length from the current viewer terminal determining the target video to the target video screen display.

[0075] Step S306: Input the selectable video gear information and the historical viewing information into a pre-built decision tree model to obtain a preference weight for each selectable video gear.

[0076] Step S307: Determine the target video gear corresponding to the current viewer terminal based on the video output duration indicator and the preference weight corresponding to each optional video gear.

[0077] As mentioned above, by dividing the model training data according to different geographical divisions and training the neural network model corresponding to each geographical division, it can effectively adapt to the regional aggregation characteristics and provide highly matched neural network models for audiences in different geographical divisions, while improving the model convergence and enhancing the accuracy of the prediction results.

[0078] FIG5 is a flowchart of a method for determining a video gear position including a decision tree model building process provided by an embodiment of the present application. As shown in FIG5 , the method includes the following steps:

[0079] Step S401: Acquire the attribute information, selectable video level information, and historical viewing information of the current viewer terminal.

[0080] Step S402: Input the attribute information and the optional video gear information into the trained neural network model to obtain the video output duration indicator corresponding to each optional video gear, wherein the video output duration indicator is used to indicate the time from the current viewer end determining the target video to the target video screen display.

[0081] Step S403: Acquire model feature data, wherein the model feature data includes viewer-side feature data and viewer-side historical data.

[0082] The viewer-side feature data may include the viewer's region, terminal model, network type, etc. The viewer-side historical data may include video viewing time, whether the viewer has manually switched gears and the previous and next gears corresponding to the gear switching behavior, ABR gear switching behavior records, network bandwidth status (bandwidth size and whether the bandwidth suddenly increases or decreases), video viewing gear, video viewing bit rate, video output speed, and freeze rate, etc.

[0083] Step S404: clean the model feature data to obtain model construction data.

[0084] Among them, data cleaning can include missing value processing, outlier processing, duplicate value processing and redundant data removal, which is not limited in this application.

[0085] Step S405: construct a decision tree model based on the model construction data.

[0086] In one embodiment, FIG6 is a flowchart of a method for constructing a decision tree model provided in an embodiment of the present application. As shown in FIG6 , the method includes the following steps:

[0087] Step S4051: Calculate the information gain of the features to be determined for the root node, and determine the feature with the largest information gain from the features to be determined as the corresponding node feature.

[0088] The information gain calculation may be performed by calculating the entropy of the parent node and the difference between the entropy of the weighted child nodes of each feature to be determined, so as to determine the feature with the largest information gain as the splitting feature of the node.

[0089] Step S4052: Create corresponding child nodes based on different values ​​of corresponding node features until the stopping condition for completing the building of the decision tree model is reached.

[0090] Among them, the stopping condition for the completion of the decision tree model construction can be that the information gain of all features is less than the set threshold or there are no new features to be selected.

[0091] Step S406: Input the selectable video gear information and the historical viewing information into a pre-built decision tree model to obtain a preference weight for each selectable video gear.

[0092] Step S407: Determine the target video gear corresponding to the current viewer terminal based on the video output duration indicator and the preference weight corresponding to each optional video gear.

[0093] As mentioned above, by collecting viewer-side feature data and viewer-side historical data to build a decision tree model, we can effectively learn users' preferences and habits, accurately predict the current viewer's preference weight, and provide a reliable reference for the decision of the target video gear.

[0094] FIG7 is a block diagram of a video gear determination device provided in an embodiment of the present application. The device is configured to execute the video gear determination method provided in the above embodiment and has the corresponding functional modules and beneficial effects of the execution method. As shown in FIG7 , the device includes:

[0095] The acquisition module 101 is configured to acquire the attribute information, selectable video gear information and historical viewing information of the current viewer terminal;

[0096] The duration indicator determination module 102 is configured to input the attribute information and the optional video gear information into the trained neural network model to obtain the video output duration indicator corresponding to each optional video gear. The video output duration indicator is used to represent the time from the current viewer end determining the target video to the target video screen display;

[0097] The preference weight determination module 103 is configured to input the optional video gear information and the historical viewing information into a pre-built decision tree model to obtain the preference weight of each optional video gear;

[0098] The target gear determination module 104 is configured to determine the target video gear corresponding to the current viewer terminal based on the video output duration indicator and the preference weight corresponding to each optional video gear.

[0099] In the above, by obtaining the attribute information, optional video gear information and historical viewing information of the current viewer end; inputting the attribute information and optional video gear information into the trained neural network model, the video output duration index corresponding to each optional video gear is obtained, wherein the video output duration index is used to represent the time from the current viewer end determining the target video to the target video screen display; inputting the optional video gear information and historical viewing information into a pre-built decision tree model, the preference weight of each optional video gear is obtained; based on the video output duration index and preference weight corresponding to each optional video gear, the target video gear corresponding to the current viewer end is determined. In the above scheme, by introducing the neural network model, the video output duration index after the viewer end determines the target video is accurately predicted. The video output duration index is conducive to improving the gear decision while increasing the second output rate and reducing the freeze rate, ensuring the user's viewing clarity. In addition, by introducing the decision tree model, the user's preference habits are effectively perceived, which is conducive to assisting the gear decision to meet the user's target needs, improving the user experience and user retention rate.

[0100] In a possible embodiment, a model training module is further included, configured to:

[0101] Obtain model training data, which includes attribute information of sample viewers and transcoding gear information;

[0102] Train the neural network model based on the model training data.

[0103] In one possible embodiment, the model training module is further configured to:

[0104] Divide the model training data into multiple regional training data based on the set geographical partitions;

[0105] Training the corresponding neural network model based on the regional training data corresponding to each geographical division;

[0106] Accordingly, the duration indicator determination module 102 is further configured to:

[0107] The current attribute information and the optional video gear information are input into the trained neural network model corresponding to the geographical partition where the current viewer terminal is located.

[0108] In a possible embodiment, a decision tree building module is further included, configured to:

[0109] Obtaining model feature data, the model feature data including viewer-side feature data and viewer-side historical data;

[0110] Perform data cleaning on the model feature data to obtain model construction data;

[0111] Build a decision tree model based on the model building data.

[0112] In a possible embodiment, the decision tree construction module is further configured to:

[0113] Calculate the information gain of the features to be determined for the root node, and determine the feature with the largest information gain from the features to be determined as the corresponding node feature;

[0114] The corresponding child nodes are established based on the different values ​​of the corresponding node features until the stopping condition for the completion of the decision tree model is reached.

[0115] In a possible embodiment, the target gear determination module 104 is further configured to:

[0116] Determine a candidate video gear that meets the bandwidth conditions of the current viewer terminal from the optional video gears based on an adaptive bit rate algorithm;

[0117] Calculate the corresponding position score based on the video length index and preference weight corresponding to each candidate video position;

[0118] The video gear with the highest gear score is determined as the target video gear from the candidate video gears.

[0119] Figure 8 is a schematic diagram of the structure of a video gear determination device provided in an embodiment of the present application. As shown in Figure 8, the device includes a processor 201, a memory 202, an input device 203, and an output device 204. The device may contain one or more processors 201, with Figure 8 citing a single processor 201 as an example. The processor 201, memory 202, input device 203, and output device 204 may be connected via a bus or other means, with Figure 8 citing a bus connection as an example. Memory 202, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the video gear determination method in the embodiment of the present application. Processor 201 executes the software programs, instructions, and modules stored in memory 202 to execute various functional applications and data processing of the device, thereby implementing the aforementioned video gear determination method. Input device 203 can be configured to receive input digital or character information and generate key signal input related to user settings and function control of the device. Output device 204 may include a display device such as a display screen.

[0120] An embodiment of the present application also provides a non-volatile storage medium containing computer-executable instructions, which, when executed by a computer processor, is configured to execute a video gear determination method described in the above embodiment, which includes: obtaining attribute information, optional video gear information and historical viewing information of the current viewer end; inputting the attribute information and optional video gear information into a trained neural network model to obtain a video output duration indicator corresponding to each optional video gear, where the video output duration indicator is used to indicate the time from the target video determined by the current viewer end to the screen display of the target video; inputting the optional video gear information and historical viewing information into a pre-built decision tree model to obtain a preference weight for each optional video gear; and determining the target video gear corresponding to the current viewer end based on the video output duration indicator and the preference weight corresponding to each optional video gear.

[0121] It is worth noting that in the embodiment of the above-mentioned video gear determination device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not configured to limit the scope of protection of the embodiments of this application.

[0122] In some possible implementations, various aspects of the methods provided herein may also be implemented in the form of a program product, which includes program code. When the program product is executed on a computer device, the program code is configured to cause the computer device to execute the steps of the methods according to the various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the video gear determination method described in the embodiments of the present application. The program product may be implemented using any combination of one or more readable media.

Claims

1. A method for determining a video gear position, wherein: include: Obtain the current viewer's attribute information, available video slot information, and historical viewing information; Input the attribute information and the optional video gear information into the trained neural network model to obtain a video output time index corresponding to each optional video gear, wherein the video output time index is used to indicate the time from the current viewer end determining the target video to the target video screen display; Inputting the selectable video position information and the historical viewing information into a pre-built decision tree model to obtain a preference weight for each selectable video position; The target video gear corresponding to the current viewer terminal is determined based on the video output duration index and the preference weight corresponding to each optional video gear.

2. The method for determining the video gear position according to claim 1, wherein: Before inputting the attribute information and the optional video gear information into the trained neural network model, the method further includes: Acquire model training data, wherein the model training data includes attribute information of a sample viewer terminal and transcoding gear information; The neural network model is trained based on the model training data.

3. The method for determining the video gear position according to claim 2, wherein: The training of the neural network model based on the model training data includes: Dividing the model training data into a plurality of regional training data based on set geographical partitions; Training a corresponding neural network model based on the regional training data corresponding to each of the geographical divisions; Accordingly, the step of inputting the attribute information and the optional video gear information into the trained neural network model includes: The current attribute information and the optional video gear information are input into the trained neural network model corresponding to the geographical partition where the current audience terminal is located.

4. The method for determining the video gear position according to any one of claims 1 to 3, wherein: Before inputting the selectable video position information and the historical viewing information into the pre-built decision tree model, the method further includes: Acquire model feature data, wherein the model feature data includes viewer-end feature data and viewer-end historical data; Performing data cleaning on the model feature data to obtain model building data; A decision tree model is constructed based on the model construction data.

5. The method for determining the video gear position according to claim 4, wherein: The construction of a decision tree model based on the model construction data includes: Calculate the information gain of the feature to be determined for the root node, and determine the feature with the largest information gain from the features to be determined as the corresponding node feature; Corresponding child nodes are established based on different values ​​of the corresponding node features until the stopping condition for completing the building of the decision tree model is reached.

6. The method for determining the video gear position according to any one of claims 1 to 5, wherein: The determining of the target video position corresponding to the current viewer terminal based on the video output duration index and the preference weight corresponding to each optional video position includes: Determine a candidate video gear that meets the bandwidth condition of the current viewer terminal from the optional video gears based on an adaptive bit rate algorithm; Calculate the corresponding position score according to the video output duration index and the preference weight corresponding to each candidate video position; The video gear with the highest gear score among the candidate video gears is determined as the target video gear.

7. A device for determining a video gear position, wherein: include: An acquisition module configured to acquire attribute information, selectable video position information, and historical viewing information of the current viewer terminal; A duration index determination module is configured to input the attribute information and the optional video gear information into the trained neural network model to obtain a video duration index corresponding to each optional video gear, wherein the video duration index is used to indicate the time from the current viewer end determining the target video to the screen display of the target video; A preference weight determination module, configured to input the selectable video gear information and the historical viewing information into a pre-built decision tree model to obtain a preference weight for each of the selectable video gears; The target gear determination module is configured to determine the target video gear corresponding to the current viewer terminal based on the video output duration indicator and the preference weight corresponding to each optional video gear.

8. A video gear determination device, the device comprising: one or more processors; A storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, enables the one or more processors to implement the video gear determination method described in any one of claims 1-6.

9. A non-volatile storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the video gear determination method according to any one of claims 1 to 6 when executed by a computer processor.

10. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the video gear determination method described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Dynamic image quality video playing method and device, electronic equipment and storage medium

    CN115776590A

  • Method for scheduling system resources and related device

    CN116077943A

  • Video code rate determination method and device, electronic equipment and storage medium thereof

    CN116634231A

  • Video gear adjustment method and device, electronic equipment and storage medium

    CN117135375A

  • Video gear determination method and device, equipment, storage medium and program product

    CN117915154A