Information processing method and device, storage medium and computer program product

CN121509705APending Publication Date: 2026-02-10MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511500226.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-02-10

Smart Images

  • Figure CN121509705A_ABST
    Figure CN121509705A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information processing method, and the method comprises the steps: obtaining a transmission bandwidth needed by a to-be-played video, and determining a target receiving bandwidth of a client within a target duration from a current moment; if the target receiving bandwidth does not meet the transmission bandwidth, dividing the image of the video to be played to obtain a plurality of groups of images; processing each image of the plurality of groups of images by adopting a target large model to obtain description information of each image; and sending the description information and the target information to the client, so that the client generates the to-be-played video based on the description information and the target information, thereby solving the problems of relatively poor video quality and relatively low video playing fluency in a video playing process in related technologies. The embodiment of the invention further discloses information processing equipment, a storage medium and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to an information processing method, device, storage medium, and computer program product. Background Technology

[0002] During video playback, the client's network conditions directly impact the user's viewing experience. Poor network conditions can cause video playback stuttering. To address this, technologies typically employ methods such as reducing the video bitrate or prompting the user to switch networks to ensure normal playback. However, while these methods alleviate playback interruptions to some extent, they significantly degrade video quality and affect playback smoothness, resulting in a poor viewing experience for the user. Summary of the Invention

[0003] To address the aforementioned technical problems, this application aims to provide an information processing method, device, storage medium, and computer program product that solves the problems of poor video quality and low smoothness during video playback in related technologies.

[0004] To achieve the above objectives, the technical solution of this application embodiment is implemented as follows: An information processing method, the method comprising: Obtain the transmission bandwidth required for the video to be played, and determine the target receiving bandwidth of the client within the target duration starting from the current moment; If the target receiving bandwidth does not meet the transmission bandwidth, the images of the video to be played are divided to obtain multiple sets of images; A target large model is used to process each image in the multiple sets of images to obtain descriptive information for each image; wherein, the descriptive information is used to explain the content of each image; The description information and target information are sent to the client so that the client can generate the video to be played based on the description information and target information; wherein, the target information characterizes the content features of the video to be played.

[0005] In the above scheme, determining the target receiving bandwidth of the client within the target duration from the current moment includes: Obtain the client's first real-time performance information and the client's network's second real-time performance information; Obtain the third real-time performance information of the target server and the fourth real-time performance information of the target server's network; Obtain the fifth real-time performance information corresponding to the transmission path between the client and the target server; The first real-time performance information, the second real-time performance information, the third real-time performance information, the fourth real-time performance information, and the fifth real-time performance information are processed using a target bandwidth prediction model to obtain the target receiving bandwidth.

[0006] In the above scheme, sending the description information and target information to the client includes: Determine the maximum receiving bandwidth of the client within the target duration; The first bandwidth required for the description information, the second bandwidth required for the audio information of the video to be played, the third bandwidth required for the control information of the target transmission protocol, and the fourth bandwidth required for the first image are determined; wherein, the first image is an image of a keyframe in the video to be played; Based on the first bandwidth, the second bandwidth, the third bandwidth, and a plurality of fourth bandwidths, a first target bandwidth is determined; If the maximum receiving bandwidth does not meet the first target bandwidth, a second target bandwidth is determined based on the first bandwidth, the second bandwidth, and the third bandwidth; If the maximum receiving bandwidth meets the second target bandwidth, the description information and the audio information are sent to the client; wherein, the target information includes the audio information; If the maximum receiving bandwidth meets the first target bandwidth, a target image is determined from multiple first images, and the description information, the audio information, and the target image are sent to the client; wherein, the target information further includes the target image.

[0007] In the above scheme, determining the first target bandwidth based on the first bandwidth, the second bandwidth, the third bandwidth, and multiple fourth bandwidths includes: Based on the size of each first image, a third target bandwidth is determined from the plurality of fourth bandwidths; The first target bandwidth is determined based on the first bandwidth, the second bandwidth, the third bandwidth, and the third target bandwidth.

[0008] In the above scheme, determining the target image from multiple first images includes: The plurality of first images are sorted based on the time of each first image in the video to be played, to obtain the arrangement order of the plurality of first images; According to the arrangement order, the similarity between each second image and the next frame image of each second image is determined; wherein, the second image is the image among the plurality of first images excluding the last frame first image; The target image is determined from the plurality of first images based on multiple similarity scores.

[0009] The method in the above scheme further includes: The sample images of the sample video are divided into multiple sets of sample images; First sample description information of a first target sample image is determined; wherein, the first sample description information is used to explain the content of the first target sample image; the first target sample image is an image of a keyframe in the sample video; Based on the first sample description information, second sample description information of other sample images in each group of sample images is determined; wherein, the other sample images are sample images in each group of sample images other than the first target sample image; Determine the second target sample image from multiple first target sample images; The target large model is obtained by training the initial large model based on the sample image, the second target sample image, the first sample description information, and the second sample description information.

[0010] An information processing method, the method further comprising: The system receives description information and target information sent by the target server; wherein, the target information characterizes the content features of the video to be played; the description information is obtained by the target server after processing each image of multiple sets of images, and is used to explain the content of each image; the multiple sets of images are obtained by dividing the video to be played into images when the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played; Based on the target large model, the description information, and the target information, multiple images are generated; The video to be played is generated based on the multiple images and the target information.

[0011] In the above scheme, the generation of multiple images based on the target large model, the description information, and the target information includes: If the target information does not include the target image, the description information is processed using the target large model to generate the multiple images; wherein, the target image is determined by the target server from the first image of the keyframe in the video to be played; If the target information includes the target image, the target large model is used to process the description information and the target image to obtain the multiple images.

[0012] In the above scheme, generating the video to be played based on the multiple images and the target information includes: The audio information of the video to be played is determined from the target information; The video to be played is generated based on the multiple images and the audio information.

[0013] An information processing device, the information processing device comprising: a processor, a memory, and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the information processing program stored in the memory to implement the steps of the above-described information processing method.

[0014] A storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described information processing method.

[0015] A computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described information processing method.

[0016] The information processing method, device, storage medium, and computer program product provided in this application embodiment can obtain the transmission bandwidth required for a video to be played, determine the target receiving bandwidth of the client within a target duration from the current moment, and when the target receiving bandwidth does not meet the transmission bandwidth, divide the images of the video to be played into multiple groups of images. Then, a target large model is used to process each image in the multiple groups of images to obtain description information for each image, and the description information and target information are sent to the client, so that the client generates the video to be played based on the description information and the target information; thus, the target receiving bandwidth of the client within a future period of time is determined. When the received bandwidth is insufficient to meet the transmission bandwidth required for the video to be played, the description information of each image of the video to be played can be determined in advance using a target large model and sent to the client. This allows the client to reconstruct the video to be played based on the description information and target information. In other words, even if the client experiences network lag in the future, it can still reconstruct the video to be played and play it based on the received description information and target information, instead of needing to compress the video bitrate or switch networks to ensure normal video playback as in related technologies. This solves the problems of poor video quality and low smoothness in video playback in related technologies, thereby improving the user's viewing experience. Attached Figure Description

[0017] Figure 1 A flowchart illustrating an information processing method provided in an embodiment of this application; Figure 2 A flowchart illustrating another information processing method provided in an embodiment of this application; Figure 3 A flowchart illustrating another information processing method provided in an embodiment of this application; Figure 4 A flowchart illustrating the determination of target transmission bandwidth in an information processing method provided in this application embodiment; Figure 5 A schematic diagram illustrating the training process of a target bandwidth prediction model in an information processing method provided in this application embodiment; Figure 6 A schematic diagram illustrating the training process of a target bandwidth prediction model in an information processing method provided in this application embodiment; Figure 7 A schematic diagram illustrating the training process of a target large model in an information processing method provided in this application embodiment; Figure 8 This application provides a schematic diagram of the optimization process of a target large model in an information processing method according to an embodiment of the present application. Figure 9 This is a schematic flowchart illustrating the image reconstruction process in an information processing method provided in an embodiment of this application. Figure 10 A schematic flowchart illustrating another image reconstruction method provided in an embodiment of this application; Figure 11 A schematic flowchart illustrating another image reconstruction method provided in an embodiment of this application; Figure 12 This is a schematic diagram of the structure of a first information processing device provided in an embodiment of this application; Figure 13 This is a schematic diagram of the structure of a second information processing device provided in an embodiment of this application; Figure 14 This is a schematic diagram of the structure of a target server provided in an embodiment of this application; Figure 15 This is a schematic diagram of the structure of a client provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0019] It should be understood that the phrases "embodiments of this application" or "foreign embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "embodiments of this application" or "in the foreign embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0020] Unless otherwise specified, any step in the embodiments of this application performed by the electronic device may be executed by the processor of the electronic device. It is also worth noting that the embodiments of this application do not limit the order in which the electronic device performs the following steps. Furthermore, the methods used to process data in different embodiments may be the same or different methods. It should also be noted that any step in the embodiments of this application can be executed independently by the electronic device; that is, when the electronic device performs any step in the following embodiments, it may not depend on the execution of other steps.

[0021] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0022] This application provides a method that can be applied to a target server, as shown in the following embodiments. Figure 1 As shown, the method may include the following steps: Step 101: Obtain the transmission bandwidth required for the video to be played, and determine the target receiving bandwidth of the client within the target duration starting from the current moment.

[0023] In this embodiment of the application, the video to be played can refer to the video that the client is about to play, which can be a TV series video or a game video; the transmission bandwidth can refer to the bandwidth required to transmit the video to be played from the target server to the client; the target receiving bandwidth can refer to the amount of data that the client can receive within the target duration from the current moment.

[0024] In this embodiment of the application, the resolution and frame rate of the video to be played can be obtained, and the resolution and frame rate of the video to be played can be calculated to obtain the bit rate (i.e., transmission bandwidth) of the video to be played. Then, a bandwidth prediction model can be used to determine the target receiving bandwidth of the client within the target duration.

[0025] Step 102: If the target receiving bandwidth does not meet the transmission bandwidth, the images of the video to be played are divided into multiple sets of images.

[0026] In this embodiment of the application, the target receiving bandwidth and the transmission bandwidth can be compared. If the target receiving bandwidth does not meet the transmission bandwidth, it means that the network condition of the client will be poor in the future. At this time, the images of the video to be played can be divided according to the preset Group of Pictures (GOP) length to obtain multiple image groups (multiple GOPs), that is, multiple sets of images.

[0027] It should be noted that the number of images in each group of images (i.e., each GOP) is the same, and each GOP includes one Intra-coded Frame (I-frame) image, multiple Predictive-coded Frame (P-frame) images, and multiple Bidirectional Predicted Frame (B-frame) images.

[0028] Step 103: Use the target large model to process each image in multiple sets of images to obtain the descriptive information of each image.

[0029] The descriptive information is used to explain the content of each image.

[0030] In this embodiment, the target large model can refer to a multimodal model. Specifically, each image from multiple sets of images (i.e., multiple Groups of Pictures) can be input as an input parameter into the target large model. Then, the target large model can process each image to obtain descriptive information for each frame of each GOP. The multimodal model can be a text-center large model, a VARGPT large model, or other large models; no specific limitation is made here.

[0031] In one feasible approach, if the video to be played is a match video, then the image of the video to be played is the match image, and the image description information can be "xx is passing the ball, xxx is about to intercept it"; if the video to be played is a video of a dog running, then the corresponding description information can be "a Pomeranian is running on green grass".

[0032] Step 104: Send description information and target information to the client so that the client can generate a video to be played based on the description information and target information.

[0033] Among them, target information represents the content features of the video to be played.

[0034] In the embodiments of this application, the target information may include the audio information of the video to be played. In one possible implementation, the target information may also include some keyframe images of the video to be played.

[0035] In this embodiment of the application, after the client obtains the description information of each image, it can encapsulate the description information and target information and send the encapsulated information to the client. After receiving the encapsulated information, the client can decapsulate the information to obtain the description information and target information. Then, the client can reconstruct multiple images based on the description information and target information using a target large model, and then generate a video to be played based on the reconstructed images and target information, that is, reconstruct the video to be played and play it.

[0036] It should be noted that if, at some point in the future, the target server determines that the client's target receiving bandwidth meets the transmission bandwidth, the target server can directly send the video to be played to the client according to the normal process, without performing operations such as image segmentation.

[0037] In this embodiment, the description information preserves the key visual semantics of the image. Thus, even in a poor network environment on the client side, the client can still reconstruct a high-quality video frame in advance based on the received description information and target information, and then stitch the video frame together to reconstruct the video to be played, thereby achieving smooth video playback under low bandwidth conditions.

[0038] The information processing method provided in this application embodiment can pre-determine the description information of each image of the video to be played by using a target large model in advance and send it to the client when the target receiving bandwidth of the client in the future does not meet the transmission bandwidth required by the video to be played. This allows the client to reconstruct the video to be played based on the description information and target information. In other words, even if the client experiences network lag in the future, it can still reconstruct the video to be played and play it based on the received description information and target information, instead of needing to compress the video bitrate or switch networks to ensure normal video playback as in related technologies. This solves the problems of poor video quality and low smoothness of video playback in related technologies, thereby improving the user's viewing experience.

[0039] Based on the foregoing embodiments, embodiments of this application provide an information processing method that can be applied to a client, as shown below. Figure 2 As shown, the method may include the following steps: Step 201: Receive the description information and target information sent by the target server.

[0040] The description information is obtained by the target server after processing each image in the multiple sets of images, and is used to explain the content of each image; the multiple sets of images are obtained by dividing the images of the video to be played when the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played; the target receiving bandwidth is the receiving bandwidth of the client within the target duration from the current moment; the target information characterizes the content features of the video to be played.

[0041] In this embodiment, the client can refer to a video playback terminal. Specifically, the client can receive description information and target information sent by the target server through a specified transmission protocol.

[0042] Step 202: Generate multiple images based on the target large model, description information, and target information.

[0043] In this embodiment of the application, it is possible to determine whether a target image exists in the target information, and generate (i.e. reconstruct) multiple images based on the determination result, description information and target information.

[0044] In this embodiment of the application, even if the client's network conditions are poor for a period of time in the future, it can still reconstruct the image of the video to be played based on the description information, thereby ensuring that the video to be played can be played normally.

[0045] Step 203: Generate a video to be played based on multiple images and target information.

[0046] In this embodiment of the application, the target information can be parsed to obtain the audio information of the video to be played. Then, multiple images and audio information can be spliced ​​together synchronously to obtain a complete video to be played.

[0047] It should be noted that if the video to be played is a competition video, the audio information may include the sounds of the competition venue and the commentator's voice; if the video to be played is a TV series video, the audio information may include the actors' voices, background music, and background sound effects.

[0048] The information processing method provided in this application embodiment addresses the issue that if the target receiving bandwidth of the client does not meet the transmission bandwidth required for the video to be played within a certain period of time, the client can receive the description information and target information of each image of the video to be played sent by the target server. Based on the target model, description information, and target information, multiple images are generated. Simultaneously, the audio information of the video to be played is determined from the target information, and the multiple images and audio information are spliced ​​together to obtain the video to be played. In this way, even if the client's network conditions are poor, it can reconstruct the video to be played and play it based on the description information and target information, instead of needing to compress the video bitrate or switch networks to ensure normal video playback as in related technologies. This solves the problems of poor video quality and low video playback smoothness in related technologies, thereby improving the user's viewing experience.

[0049] Based on the foregoing embodiments, embodiments of this application provide an information processing method, referring to... Figure 3 As shown, the method may include the following steps: Step 301: The target server obtains the transmission bandwidth required for the video to be played.

[0050] In the embodiments of this application, the target server may refer to a cloud server. In one possible implementation, the target server may be a Content Delivery Network (CDN) server.

[0051] In this embodiment of the application, the resolution and frame rate of the video to be played can be obtained, and the bit rate of the video to be played, i.e. the transmission bandwidth, can be calculated according to the following formula (1).

[0052] Formula (1) Where M represents the transmission bandwidth; Width indicating resolution; H represents the height of the resolution; H represents the frame rate. This represents the target coefficient, which can be pre-set.

[0053] It should be noted that there are various existing methods for calculating transmission bandwidth, which will not be elaborated here.

[0054] Step 302: The target server obtains the client's first real-time performance information and the client's second real-time network performance information.

[0055] In this embodiment, the first real-time performance information includes the client's central processing unit (CPU) utilization, memory usage, disk read / write speed, and graphics processing unit (GPU) load, etc.; the second real-time performance information includes the client's current bandwidth usage, transmission latency, packet loss rate, error rate, jitter, and signal strength, etc.

[0056] Step 303: The target server obtains the third real-time performance information of the target server and the fourth real-time performance information of the target server's network.

[0057] In this embodiment, the third real-time performance information may include the target server's CPU utilization, memory usage, disk input / output (I / O) throughput, and disk queue length. It should be noted that the third real-time performance information can characterize the target server's load and its current request processing capacity.

[0058] In this embodiment of the application, the fourth real-time performance information may include the target server's current bandwidth usage, transmission latency, packet loss rate, error rate, user volume, network congestion, etc.

[0059] Step 304: The target server obtains the fifth real-time performance information corresponding to the transmission path between the client and the target server.

[0060] In this embodiment, the fifth real-time performance information may include the performance information of the transmission path itself, the performance information of intermediate servers on the transmission path, and the performance information of the intermediate server's network. Specifically, the performance information of the transmission path itself may include link distance, number of users, path hop count, and link type; the performance information of the intermediate servers includes CPU utilization and memory usage; and the performance information of the intermediate server's network includes current bandwidth usage, transmission latency, and packet loss rate.

[0061] It should be noted that an intermediate server can refer to a server that forwards data during communication between a client and a server, and there can be multiple intermediate servers.

[0062] Step 305: The target server uses a target bandwidth prediction model to process the first, second, third, fourth, and fifth real-time performance information to obtain the target receiving bandwidth.

[0063] In this embodiment, the target receiving bandwidth can refer to the average receiving bandwidth of the client within a target duration. For example, if the target duration is 10 seconds, then the target receiving bandwidth is the average receiving bandwidth of the client within the next 10 seconds.

[0064] In the embodiments of this application, such as... Figure 4 As shown, the first, second, third, fourth, and fifth real-time performance information obtained at multiple times can be used as input parameters into the target bandwidth prediction model to obtain multiple initial bandwidths of the client within the target duration (e.g., 10s). Then, the average value of the multiple initial bandwidths can be calculated to obtain the target received bandwidth.

[0065] It should be noted that each moment corresponds to an initial bandwidth.

[0066] In this embodiment of the application, by using the target bandwidth prediction model to analyze the performance information corresponding to the client, the target server, and the transmission path, the receiving bandwidth of the client in the future can be accurately predicted. This allows the target server to promptly determine whether the client will experience network lag in the future and send the image description information to the client in advance when lag is predicted, so that the client can play the video normally when the network conditions are poor, thereby improving the user's viewing experience.

[0067] In this embodiment of the application, the target bandwidth prediction model can be trained in the following manner: A1. Obtain the client's first historical performance information, the client's second historical performance information, the target server's third historical performance information, the target server's fourth historical performance information, and the fifth historical performance information corresponding to the transmission path.

[0068] In this embodiment of the application, taking the client as an example, unlike real-time performance information, the first historical performance information may include the client's CPU usage, memory usage, and disk usage during the playback of the target video over a past period of time; the second historical performance information may include the client's bandwidth usage, transmission latency, packet loss rate, error rate, jitter, and signal strength over a past period of time.

[0069] A2. Based on the first, second, third, fourth, and fifth historical performance information, the initial bandwidth prediction model is trained to obtain the target bandwidth prediction model.

[0070] In this embodiment of the application, the initial bandwidth prediction model can be a Long Short-Term Memory (LSTM) network model. Specifically, as shown below... Figure 5 As shown, the first, second, third, fourth, and fifth historical performance information can be cleaned and normalized respectively. Then, based on the ID of the target video, the first, second, third, fourth, and fifth historical performance information can be associated and combined. Furthermore, the combined information can be divided according to a preset ratio to obtain training data and sampling data. Then, the training data can be used as the input parameters of the initial bandwidth prediction model, that is, the training data is input into the initial bandwidth prediction model for processing to achieve model training of the initial bandwidth prediction model, thereby obtaining the target bandwidth prediction model. Then, the prediction accuracy of the target bandwidth prediction model can be verified by sampling data, and the model can be optimized by adjusting the parameters of the target bandwidth prediction model based on the prediction accuracy.

[0071] It should be noted that the target video's ID is carried in the header of the transmission protocol. This allows the target server to associate and combine relevant information based on the target video's ID when it receives the first historical performance information, etc., through the transmission protocol.

[0072] Step 306: If the target receiving bandwidth does not meet the transmission bandwidth, the target server will divide the images of the video to be played into multiple sets of images.

[0073] In this embodiment of the application, if the target receiving bandwidth of the client does not meet the transmission bandwidth, it means that the network conditions of the client will be poor in the future period. This means that the client will not be able to receive the video to be played normally. At this time, the target server can divide the images of the video to be played into multiple GOPs, i.e. multiple groups of images, in units of GOP.

[0074] It should be noted that the first image in each group is an I-frame image.

[0075] Step 307: The target server uses the target large model to process each image in multiple sets of images to obtain the descriptive information of each image.

[0076] The descriptive information is used to explain the content of each image.

[0077] In the embodiments of this application, such as Figure 6 As shown, each GOP (i.e., each group of images) can be used as an input parameter, that is, each group of images is input into the target large model, and the target large model will process each image in each group of images to obtain the descriptive information of each image.

[0078] In the embodiments of this application, such as Figure 7As shown, the target large model can be trained in the following way: B1. The sample images of the sample video are divided into multiple sets of sample images.

[0079] In this embodiment, the sample images of the sample video can be divided into multiple groups of sample images, using Groups of Pictures (GOPs) as the unit. Each group of sample images includes an I-frame, a P-frame, and a B-frame.

[0080] It should be noted that the first image in each set of sample images is an I-frame image.

[0081] B2. Determine the first sample description information of the first target sample image.

[0082] The first sample description information is used to parse the content of the first target sample image; the first target sample image is an image of a key frame in the sample video.

[0083] In this embodiment, the first target sample image is an image of all keyframes (i.e., I-frames) in the sample video, and the first target sample image is also the first image in each group of sample images. Specifically, an Artificial Intelligence Generated Content (AIGC) model can be used to process the first target sample image to generate first sample description information of the first target sample image.

[0084] For example, when the AIGC model processes a sample image of a person walking, the first sample description information generated could be "a man walking on a park path".

[0085] B3. Based on the first sample description information, determine the second sample description information of other sample images in each group of sample images.

[0086] Among them, the other sample images are the sample images in each group of sample images other than the first target sample image.

[0087] In this embodiment of the application, other sample images may include P-frame images and B-frame images in each group of sample images, excluding I-frame images. Since the relevant information of I-frame images is required when decoding B-frame images and P-frame images, the sample description information of P-frame images and B-frame images in each group of sample images can be determined based on the sample description information of I-frame images (i.e., the first sample description information).

[0088] Specifically, the AIGC model can be used to process each other sample image to generate initial sample description information for each other sample image. Then, for each other sample image, the first sample description information and the initial sample description information can be combined to obtain the second sample description information for each other sample image.

[0089] For example, for the P-frame image in each group of sample images, the second sample description information = the first sample description information of the I-frame + the initial sample description information of the P-frame; for the B-frame image in each group of sample images, the second sample description information = the first sample description information of the I-frame + the initial sample description information of the B-frame.

[0090] In this embodiment of the application, by supplementing the sample description information of other sample images with the first sample description information, a more complete image semantic chain can be constructed, thereby improving the effect of subsequent training of the initial large model.

[0091] B4. Determine the second target sample image from multiple first target sample images.

[0092] In this embodiment, the second target sample image may be an image of a portion of keyframes in the video to be played. Specifically, the second target sample image may be randomly determined from a plurality of first target sample images.

[0093] In this embodiment of the application, by determining images of some key frames (i.e., second target sample images) from all key frame images of the sample video (i.e., multiple first target sample images), the sample image data used for model training can be further optimized, thereby improving the generalization ability of the model and enhancing the adaptability and stability of the model in practical applications.

[0094] B5. Based on the sample image, the second target sample image, the first sample description information, and the second sample description information, the initial large model is trained to obtain the target large model.

[0095] In this embodiment, a Contrastive Language-Image Pretraining (Clip) text encoder can be used to encode the first sample description information and the second sample description information respectively, so as to convert the first sample description information into a first vector and the second sample description information into a second vector. Then, Gaussian noise can be applied to each sample image and each second target sample image to obtain multiple low-dimensional noise images. Further, the multiple low-dimensional noise images can be used as input parameters to the Universal Network (Unet) model for multiple rounds of iterative processing to obtain multiple processed images. Then, the multiple processed images, the first sample description information and the second sample description information can be used as input parameters, that is, the multiple processed images, the first sample description information and the second sample description information can be input into the initial large model for processing to realize the model training of the initial large model, thereby obtaining the target large model.

[0096] In the embodiments of this application, after training to obtain the target large model, as follows: Figure 8 As shown, the target database can be used to obtain the following data: first sample image from other sample videos, images of all keyframes from other sample videos (i.e., the third target sample image), images of some keyframes from other sample videos (i.e., the fourth target sample image), sample description information of all keyframe images (i.e., the third sample description information), and sample description information of images other than keyframe images from other sample videos (i.e., the fourth sample description information). Then, the third and fourth sample description information can be encoded using a Clip text encoder to obtain the third and fourth vectors. Further, the first sample image and the third target sample image can be denoised, and multiple denoised images can be input into the Unet model. Then, the comparison result between the images generated by the Unet model and the original images can be used as a loss function, and the parameters of the target large model can be adjusted according to this loss function to achieve the goal of optimizing the target large model.

[0097] Step 308: The target server determines the maximum receiving bandwidth of the client within the target duration.

[0098] In this embodiment of the application, the maximum receiving bandwidth can refer to the maximum amount of data that the client can receive within the target duration. Specifically, the maximum receiving bandwidth can be determined from the multiple initial bandwidths determined in step 305, or it can be determined by other existing algorithms.

[0099] Step 309: The target server determines the first bandwidth required for the description information, the second bandwidth required for the audio information of the video to be played, the third bandwidth required for the control information of the target transmission protocol, and the fourth bandwidth required for the first image.

[0100] The first image is a keyframe from the video to be played.

[0101] In this embodiment, the first bandwidth may be the bandwidth required for transmitting description information, the second bandwidth may be the bandwidth required for transmitting audio information, the third bandwidth may be the bandwidth required for transmitting control information, and the fourth bandwidth may be the bandwidth required for transmitting the first image; the control information is used to establish a communication connection and ensure the transmission and reception of data, and may include source address, destination address, protocol header and protocol trailer, etc.

[0102] It should be noted that control information is necessary for communication between the client and the target server.

[0103] Step 310: The target server determines the first target bandwidth based on the first bandwidth, the second bandwidth, the third bandwidth, and multiple fourth bandwidths.

[0104] In the embodiments of this application, step 310 can be implemented by steps 310a to 310b.

[0105] Step 310a: The target server determines the third target bandwidth from multiple fourth bandwidths based on the size of each first image.

[0106] In this embodiment of the application, the sizes of any two first images can be compared, and the largest first image can be determined from multiple first images based on the comparison results. The fourth bandwidth of the largest first image is then determined as the third target bandwidth.

[0107] For example: if there are 3 first images, image A has a size of 10k and a fourth bandwidth, image B has a size of 103k and a fourth bandwidth of b, and image C has a size of 1M and a fourth bandwidth of c, then the largest first image can be determined to be image C, and the third target bandwidth is c.

[0108] Step 310b: The target server determines the first target bandwidth based on the first bandwidth, the second bandwidth, the third bandwidth, and the third target bandwidth.

[0109] In this embodiment of the application, the first bandwidth, the second bandwidth, the third bandwidth and the third target bandwidth can be summed to obtain the first target bandwidth.

[0110] In the embodiments of this application, steps 311-314 or steps 315-317 can be executed after step 310.

[0111] Step 311: If the maximum receiving bandwidth does not meet the first target bandwidth, the target server determines the second target bandwidth based on the first bandwidth, the second bandwidth, and the third bandwidth.

[0112] In this embodiment of the application, if the maximum receiving bandwidth of the client meets the first target bandwidth, it means that the network condition of the client will be average in the future, that is, the client cannot receive the description information, audio information and target image at the same time. At this time, the first bandwidth, the second bandwidth and the third bandwidth can be summed to obtain the second target bandwidth.

[0113] Step 312: If the maximum receiving bandwidth meets the second target bandwidth, the target server sends description information and audio information to the client.

[0114] The target information includes audio information.

[0115] In this embodiment of the application, the maximum receiving bandwidth can be compared with the second target bandwidth. If the maximum receiving bandwidth meets the second target bandwidth, it means that the client can receive description information and audio information at the same time. At this time, description information and audio information can be sent to the client.

[0116] It should be noted that the control information of the target transmission protocol must be sent to the client. In other words, the target server essentially encapsulates the description information, audio information, and control information, and then sends the encapsulated information to the client.

[0117] In this embodiment of the application, if the maximum receiving bandwidth still does not meet the second target receiving bandwidth, it indicates that the client's network condition is extremely poor, which means that the client cannot receive any information. In this case, there is no need to send any information to the client.

[0118] Step 313: The client receives the description information and target information sent by the target server.

[0119] Among them, the target information represents the content characteristics of the video to be played; the description information is obtained by the target server after processing each image of the multiple sets of images, and is used to explain the content of each image; the multiple sets of images are obtained after dividing the video to be played into images when the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played.

[0120] In this embodiment of the application, the target information may include the audio information of the video to be played.

[0121] In this embodiment, the client can refer to a video playback terminal. Specifically, the client can receive encapsulated information sent by the target server, and then the client can decapsulate the information to determine what specific information was received, whether it is descriptive information and audio information, or descriptive information, audio information, and target image.

[0122] Step 314: The client uses the target large model to process the description information and generate multiple images.

[0123] The target image is determined by the target server from the first image of the keyframes in the video to be played.

[0124] In the embodiments of this application, such as Figure 9 As shown, the obtained multiple descriptive information can be processed by combining the descriptive information of all images except the I-frame in each group of images (i.e., each GOP) with the descriptive information of the I-frame, thereby obtaining multiple processed descriptive information.

[0125] For example, for multiple images in the Nth group of images, there are multiple descriptive information Gn={d1, d2, d3, d4…dn}, where d1 is the descriptive information of the I-frame image. Then, d1 and d2 can be added together and replaced with the original d2, d1 and d3 can be added together and replaced with the original d3, d1 and d4 can be added together and replaced with the original d4, and d1 and dn can be added together and replaced with the original dn, resulting in multiple processed descriptive information G'n={d1, d1+d2, d1+d3, d1+d4…d1+dn}.

[0126] Then, multiple processed descriptive information from each set of images can be input as input parameters into the target large model. The target large model can then process the multiple processed descriptive information to reconstruct multiple images.

[0127] Step 315: If the maximum receiving bandwidth meets the first target bandwidth, the target server determines the target image from multiple first images and sends description information, audio information and the target image to the client.

[0128] The target information includes audio information and target images.

[0129] In this embodiment of the application, if the maximum receiving bandwidth of the client meets the first target bandwidth, it means that the network condition of the client will be good in the future, that is, the client can receive description information, audio information and target image at the same time. At this time, description information, audio information and target image can be sent to the client.

[0130] In this embodiment of the application, by flexibly adjusting the transmitted information according to different network conditions, the client can provide the best visual effect in a limited network environment, thereby significantly improving the user's viewing experience.

[0131] In this embodiment of the application, the "target server determines the target image from multiple first images" in step 315 can be implemented through steps 315a to 315c.

[0132] Step 315a: The target server sorts the multiple first images based on the time of each first image in the video to be played, and obtains the arrangement order of the multiple first images.

[0133] In this embodiment of the application, multiple first images can be sorted according to the time of each first image in the video to be played, that is, the first image with the earlier time is placed first, and the first image with the later time is placed last, thereby obtaining the arrangement order of multiple first images.

[0134] Step 315b: The target server determines the similarity between each second image and the next frame image of each second image according to the arrangement order.

[0135] The second image is any image other than the last frame of the first images among a plurality of first images.

[0136] In the embodiments of this application, such as Figure 10 As shown, the similarity between each second image and its next frame can be determined sequentially according to the arrangement order, resulting in multiple similarity scores. For example, if multiple first images A include {a1, a2, a3, a4, a5}, then a5 is the last frame of the first images, and a1, a2, a3, and a4 are all second images. Subsequently, the similarity between a1 and a2, a2 ​​and a3, a3 and a4, and a4 and a5 can be calculated sequentially according to the arrangement order, resulting in multiple similarity scores S including {s1, s2, s3, s4}.

[0137] Step 315c: The target server determines the target image from multiple first images based on multiple similarities.

[0138] In this embodiment of the application, after obtaining multiple similarities, the number of targets in the target image to be extracted can be calculated according to the following formula (2).

[0139] Formula (2) Where sum represents the number of targets in the target image to be extracted; Indicates the maximum receiving bandwidth; Indicates the second target bandwidth; Indicates the third target bandwidth; Indicates a unit of time.

[0140] Next, the average similarity (avg) of multiple similarities can be calculated. Then, each of the multiple similarities is sequentially compared with the average similarity. That is, if the multiple similarities S include {s1, s2, s3, s4…sn}, the parameter a=0 is preset. Then, s1, s2, s3, s4…sn are compared with avg in sequence. If s1 is greater than or equal to avg, a=0+1=1. Then, the comparison of s2 and sn continues until a similarity less than avg is determined. At this point, the value of a is recorded. For example, if s4 is determined to be less than avg, then a is determined to be 3.

[0141] Then, the parameter a is reset to 0, and the similarity {s5, s6...sn} is traversed again. That is, it is determined whether s6 is greater than or equal to avg. If s6 is greater than or equal to avg, a=0+1=1, until a similarity less than avg is determined again, and the value of a is recorded at this time. Then, the process is repeated until all similarities are traversed.

[0142] Furthermore, the obtained multiple 'a' values ​​can be sorted in descending order of numerical value. Then, the target number of 'a' values ​​can be determined from the multiple 'a' values. For example, if a = {3, 4, 7, 10} and the target number is 2, then after sorting 'a' in descending order, we can get a = {10, 7, 4, 3}. Then, we can extract the two values ​​10 and 7 from the multiple 'a' values. These two values ​​are the image sequence numbers. Then, according to the sorting order, we can determine the 7th and 10th images from the multiple first images as the target images.

[0143] It should be noted that, in one feasible approach, to ensure that the image reconstructed by the client is completely consistent with the original image, multiple similarity scores, descriptive information, audio information, and the target image can be sent to the client together.

[0144] Step 316: The client receives the description information and target information sent by the target server.

[0145] Among them, the target information represents the content characteristics of the video to be played; the description information is obtained by the target server after processing each image of the multiple sets of images, and is used to explain the content of each image; the multiple sets of images are obtained after dividing the video to be played into images when the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played.

[0146] In this embodiment of the application, the target information includes the audio information of the video to be played and the target image.

[0147] Step 317: The client uses the target large model to process the description information and target image to obtain multiple images.

[0148] In the embodiments of this application, such as Figure 10 As shown, multiple descriptive information can be processed according to step 313 to obtain multiple processed descriptive information, and the multiple processed descriptive information and the target image can be used as input parameters to input into the target large model. The target large model can process the multiple processed descriptive information and the target image to reconstruct multiple images.

[0149] It should be noted that, in one feasible approach, if the client receives descriptive information, audio information, the target image, and similarity scores between multiple first images, then as follows: Figure 11 As shown, multiple descriptive information can be processed according to step 314 to obtain multiple processed descriptive information. Then, multiple similarities, target images and multiple processed descriptive information can be used as input parameters to input into the target large model to reconstruct multiple images of the video to be played.

[0150] In the embodiments of this application, steps 318 to 319 can be executed after steps 314 and 317.

[0151] Step 318: The client determines the audio information of the video to be played from the target information.

[0152] In this embodiment of the application, the target information can be parsed to obtain the audio information of the video to be played from the target information.

[0153] Step 319: The client generates and plays the video based on multiple image and audio information.

[0154] In the embodiments of this application, multiple image and audio information can be processed by a video generation model to reconstruct the video to be played. Alternatively, other existing methods can be used to generate the video to be played, which will not be elaborated here.

[0155] In this embodiment, the cloud server can predict the network lag situation of the client in advance, and extract the image description information of the video in advance when the lag is predicted, and then send multiple description information to the playback end (i.e. the client) so that the client can reconstruct the video through the target large model based on the description information.

[0156] It should be noted that the descriptions of the same steps and contents as in other embodiments in this embodiment can be found in the descriptions in other embodiments, and will not be repeated here.

[0157] The information processing method provided in this application embodiment can pre-determine the description information of each image of the video to be played by using a target large model in advance and send it to the client when the target receiving bandwidth of the client in the future does not meet the transmission bandwidth required by the video to be played. This allows the client to reconstruct the video to be played based on the description information and target information. In other words, even if the client experiences network lag in the future, it can still reconstruct the video to be played and play it based on the received description information and target information, instead of needing to compress the video bitrate or switch networks to ensure normal video playback as in related technologies. This solves the problems of poor video quality and low smoothness of video playback in related technologies, thereby improving the user's viewing experience.

[0158] Based on the foregoing embodiments, embodiments of this application provide a first information processing device 4, which can be applied to... Figure 1 and 3 In the information processing method provided in the corresponding embodiment, refer to Figure 12 As shown, the first information processing device 4 may include: an acquisition unit 41, a first processing unit 42, a second processing unit 43, and a sending unit 44, wherein: The acquisition unit 41 is used to acquire the transmission bandwidth required for the video to be played and to determine the target receiving bandwidth of the client within the target duration from the current moment. The first processing unit 42 is used to divide the images of the video to be played into multiple groups of images if the target receiving bandwidth does not meet the transmission bandwidth. The second processing unit 43 is used to process each image of multiple sets of images using the target large model to obtain descriptive information for each image; wherein, the descriptive information is used to explain the content of each image. The sending unit 44 is used to send description information and target information to the client so that the client can generate a video to be played based on the description information and target information; wherein, the target information represents the content features of the video to be played.

[0159] In other embodiments of this application, the acquisition unit 41 is further configured to perform the following steps: Obtain the client's first real-time performance information and the client's network's second real-time performance information; Obtain the third real-time performance information of the target server and the fourth real-time performance information of the target server's network; Obtain the fifth real-time performance information corresponding to the transmission path between the client and the target server; The target receiving bandwidth is obtained by processing the first, second, third, fourth, and fifth real-time performance information using a target bandwidth prediction model.

[0160] In other embodiments of this application, the sending unit 44 is further configured to perform the following steps: Determine the maximum receiving bandwidth for the client within the target duration; The first bandwidth required for the description information, the second bandwidth required for the audio information of the video to be played, the third bandwidth required for the control information of the target transmission protocol, and the fourth bandwidth required for the first image are determined; wherein, the first image is an image of a key frame in the video to be played. The first target bandwidth is determined based on the first bandwidth, the second bandwidth, the third bandwidth, and multiple fourth bandwidths; If the maximum receiving bandwidth does not meet the first target bandwidth, the second target bandwidth is determined based on the first bandwidth, the second bandwidth, and the third bandwidth; If the maximum receiving bandwidth meets the second target bandwidth, send description information and audio information to the client; wherein, the target information includes audio information; If the maximum receiving bandwidth meets the first target bandwidth, the target image is determined from multiple first images, and description information, audio information, and the target image are sent to the client; wherein, the target information also includes the target image.

[0161] In other embodiments of this application, the sending unit 44 is further configured to perform the following steps: Based on the size of each first image, a third target bandwidth is determined from multiple fourth bandwidths; The first target bandwidth is determined based on the first bandwidth, the second bandwidth, the third bandwidth, and the third target bandwidth.

[0162] In other embodiments of this application, the sending unit 44 is further configured to perform the following steps: The multiple first images are sorted based on the time of each first image in the video to be played, resulting in the arrangement order of the multiple first images; According to the order of arrangement, determine the similarity between each second image and the next frame image of each second image; wherein, the second image is the image other than the last frame of the first images among a plurality of first images; The target image is determined from multiple first images based on multiple similarity scores.

[0163] In other embodiments of this application, the second processing unit 43 is further configured to perform the following steps: The sample images of the sample video are divided into multiple sets of sample images; First sample description information of the first target sample image is determined; wherein, the first sample description information is used to explain the content of the first target sample image; the first target sample image is an image of a keyframe in the sample video; Based on the first sample description information, the second sample description information of other sample images in each group of sample images is determined; wherein, other sample images are sample images in each group of sample images other than the first target sample image; Determine the second target sample image from multiple first target sample images; The target large model is obtained by training the initial large model based on the sample image, the second target sample image, the first sample description information, and the second sample description information.

[0164] It should be noted that the specific implementation process of the steps performed by each unit in the embodiments of this application can be referred to Figure 1 and Figure 3 The implementation process of the information processing method provided in the corresponding embodiments will not be described in detail here.

[0165] The first information processing device provided in the embodiments of this application can pre-determine the description information of each image of the video to be played using a target large model and send it to the client when the target receiving bandwidth of the client in the future does not meet the transmission bandwidth required by the video to be played. This allows the client to reconstruct the video to be played based on the description information and target information. In other words, even if the client experiences network lag in the future, it can still reconstruct the video to be played and play it based on the received description information and target information, instead of needing to compress the video bitrate or switch networks to ensure normal video playback as in related technologies. This solves the problems of poor video quality and low video playback smoothness in related technologies, thereby improving the user's viewing experience.

[0166] Based on the foregoing embodiments, embodiments of this application provide a second information processing device 5, which can be applied to... Figure 2 and 3 In the information processing method provided in the corresponding embodiment, refer to Figure 13 As shown, the second information processing device 5 may include: a receiving unit 51, a third processing unit 52, and a fourth processing unit 53, wherein: The receiving unit 51 is used to receive description information and target information sent by the target server; wherein, the target information represents the content features of the video to be played; the description information is obtained by the target server after processing each image of multiple sets of images, and is used to explain the content of each image; the multiple sets of images are obtained by dividing the images of the video to be played when the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played. The third processing unit 52 is used to generate multiple images based on the target large model, description information and target information; The fourth processing unit 53 is used to generate a video to be played based on multiple images and target information.

[0167] In other embodiments of this application, the third processing unit 52 is further configured to perform the following steps: If the target information does not include the target image, the target large model is used to process the description information to generate multiple images; among them, the target image is determined by the target server from the first image of the keyframe in the video to be played. If the target information includes the target image, a large target model is used to process the description information and the target image to obtain multiple images.

[0168] In other embodiments of this application, the fourth processing unit 53 is further configured to perform the following steps: Determine the audio information of the video to be played from the target information; A video to be played is generated based on multiple image and audio information.

[0169] It should be noted that the specific implementation process of the steps performed by each unit in the embodiments of this application can be referred to Figure 2 and Figure 3 The implementation process of the information processing method provided in the corresponding embodiments will not be described in detail here.

[0170] The second information processing device provided in the embodiments of this application can receive description information and target information of each image of the video to be played from the target server if the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played within a certain period of time. Based on the target big model, description information and target information, it generates multiple images. At the same time, it determines the audio information of the video to be played from the target information and splices the multiple images and audio information to obtain the video to be played. In this way, even if the client's network condition is poor, it can reconstruct the video to be played and play it based on the description information and target information, instead of needing to compress the video bitrate or switch networks to ensure that the video can be played normally as in related technologies. This solves the problem of poor video quality and low smoothness of video playback in related technologies, thereby improving the user's viewing experience.

[0171] Based on the foregoing embodiments, embodiments of this application provide an information processing device, with reference to... Figure 14 As shown, the information processing device may include a target server 6, a processor may include a first processor 61, a memory may include a first memory 62, and a communication bus may include a first communication bus 63. The target server 6 can be applied to... Figure 1 and 3 In the information processing method provided in the corresponding embodiment, wherein: The first communication bus 63 is used to realize the communication connection between the first processor 61 and the first memory 62; The first processor 61 is used to execute the information processing program in the first memory 62 to perform the following steps: Obtain the transmission bandwidth required for the video to be played, and determine the target receiving bandwidth of the client within the target duration starting from the current moment; If the target receiving bandwidth does not meet the transmission bandwidth, the images of the video to be played are divided into multiple sets of images. A target large model is used to process each image in multiple sets of images to obtain descriptive information for each image; the descriptive information is used to explain the content of each image. The system sends description information and target information to the client, enabling the client to generate a video to be played based on the description information and target information; whereby the target information represents the content characteristics of the video to be played.

[0172] In other embodiments of this application, the first processor 61 is used to execute an information processing program in the first memory 62 to perform the following steps: Obtain the client's first real-time performance information and the client's network's second real-time performance information; Obtain the third real-time performance information of the target server and the fourth real-time performance information of the target server's network; Obtain the fifth real-time performance information corresponding to the transmission path between the client and the target server; The target receiving bandwidth is obtained by processing the first, second, third, fourth, and fifth real-time performance information using a target bandwidth prediction model.

[0173] In other embodiments of this application, the first processor 61 is used to execute an information processing program in the first memory 62 to perform the following steps: Determine the maximum receiving bandwidth for the client within the target duration; The first bandwidth required for the description information, the second bandwidth required for the audio information of the video to be played, the third bandwidth required for the control information of the target transmission protocol, and the fourth bandwidth required for the first image are determined; wherein, the first image is an image of a key frame in the video to be played. The first target bandwidth is determined based on the first bandwidth, the second bandwidth, the third bandwidth, and multiple fourth bandwidths; If the maximum receiving bandwidth does not meet the first target bandwidth, the second target bandwidth is determined based on the first bandwidth, the second bandwidth, and the third bandwidth; If the maximum receiving bandwidth meets the second target bandwidth, send description information and audio information to the client; wherein, the target information includes audio information; If the maximum receiving bandwidth meets the first target bandwidth, the target image is determined from multiple first images, and description information, audio information, and the target image are sent to the client; wherein, the target information also includes the target image.

[0174] In other embodiments of this application, the first processor 61 is used to execute an information processing program in the first memory 62 to perform the following steps: Based on the size of each first image, a third target bandwidth is determined from multiple fourth bandwidths; The first target bandwidth is determined based on the first bandwidth, the second bandwidth, the third bandwidth, and the third target bandwidth.

[0175] In other embodiments of this application, the first processor 61 is used to execute an information processing program in the first memory 62 to perform the following steps: The multiple first images are sorted based on the time of each first image in the video to be played, resulting in the arrangement order of the multiple first images; According to the order of arrangement, determine the similarity between each second image and the next frame image of each second image; wherein, the second image is the image other than the last frame of the first images among a plurality of first images; The target image is determined from multiple first images based on multiple similarity scores.

[0176] In other embodiments of this application, the first processor 61 is used to execute an information processing program in the first memory 62 to perform the following steps: The sample images of the sample video are divided into multiple sets of sample images; First sample description information of the first target sample image is determined; wherein, the first sample description information is used to explain the content of the first target sample image; the first target sample image is an image of a keyframe in the sample video; Based on the first sample description information, the second sample description information of other sample images in each group of sample images is determined; wherein, other sample images are sample images in each group of sample images other than the first target sample image; Determine the second target sample image from multiple first target sample images; The target large model is obtained by training the initial large model based on the sample image, the second target sample image, the first sample description information, and the second sample description information.

[0177] It should be noted that a detailed explanation of the steps performed by the first processor 61 can be found in [reference needed]. Figure 1 and 3 The information processing methods provided in the corresponding embodiments will not be described in detail here.

[0178] The target server provided in the embodiments of this application can pre-determine the description information of each image of the video to be played using a target large model and send it to the client when the target receiving bandwidth of the client in the future does not meet the transmission bandwidth required by the video to be played. This allows the client to reconstruct the video to be played based on the description information and target information. In other words, even if the client experiences network lag in the future, it can still reconstruct the video to be played and play it based on the received description information and target information, instead of needing to compress the video bitrate or switch networks to ensure normal video playback as in related technologies. This solves the problems of poor video quality and low video playback smoothness in related technologies, thereby improving the user's viewing experience.

[0179] Based on the foregoing embodiments, embodiments of this application provide another information processing device, referred to Figure 15 As shown, the information processing device may further include a client 7, a processor may further include a second processor 71, a memory may further include a second memory 72, and a communication bus may further include a second communication bus 73. The client 7 can be applied to… Figure 2 and Figure 3 In the information processing method provided in the corresponding embodiment, wherein: The second communication bus 73 is used to realize the communication connection between the second processor 71 and the second memory 72; The second processor 71 is used to execute the information processing program in the second memory 72 to perform the following steps: The system receives description information and target information sent by the target server. The target information represents the content characteristics of the video to be played. The description information is obtained by the target server after processing each image in multiple sets of images and is used to explain the content of each image. The multiple sets of images are obtained by dividing the video to be played into images when the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played. Multiple images are generated based on the target model, descriptive information, and target information. Based on multiple images and target information, a video to be played is generated.

[0180] In other embodiments of this application, the second processor 71 is used to execute an information processing program in the second memory 72 to perform the following steps: If the target information does not include the target image, the target large model is used to process the description information to generate multiple images; among them, the target image is determined by the target server from the first image of the keyframe in the video to be played. If the target information includes the target image, a large target model is used to process the description information and the target image to obtain multiple images.

[0181] In other embodiments of this application, the second processor 71 is used to execute an information processing program in the second memory 72 to perform the following steps: Determine the audio information of the video to be played from the target information; A video to be played is generated based on multiple image and audio information.

[0182] It should be noted that a detailed description of the steps performed by the second processor 71 can be found in [reference needed]. Figure 2 and 3 The information processing methods provided in the corresponding embodiments will not be described in detail here.

[0183] The client provided in the embodiments of this application can receive description information and target information of each image of the video to be played from the target server if the target receiving bandwidth of the client does not meet the transmission bandwidth required for the video to be played within a certain period of time. Based on the target large model, description information and target information, it generates multiple images. At the same time, it determines the audio information of the video to be played from the target information and splices the multiple images and audio information to obtain the video to be played. In this way, even if the client's network condition is poor, it can reconstruct the video to be played and play it based on the description information and target information, instead of needing to compress the video bitrate or switch networks to ensure that the video can play normally as in related technologies. This solves the problem of poor video quality and low video playback smoothness in related technologies, thereby improving the user's viewing experience.

[0184] Based on the foregoing embodiments, embodiments of this application provide a storage medium storing a computer program thereon, which is implemented when executed by a processor. Figures 1-3 The corresponding embodiments provide the steps of the information processing method.

[0185] Based on the foregoing embodiments, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements... Figures 1-3 The corresponding embodiments provide the steps of the information processing method.

[0186] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0187] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0188] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0189] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0190] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An information processing method, characterized in that, The method includes: Obtain the transmission bandwidth required for the video to be played, and determine the target receiving bandwidth of the client within the target duration starting from the current moment; If the target receiving bandwidth does not meet the transmission bandwidth, the images of the video to be played are divided to obtain multiple sets of images; A target large model is used to process each image in the multiple sets of images to obtain descriptive information for each image; wherein, the descriptive information is used to explain the content of each image; The description information and target information are sent to the client so that the client can generate the video to be played based on the description information and target information; wherein, the target information characterizes the content features of the video to be played.

2. The method according to claim 1, characterized in that, Determining the target receiving bandwidth for the client within the target duration starting from the current moment includes: Obtain the client's first real-time performance information and the client's network's second real-time performance information; Obtain the third real-time performance information of the target server and the fourth real-time performance information of the target server's network; Obtain the fifth real-time performance information corresponding to the transmission path between the client and the target server; The first real-time performance information, the second real-time performance information, the third real-time performance information, the fourth real-time performance information, and the fifth real-time performance information are processed using a target bandwidth prediction model to obtain the target receiving bandwidth.

3. The method according to claim 1, characterized in that, Sending the description information and target information to the client includes: Determine the maximum receiving bandwidth of the client within the target duration; The first bandwidth required for the description information, the second bandwidth required for the audio information of the video to be played, the third bandwidth required for the control information of the target transmission protocol, and the fourth bandwidth required for the first image are determined; wherein, the first image is an image of a keyframe in the video to be played; Based on the first bandwidth, the second bandwidth, the third bandwidth, and a plurality of fourth bandwidths, a first target bandwidth is determined; If the maximum receiving bandwidth does not meet the first target bandwidth, a second target bandwidth is determined based on the first bandwidth, the second bandwidth, and the third bandwidth; If the maximum receiving bandwidth meets the second target bandwidth, the description information and the audio information are sent to the client; wherein, the target information includes the audio information; If the maximum receiving bandwidth meets the first target bandwidth, a target image is determined from multiple first images, and the description information, the audio information, and the target image are sent to the client; wherein, the target information further includes the target image.

4. The method according to claim 3, characterized in that, The step of determining the first target bandwidth based on the first bandwidth, the second bandwidth, the third bandwidth, and a plurality of fourth bandwidths includes: Based on the size of each first image, a third target bandwidth is determined from the plurality of fourth bandwidths; The first target bandwidth is determined based on the first bandwidth, the second bandwidth, the third bandwidth, and the third target bandwidth. Accordingly, determining the target image from a plurality of first images includes: The plurality of first images are sorted based on the time of each first image in the video to be played, to obtain the arrangement order of the plurality of first images; According to the arrangement order, the similarity between each second image and the next frame image of each second image is determined; wherein, the second image is the image among the plurality of first images excluding the last frame first image; The target image is determined from the plurality of first images based on multiple similarity scores.

5. The method according to claim 1, characterized in that, The method further includes: The sample images of the sample video are divided into multiple sets of sample images; First sample description information of a first target sample image is determined; wherein, the first sample description information is used to explain the content of the first target sample image; the first target sample image is an image of a keyframe in the sample video; Based on the first sample description information, second sample description information of other sample images in each group of sample images is determined; wherein, the other sample images are sample images in each group of sample images other than the first target sample image; Determine the second target sample image from multiple first target sample images; The target large model is obtained by training the initial large model based on the sample image, the second target sample image, the first sample description information, and the second sample description information.

6. An information processing method, characterized in that, The method further includes: The system receives description information and target information sent by the target server; wherein, the target information characterizes the content features of the video to be played; the description information is obtained by the target server after processing each image of multiple sets of images, and is used to explain the content of each image; the multiple sets of images are obtained by dividing the video to be played into images when the target receiving bandwidth of the client does not meet the transmission bandwidth required by the video to be played; Based on the target large model, the description information, and the target information, multiple images are generated; The video to be played is generated based on the multiple images and the target information.

7. The method according to claim 6, characterized in that, Based on the target large model, the description information, and the target information, multiple images are generated, including: If the target information does not include the target image, the description information is processed using the target large model to generate the multiple images; wherein, the target image is determined by the target server from the first image of the keyframe in the video to be played; If the target information includes the target image, the target large model is used to process the description information and the target image to obtain the plurality of images; Accordingly, generating the video to be played based on the plurality of images and the target information includes: The audio information of the video to be played is determined from the target information; The video to be played is generated based on the multiple images and the audio information.

8. An information processing device, characterized in that, The information processing device includes: a processor, a memory, and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the information processing program in the memory to implement the steps of the information processing method according to any one of claims 1 to 5 or 6 to 7.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the information processing method according to any one of claims 1 to 5 or 6 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the information processing method according to any one of claims 1 to 5 or 6 to 7.