A Browser-Based Chat Video Transcoding Method, System, Device, and Medium

The browser-based video transcoding method addresses server load and latency issues by using intelligent transcoding and adaptive playback strategies, ensuring smooth video playback across diverse browsers and networks.

CN119135825BActive Publication Date: 2025-07-15GUANGZHOU SANQI DREAM NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411220367.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-07-15
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

The existing technology has compatibility problems caused by inconsistent support for video formats in real-time chat tools, which have high server load, large latency and serious resource waste, and cannot meet the immediate requirements.

Method used

On the browser side, by obtaining video links, constructing video source data, dynamically adjusting encoding parameters for transcoding, and using pre-training neural network models and video playback strategy libraries to optimize video playback quality.

Benefits of technology

It solves the problem of video format compatibility between browsers, reduces server load and latency, improves video transcoding efficiency and user experience, and has good scalability and compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119135825B_ABST
    Figure CN119135825B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of browser technology and provides a browser-based chat video transcoding method, including: obtaining the source video link of a video file and constructing video source data based on the source video link; transcoding the video file based on the source video link to obtain transcoded video data; writing the transcoded video data into a callback queue and caching the transcoded video data to a chat page by calling the callback queue; loading a pre-trained neural network model, where the neural network model is used to analyze the transcoded video data to enable the video file to play; constructing a video playback policy library according to the performance metrics of the browser and network environment parameters, where the video playback policy library includes multiple video playback policies, and dynamically adjusting the video playback quality by dynamically selecting a video playback policy. This application can effectively solve the problem of inconsistent support for different video formats by different browsers, and improve the efficiency and user experience of video transcoding in the chat tool of the browser.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of browsers, and particularly relates to a chat video transcoding method, system, device, and medium based on browsers. Background Art

[0002] With the rapid development of Internet technology, browser-based real-time chat tools have become an indispensable part of social media interaction, customer service response, and efficient team collaboration. These real-time chat tools greatly facilitate the circulation and sharing of information by providing instant messaging functions, enabling users to easily send multimedia content including text, pictures, audio, and video on the platform.

[0003] However, while enjoying this convenience, the transmission and display of video files usually face compatibility issues with different browsers. Specifically, different browsers have significant differences in supporting video encoding formats. For example, MOV format videos, as a relatively professional video encapsulation format, are popular in some professional application scenarios due to their high-quality encoding characteristics. However, in the browser environment of ordinary users, videos often cannot be played normally or even appear black screens due to lack of native support or inconsistent decoders, seriously affecting the user experience.

[0004] To alleviate this problem, the prior art usually adopts the method of server-side video transcoding, that is, after the video file is uploaded to the server, the server uniformly converts MOV format videos into formats widely supported by browsers, such as MP4, and then distributes them to the recipients. Although this method solves the browser compatibility problem to a certain extent, its inherent drawbacks cannot be ignored: First, the server needs to undertake additional transcoding tasks, significantly increasing the server's load pressure; second, the transcoding process takes a certain amount of time, resulting in obvious delays in video transmission and unable to meet the strict requirements of real-time chat tools for immediacy; finally, a large number of transcoding operations not only consume valuable computing resources but also may exacerbate hardware wear due to frequent disk reads and writes, causing resource waste. In summary, the server-side video transcoding method of the prior art has problems such as high server load, large delay, and resource waste, and cannot meet the requirements of real-time chat tools for efficiency and immediacy. Summary of the Invention

[0005] The embodiments of this application provide a chat video transcoding method, system, device, and medium based on browsers, which can solve one of the above-mentioned prior art problems.

[0006] In a first aspect, the embodiments of this application provide a chat video transcoding method based on browsers, including:

[0007] Obtain the source video link of the video file and construct video source data based on the source video link;

[0008] Transcode the video file based on the source video link to obtain transcoded video data;

[0009] Write the transcoded video data into a callback queue, and cache the transcoded video data to the chat page by calling the callback queue;

[0010] Load a pre-trained neural network model, which is used to analyze the transcoded video data to play the video file;

[0011] Construct a video playback policy library according to the performance metrics of the browser and network environment parameters, where the video playback policy library includes multiple video playback policies, and dynamically adjust the video playback quality by dynamically selecting video playback policies.

[0012] Further, before obtaining the source video link of the video file and constructing video source data based on the source video link, it includes:

[0013] Judge whether the video format of the video file is a preset format according to the file name of the video file;

[0014] If the video format of the video file is a preset format, cache the video file to the chat page;

[0015] If the video format of the video file is not a preset format, construct video source data, write the video source data into a storage queue, and the video source data includes a source video link, a callback queue name, and a video source flag.

[0016] Further, the transcoding the video file based on the source video link to obtain transcoded video data includes:

[0017] Obtain a video data stream through the source video link, and split the video data stream into multiple first data blocks according to a preset splitting rule, where each first data block includes a video segment with a preset duration;

[0018] Perform transcoding processing on each first data block by dynamically adjusting encoding parameters to obtain a second data block, where the second data block includes transcoded video data.

[0019] Further, the performing transcoding processing on each first data block by dynamically adjusting encoding parameters to obtain a second data block includes:

[0020] For each first data block, extract all key frames of the first data block, calculate the similarity between adjacent segment frames, and obtain the inter-frame similarity;

[0021] Judge the content complexity of the first data block according to the key frames and the inter-frame similarity;

[0022] If the key frames are dense and the inter-frame similarity is low, it indicates that the content of the first data block is complex;

[0023] If the key frames are sparse and the inter-frame similarity is high, it indicates that the content of the first database is simple;

[0024] Adopt a decision tree algorithm, and use the content complexity of the first data block as a judgment condition to dynamically select appropriate encoding parameters;

[0025] By comparing the video quality and the transcoding time between the first data block and the second data block, adopt a reinforcement learning algorithm to dynamically optimize the encoding parameter selection strategy.

[0026] Further, the step of writing the transcoded video data into the callback queue and caching the transcoded video data to the chat page by calling the callback queue includes:

[0027] Store multiple second data blocks into the callback queue in the transcoding order, and establish callback information for each second data block. The callback information includes the start timestamp, the duration, and the video source data;

[0028] When calling the callback queue, update the cached video duration of the video file in real time through the callback information;

[0029] Obtain the caching progress of the video file based on the cached video duration.

[0030] Further, after caching the transcoded video data to the chat page by calling the callback queue, it includes:

[0031] Obtain the play request of the video file. According to the play request, traverse the callback queue to obtain all the second data blocks and the caching progress in the video file;

[0032] If the video file has been transcoded completely, play the transcoded video data of the second data blocks in the play order;

[0033] If there are untranscoded first data blocks in the video file, for the transcoded first data blocks, obtain the second data blocks, and play the corresponding video segments according to the transcoded video data of the second data blocks;

[0034] For the untranscoded first data blocks, obtain the video source data of the untranscoded video segments according to the callback information of the second data blocks, and input the video source data into the video transcoder for real-time transcoding.

[0035] Further, for loading the pre-trained neural network model, the neural network model is used to analyze the transcoded video data to play the video file, including:

[0036] Compiling the neural network model to convert the neural network model into a compiled model;

[0037] Creating a WebWorker thread, loading the compiled model into the WebWorker, and performing asynchronous execution;

[0038] Obtaining the input requirement data of the compiled model, and processing the video segment corresponding to the transcoded video data according to the input requirement data to obtain a standard input segment;

[0039] Inputting the standard input segment into the compiled model to play the corresponding video segment.

[0040] In a second aspect, an embodiment of the present application provides a browser-based chat video transcoding system, including:

[0041] A first processing module: used to obtain the source video link of the video file and construct video source data based on the source video link;

[0042] A second processing module: used to transcode the video file based on the source video link to obtain transcoded video data;

[0043] A third processing module: used to write the transcoded video data into a callback queue, and cache the transcoded video data to the chat page by calling the callback queue;

[0044] A fourth processing module: used to load a pre-trained neural network model, and the neural network model is used to analyze the transcoded video data to play the video file;

[0045] A fifth processing module: used to construct a video playback policy library according to the performance metrics of the browser and network environment parameters, where the video playback policy library includes multiple video playback policies, and the video playback quality is dynamically adjusted by dynamically selecting a video playback policy.

[0046] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above browser-based chat video transcoding method is implemented.

[0047] Fourthly, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned browser-based chat video transcoding method is implemented.

[0048] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0049] The browser-based chat video transcoding method of the present application can effectively solve the problem of inconsistent support for different video formats by different browsers and avoid the video black screen phenomenon by transcoding video files on the browser side. At the same time, in order to avoid the high load and high latency problems brought by server-side transcoding, the present application also sets up a neural network model and a video playback strategy library, improving the efficiency and user experience of browser chat tool video transcoding. Therefore, the present application has good scalability and compatibility and can meet the requirements of different browsers and versions. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0051] Figure 1 FIG. is a schematic flow chart of a browser-based chat video transcoding method provided by an embodiment of the present invention;

[0052] Figure 2 FIG. is a schematic structural diagram of a browser-based chat video transcoding system provided by an embodiment of the present invention;

[0053] Figure 3 FIG. is a schematic structural diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0055] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0056] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0057] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0058] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0059] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0060] Please refer to Figure 1 As shown, the present invention is a browser-based chat video transcoding method, including the following steps:

[0061] S100. Obtain the source video link of the video file and construct video source data based on the source video link;

[0062] The chat video transcoding method of this application is applied to the real-time chat tool of the browser. When a user sends a video file during a chat, a network protocol and parsing technology are used to obtain the video file and the source video link of the video file. By constructing video source data, the video file has a unique data source, ensuring data security and providing a stable and reliable input for subsequent video transcoding, analysis, processing, etc., and ensuring the continuity and accuracy of the entire video processing process.

[0063] In some of these embodiments, before the above step S100, it includes:

[0064] According to the file name of the video file, determine whether the video format of the video file is a preset format;

[0065] If the video format of the video file is a preset format, then cache the video file to the chat page;

[0066] If the video format of the video file is not a preset format, then construct video source data and write the video source data into a storage queue. The video source data includes a source video link, a callback queue name, and a video source flag.

[0067] In this application, before performing a transcoding operation on the received video file, it is necessary to determine whether the video format of the video file is a preset format, and then determine whether a transcoding operation is required. If the video file is already in the preset format, then subsequent caching operations can be directly performed without unnecessary transcoding processes, thus saving computing resources and time. For video files that need to be transcoded, they can be screened out in advance through this step, enabling the browser to concentrate resources for efficient transcoding processing of these video files without attempting to transcode all received video files, further improving the processing efficiency. At the same time, by providing timely transcoding services for video files that do not meet the preset format, the user chat process becomes smoother and more efficient, thereby enhancing the overall user experience. In addition, performing file format judgment before transcoding can also avoid transcoding failures or system errors caused by format errors or unsupported formats. This helps reduce the system burden and failure rate caused by processing incorrect video files.

[0068] Since video files are usually stored and transmitted in the form of binary data in a computer, in this application, according to the received video file, obtain the binary data stream of the video file; by parsing the binary data stream, extract the file name of the video file. Specifically, use a pattern matching algorithm to match a file name string from the binary data stream of the video file, determine the start position and end position of the file name string, and extract the complete file name.

[0069] In this embodiment, a file name string is obtained and stored in a string variable. By traversing the above string variable from right to left, the position index of the last dot is obtained. According to the above position index, the substring after the dot is intercepted and stored in a format string variable. The above format string variable is compared with a preset format to determine whether the video format of the video file is the preset format, and the judgment result is returned.

[0070] In this embodiment, the preset format is usually set based on the compatibility requirements of the browser to ensure that all video files to be processed meet the playback requirements of the browser and reduce problems such as black screens caused by format incompatibility. Therefore, the number of preset formats is usually multiple. The multiple preset formats are stored in a preset format list. By traversing the above preset format list, it is determined whether the video format of the video file meets the playback requirements of the browser, and then it is judged whether the video file needs to be transcoded.

[0071] During the chat process, the user may send more than one video file. Therefore, in this embodiment, by timely writing the unique video source data constructed by each video file into the storage queue, the system can manage the transcoding tasks of multiple video files in an orderly manner. The orderly queuing of the video source data ensures that the transcoding tasks can be carried out in the receiving order or priority, avoiding resource conflicts and efficiency degradation caused by processing multiple videos simultaneously, ensuring that each video is processed timely and orderly, reducing the waiting time of the user, improving the processing efficiency, and at the same time enhancing the overall user experience. In addition, through queue management, the browser can allocate resources more flexibly, avoiding browser crashes or performance degradation caused by sudden large numbers of video transcoding requests, and enhancing the stability and reliability of the browser.

[0072] S200. Transcode the video file based on the source video link to obtain transcoded video data;

[0073] In some of the embodiments, the above step S200 includes:

[0074] Obtain a video data stream through the source video link, and split the video data stream into multiple first data blocks according to a preset splitting rule, where each of the first data blocks includes a video segment with a preset duration;

[0075] Perform transcoding processing on each of the first data blocks by dynamically adjusting encoding parameters to obtain a second data block, where the second data block includes transcoded video data.

[0076] In this embodiment, the video data stream of the video file is obtained through the Web Streams API, the format and encoding method of the video data stream are judged, and the appropriate segmentation rule and data block size are determined. Among them, the Web Streams API is an interface for processing stream data in a web browser. It provides a mechanism to transfer the data stream from one place to another without loading the entire data set into memory. Therefore, the Web Streams API includes the data stream of the video file.

[0077] In this embodiment, the preset duration of the first data block is determined according to the total video duration and the preset number of data blocks. The size of the first data block needs to evaluate the impact of data blocks of different sizes on the storage space and transcoding efficiency. Specifically, machine learning algorithms such as decision trees or support vector machines are used to obtain the optimal size of the first data block, minimizing the number of transcoding times while meeting the progressive caching. For example, when the size of the first data block is 2MB, the storage space is 200GB and the transcoding time is 10 minutes. When the size of the first data block is 4MB, the storage space is 100GB and the transcoding time is 8 minutes. Through the decision tree algorithm, the optimal data block size can be obtained as 4MB.

[0078] In this embodiment, for each first data block, the duration of the video segment of the transcoded second data block should be the same as that of the first data block to ensure the continuity of video playback.

[0079] In some of these embodiments, the transcoding process of each of the first data blocks by dynamically adjusting the encoding parameters to obtain the second data blocks includes:

[0080] For each of the first data blocks, all key frames of the first data block are extracted, the similarity between adjacent segment frames is calculated, and the inter-frame similarity is obtained;

[0081] According to the key frames and the inter-frame similarity, the content complexity of the first data block is judged;

[0082] If the key frames are dense and the inter-frame similarity is low, it indicates that the content of the first data block is complex;

[0083] If the key frames are sparse and the inter-frame similarity is high, it indicates that the content of the first database is simple;

[0084] Using the decision tree algorithm, with the content complexity of the first data block as the judgment condition, appropriate encoding parameters are dynamically selected;

[0085] By comparing the video quality and the transcoding time between the first data block and the second data block, the encoding parameter selection strategy is dynamically optimized using the reinforcement learning algorithm.

[0086] In this embodiment, for each first data block, it is segmented to obtain multiple segment frames. Among them, the segmentation granularity can be dynamically adjusted according to the total duration and content characteristics of the first data block. For example, a first data block with a relatively short predetermined duration can be segmented into 5-10 segment frames, while a first data block with a longer duration can be segmented into 20-30 segment frames. In addition, for the segmented segment frames, the Harris corner detection algorithm is used to extract the key frames therein. Specifically, all the corner points in each segment frame are obtained, and by counting the number of corner points in each segment frame and according to the distribution characteristics of the number of corner points, the segment frames with the number of corner points exceeding the third quartile are selected as the key frames.

[0087] In this embodiment, the structural similarity method is used to calculate the similarity between two adjacent segment frames. Specifically, the calculation formula for the inter-frame similarity is as follows:

[0088]

[0089] where SSIM is the inter-frame similarity, and the value range of the inter-frame similarity is from -1 to 1. The larger the value, the more similar the two adjacent segment frames are. x and y respectively represent two adjacent segment frames, μ x and μ y respectively represent the means of x and y, and respectively represent the variances of x and y, σ xy represents the covariance between x and y, and C1 and C2 both represent constants used to prevent the denominator from being zero.

[0090] In this embodiment, the content complexity of the corresponding first data block is judged by the number of key frames and the inter-frame similarity, and then the encoding parameters during transcoding of the first data block are determined to ensure the quality of the video file after transcoding. Specifically, if the number of key frames is large, it indicates that the video segments of the first data block contain more important information or significant changes, thus increasing the content complexity. In addition, if the inter-frame similarity is generally low, it means that the difference between adjacent segment frames is large, and the content in the video segments of the first data block changes frequently, which also increases the content complexity synchronously. On the contrary, if the inter-frame similarity is generally high and the number of key frames is low, it means that the video content is relatively stable and the complexity is low.

[0091] In this embodiment, a CART decision tree is constructed, taking the number of key frames and the inter-frame similarity as input features, and training to generate a decision tree capable of outputting encoding parameters, so as to determine the encoding parameters according to the content complexity of the first data block during transcoding, where the encoding parameters include quantization parameter (QP), frame rate (FPS), resolution, bit rate, etc. In addition, during transcoding, the Q-Learning reinforcement learning algorithm is used to dynamically optimize the encoding parameter selection strategy, so that the encoding parameter selection strategy achieves a better balance between transcoding quality and transcoding speed. Specifically, when constructing an encoding parameter optimization strategy based on Q-Learning reinforcement learning, the state and action spaces need to be clearly defined. Among them, the state space is used to reflect various characteristics of the current encoding environment, including but not limited to the resolution and complexity of the current frame, as well as the average PSNR value and average time consumption of the past encoded frames. At the same time, it is also necessary to record the settings of the previous encoding parameters and their effects, such as the quantization parameter QP, frame rate setting, and selection of encoding mode. The action space is the adjustment operations that can be performed when optimizing the encoding parameters, such as changing the quantization parameter QP (affecting the compression ratio and quality), changing the encoding mode (such as switching from I-frame to P / B-frame), adjusting the encoding rate, etc. After that, when designing the reward function, for example, according to the optimization goal, when the PSNR value significantly increases by 1 dB and the encoding time consumption decreases by 10%, a significant reward (+10) should be given to encourage such positive changes. On the contrary, if the PSNR value drops by 1 dB and the time consumption increases by 10%, a severe penalty (-10) should be imposed to avoid such adverse situations. For intermediate situations between the two, the reward function can be flexibly designed, and a linear interpolation or a more complex function can be used to reflect the trade-off between PSNR improvement and time consumption reduction, so as to guide the learning algorithm to find the best balance point between quality and speed.

[0092] S300. Write the transcoded video data into the callback queue, and cache the transcoded video data to the chat page by calling the callback queue;

[0093] In some of these embodiments, the above step S300 includes:

[0094] Store multiple second data blocks into the callback queue in the transcoding order, and establish callback information for each second data block, where the callback information includes a start timestamp, a duration, and video source data;

[0095] When calling the callback queue, the cached video duration of the video file is updated in real time through the callback information;

[0096] Based on the cached video duration, obtain the caching progress of the video file.

[0097] In order to improve the transcoding efficiency of video files, in this embodiment, each first data block is transcoded independently, and then the multiple second data blocks obtained by transcoding are sequentially stored in the callback queue. By calling the callback queue, the second data blocks are cached to the chat page. At the same time, when the callback queue is called, callback information is returned, and the caching progress of the video file can be obtained.

[0098] In some of these embodiments, after caching the transcoded video data to the chat page by calling the callback queue, it includes:

[0099] Obtain a play request for the video file. According to the play request, traverse the callback queue to obtain all the second data blocks in the video file and the caching progress;

[0100] If the video file has been completely transcoded, then play the transcoded video data of the second data blocks in sequence according to the play order;

[0101] If there are first data blocks in the video file that have not been transcoded, for the first data blocks that have been transcoded, obtain the second data blocks, and play the corresponding video segments according to the transcoded video data of the second data blocks;

[0102] For the first data blocks that have not been transcoded, obtain the video source data of the video segments that have not been transcoded according to the callback information of the second data blocks, and input the video source data into the video transcoder for real-time transcoding.

[0103] In this example, according to the obtained play request, the video file requested to be played is obtained from the storage queue, and the callback queue is traversed to obtain all the first data blocks that have not been transcoded and cached and the transcoded video data obtained by transcoding for the corresponding video file. According to the transcoded video data, play is performed in the time order of the video file. Thus, when a play request for a video file is obtained, the transcoding and caching status of the video file can be obtained in a timely manner, and the transcoded video data that has been transcoded is immediately played in response, thereby reducing the user's waiting time and improving the user's chat and viewing experience.

[0104] It can be understood that in this embodiment, the first data blocks are transcoded into second data blocks in the time order of the video file. Thus, when there are first data blocks in the video file that have not been transcoded, the system obtains the video segments that have been cached in the video file according to the transcoded video data that has been transcoded and the callback information, and plays the cached video segments in time order.

[0105] In this embodiment, the video transcoder is a pre-trained model that can transcode the untranscoded first data block into a second data block according to the performance metrics of the browser and the network environment parameters. At the same time, it calls the transcoded video data obtained by transcoding and caches it in a timely manner to improve the smoothness of video playback. In a preferred embodiment, the video transcoder uses the built-in WebAssembly technology of the browser to transcode the first video block without relying on third-party tools. The transcoding process is completed on the browser side, ensuring real-time performance.

[0106] S400. Load a pre-trained neural network model, where the neural network model is used to analyze the transcoded video data and play the video file.

[0107] In some of these embodiments, the above step S400 includes:

[0108] Compile the neural network model to convert it into a compiled model.

[0109] Create a WebWorker thread, load the compiled model into the WebWorker, and perform asynchronous execution.

[0110] Obtain the input requirement data of the compiled model, and process the video segment corresponding to the transcoded video data according to the input requirement data to obtain a standard input segment.

[0111] Input the standard input segment into the compiled model and play the corresponding video segment.

[0112] In this embodiment, the pre-trained neural network model is a network model that can perform video analysis and play videos, which is obtained by training with a large amount of training data. By loading the pre-trained neural network model, the transcoded video data can be efficiently analyzed quickly, thereby reducing the waiting time and improving the user's viewing experience. In addition, the pre-trained neural network model has strong generalization ability and robustness, can work stably in different scenarios and environments, reduces the dependence on manual intervention and customized development, and thus reduces the operation costs of video file processing and chat tools.

[0113] In this embodiment, WebAssembly technology is used for compilation to convert the neural network model into a binary format that can be directly executed by the browser, enabling the video to be played in the browser environment. In addition, a Web Worker thread is created, and the compiled model after compilation is loaded into the Web Worker to achieve asynchronous execution of the compiled model, avoiding blocking the main thread and enhancing the user experience.

[0114] In this embodiment, it is necessary to process the video segments corresponding to the transcoded video data to meet the input requirements of the neural network model, so that the corresponding video segments can be played after the transcoded video data is input into the neural network model. Specifically, the input requirement data of the neural network model is obtained, where the input requirement data includes the dimensions and data types of the input tensors, etc. The size of the video file is obtained through the video source data of the video file. If the size of the video file is larger than the size requirement of the input layer of the neural network model, a cropping operation is performed on the video segments corresponding to the transcoded video data to remove the redundant pixels and obtain the standard input segments to match the model input size; if the size of the video frame is smaller than the size requirement of the model input layer, a padding operation is performed on the corresponding video segments to add pixels around to reach the size requirement of the model input. Taking the ResNet-50 neural network model as an example, ResNet-50 requires an input tensor with a size of 224x224x3 and a data type of float32. If the size of the video file obtained through the video source data is 1920x1080, the video segments of the transcoded video data need to be scaled to 224x224, where the scaling ratio is 117, and at the same time, the central region of 224x224 is cropped out; if the video frame size is smaller than the model input size, a padding operation is performed, such as adding black pixels around the video frame to make it reach the size of 224x224.

[0115] S500. Construct a video playback policy library according to the performance metrics of the browser and the network environment parameters. Among them, the video playback policy library includes multiple video playback policies, and the video playback quality is dynamically adjusted by dynamically selecting video playback policies.

[0116] In this embodiment, the performance metrics of the current browser and the network environment parameters are obtained. Among them, the performance metrics are obtained through the performance monitoring interface, and the performance metrics include the CPU occupancy rate and the memory occupancy rate. The network environment parameters are obtained in real time through the network detection tool, and the network environment parameters include the network bandwidth, latency, and packet loss rate. The network environment is divided into three levels: excellent, general, and poor according to the preset network threshold, and the performance metrics are divided into three levels: high, medium, and low according to the preset performance threshold; the network environment type and the performance metrics are combined to establish in the video playback policy library, and the video playback quality is determined. For example, by combining the above three levels of network environment and three levels of performance metrics, 9 combination methods can be obtained, and each combination method is a video playback policy. Among them, the video playback policy is correspondingly provided with parameters such as video coding format, bit rate, and resolution to measure the video playback quality.

[0117] In another embodiment, real-time network speed change data during video playback is obtained, and through a time series analysis algorithm, the network speed change trend in the next period of time is predicted. If the predicted future network speed change trend shows a significant decline, the current video playback strategy is dynamically adjusted to reduce the playback quality and avoid video stuttering. If the predicted future network speed change trend remains stable or increases, the current video playback strategy is maintained, or the video playback quality is moderately improved to enhance video clarity. By continuously monitoring network environment parameters, device performance parameters, and video playback status parameters, the video playback strategy is adjusted in real time to ensure a smooth video playback experience under different network conditions and device capabilities.

[0118] Please refer to Figure 2 As shown, the present invention also provides a browser-based chat video transcoding system, and the system includes:

[0119] The first processing module 201: is used to obtain the source video link of the video file and construct video source data based on the source video link;

[0120] The second processing module 202: is used to transcode the video file based on the source video link to obtain transcoded video data;

[0121] The third processing module 203: is used to write the transcoded video data into a callback queue and cache the transcoded video data to the chat page by calling the callback queue;

[0122] The fourth processing module 204: is used to load a pre-trained neural network model, and the neural network model is used to analyze the transcoded video data to play the video file;

[0123] The fifth processing module 205: is used to construct a video playback strategy library according to the performance metrics of the browser and network environment parameters, wherein the video playback strategy library includes multiple video playback strategies, and the video playback quality is dynamically adjusted by dynamically selecting a video playback strategy.

[0124] It can be understood that the content in the embodiment of the browser-based chat video transcoding method as Figure 1 shown is applicable to the embodiment of this browser-based chat video transcoding system. The functions specifically implemented in the embodiment of this browser-based chat video transcoding system are the same as those in the embodiment of the browser-based chat video transcoding method as Figure 1 shown, and the beneficial effects achieved are also the same as those achieved in the embodiment of the browser-based chat video transcoding method as Figure 1 shown.

[0125] It should be noted that, regarding the information interaction, execution process, etc. between the above systems, since they are based on the same concept as the method embodiments of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be repeated here.

[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments, and details will not be repeated here.

[0127] Please refer to Figure 3 As shown, an embodiment of the present invention further provides a computer device 3, including: a memory 302, a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the browser-based chat video transcoding method as described in any one of the above methods.

[0128] The computer device 3 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art can understand that Figure 3 This is only an example of the computer device 3 and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0129] The so-called processor 301 may be a Central Processing Unit (CPU), and this processor 301 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0130] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In other embodiments, the memory 302 may also be an external storage device of the computer device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 3. Further, the memory 302 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 302 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 302 may also be used to temporarily store data that has been output or is to be output.

[0131] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the browser-based chat video transcoding method as described in any one of the above methods.

[0132] In this embodiment, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the method of the above embodiment in this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0133] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A browser-based chat video transcoding method, characterized in that including: Obtain the source video link of the video file and construct video source data based on the source video link; Transcode the video file based on the source video link to obtain transcoded video data; Write the transcoded video data into a callback queue and cache the transcoded video data to the chat page by calling the callback queue; Load a pre-trained neural network model, which is used to analyze the transcoded video data to play the video file; Construct a video playback policy library according to the performance metrics of the browser and network environment parameters, where the video playback policy library includes multiple video playback policies, and the video playback quality is dynamically adjusted by dynamically selecting video playback policies; Wherein, the transcoding the video file based on the source video link to obtain transcoded video data includes: obtaining a video data stream through the source video link, and splitting the video data stream into multiple first data blocks according to a preset splitting rule, where each of the first data blocks includes a video segment with a preset duration; performing transcoding processing on each of the first data blocks by dynamically adjusting encoding parameters to obtain second data blocks, where the second data blocks include transcoded video data; Wherein, the performing transcoding processing on each of the first data blocks by dynamically adjusting encoding parameters to obtain second data blocks includes: for each of the first data blocks, extracting all key frames of the first data block, calculating the similarity between adjacent segment frames to obtain the inter-frame similarity; judging the content complexity of the first data block according to the key frames and the inter-frame similarity; if the key frames are dense and the inter-frame similarity is low, it indicates that the content of the first data block is complex; if the key frames are sparse and the inter-frame similarity is high, it indicates that the content of the first data block is simple; using a decision tree algorithm, taking the content complexity of the first data block as a judgment condition, dynamically select appropriate encoding parameters; by comparing the video quality and transcoding processing time between the first data block and the second data block, dynamically optimize the encoding parameter selection strategy using a reinforcement learning algorithm.

2. The method according to claim 1, characterized in that, Before obtaining the source video link of the video file and constructing video source data based on the source video link, it includes: Judge whether the video format of the video file is a preset format according to the file name of the video file; If the video format of the video file is a preset format, cache the video file to the chat page; If the video format of the video file is not a preset format, construct video source data and write the video source data into a storage queue, where the video source data includes a source video link, a callback queue name, and a video source marker.

3. The method according to claim 1, wherein The writing the transcoded video data into a callback queue and caching the transcoded video data to the chat page by calling the callback queue includes: Sequentially store multiple of the second data blocks into the callback queue according to the transcoding order, and establish callback information for each of the second data blocks, where the callback information includes a start timestamp, a duration, and video source data; When calling the callback queue, update the cached video duration of the video file in real time through the callback information; Obtain the caching progress of the video file based on the cached video duration.

4. The method according to claim 3, characterized in that, After caching the transcoded video data to the chat page by calling the callback queue, it includes: Obtain the play request of the video file, traverse the callback queue according to the play request, and obtain all the second data blocks and the caching progress in the video file; If the video file has been transcoded completely, play the transcoded video data of the second data blocks in sequence according to the play order; If there are untranscoded first data blocks in the video file, for the transcoded first data blocks, obtain the second data blocks, and play the corresponding video segments according to the transcoded video data of the second data blocks; For the untranscoded first data blocks, obtain the video source data of the untranscoded video segments according to the callback information of the second data blocks, and input the video source data into a video transcoder for real-time transcoding.

5. The method according to claim 1, wherein Loading the pre-trained neural network model, where the neural network model is used to analyze the transcoded video data to play the video file, includes: Compile the neural network model and convert it into a compiled model; Create a WebWorker thread, load the compiled model into the WebWorker, and perform asynchronous execution; Obtain the input requirement data of the compiled model, and process the video segments corresponding to the transcoded video data according to the input requirement data to obtain standard input segments; Input the standard input segments into the compiled model and play the corresponding video segments.

6. A browser-based chat video transcoding system, characterized in that, It includes: The first processing module: used to obtain the source video link of the video file and construct video source data based on the source video link; The second processing module: used to transcode the video file based on the source video link to obtain transcoded video data; The third processing module: used to write the transcoded video data into the callback queue and cache the transcoded video data to the chat page by calling the callback queue; The fourth processing module: used to load the pre-trained neural network model, where the neural network model is used to analyze the transcoded video data to play the video file; The fifth processing module: used to construct a video play policy library according to the performance metrics of the browser and network environment parameters, where the video play policy library includes multiple video play policies, and dynamically adjust the video play quality by dynamically selecting video play policies; Among them, transcoding the video file based on the source video link to obtain transcoded video data includes: obtaining a video data stream through the source video link, splitting the video data stream into multiple first data blocks according to a preset splitting rule, where each first data block includes a video segment with a preset duration; performing transcoding processing on each first data block by dynamically adjusting encoding parameters to obtain second data blocks, where the second data blocks include transcoded video data; Among them, the transcoding process of each of the first data blocks by dynamically adjusting encoding parameters to obtain second data blocks includes: for each of the first data blocks, extracting all key frames of the first data block, calculating the similarity between adjacent segment frames to obtain the inter-frame similarity; judging the content complexity of the first data block according to the key frames and the inter-frame similarity; if the key frames are dense and the inter-frame similarity is low, it indicates that the content of the first data block is complex; if the key frames are sparse and the inter-frame similarity is high, it indicates that the content of the first data block is simple; using a decision tree algorithm, with the content complexity of the first data block as the judgment condition, dynamically selecting appropriate encoding parameters; by comparing the video quality and the transcoding time between the first data block and the second data block, using a reinforcement learning algorithm to dynamically optimize the encoding parameter selection strategy.

7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Web-end plug-in-free monitoring video playing method based on Wasm

    CN114827751A

  • Multi-channel video playing method, device and system for property monitoring and medium

    CN116405725A