Comprehensive signal data transmission method based on large language model

Through a comprehensive signal data transmission method based on a large language model, multimodal information in the video streaming service is integrated and transmission parameters are dynamically adjusted, which solves the problems of incomplete information processing and low transmission efficiency in the video streaming service, and achieves a high-definition and smooth video experience.

CN119232979BActive Publication Date: 2025-05-06SHANGHAI SOUTH DATA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411758416.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing video streaming services have challenges in incomplete information processing, low transmission efficiency and quality, and lack of dynamic adjustment mechanisms, which leads to inability to meet users' needs for high-definition and smooth video experience.

Method used

The comprehensive signal data transmission method based on the large language model is adopted to obtain the fusion feature matrix by integrating multimodal information in the video streaming service, build a high-level language model to generate the transmission signal matrix, and optimize the transmission strategy by dynamically adjusting the signal transmission parameters and adaptive feedback loop mechanism.

Benefits of technology

It improves the comprehensiveness and accuracy of information processing, enhances the efficiency and quality of data transmission, ensures the fast and accurate transmission of video content, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119232979B_ABST
    Figure CN119232979B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information technology and media communication, and in particular to a comprehensive signal data transmission method based on a large language model, including obtaining a fusion feature matrix by fusing multimodal information in a video streaming service, thereby improving the comprehensiveness and accuracy of information processing; constructing a high-level language model based on the obtained fusion feature matrix, and generating a transmission signal matrix, which is beneficial to subsequent transmission or processing tasks; performing calculation and analysis by combining the obtained fusion feature matrix and the generated transmission signal matrix, and dynamically adjusting signal transmission parameters, optimizing the performance of signal transmission, and improving transmission efficiency and quality; and creating an adaptive feedback loop mechanism, which can cope with different network conditions and video content changes, and exhibits strong robustness. The present invention is used to solve the technical problems of low signal transmission efficiency and insufficient accuracy of video streaming services in existing solutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of information technology and media communication, and in particular to a comprehensive signal data transmission method based on a large language model. Background Art

[0002] Information technology refers to the technology used to manage and process information, including computer hardware, software, network communications, etc. In the field of media communication, information technology is widely used to improve the efficiency and quality of information acquisition, storage, processing and dissemination.

[0003] Video streaming services in the current information technology and media communication fields still face some challenges, including the following: incomplete information processing. When processing video content, they may only focus on information of a single modality (such as visual information) and ignore information of other modalities (such as auditory information and text information), resulting in insufficient comprehensiveness of information processing; low transmission efficiency and quality. Current data transmission methods may lack effective extraction and encoding of video content features, resulting in low transmission efficiency and quality, and unable to meet users' demand for high-definition and smooth video experience; lack of dynamic adjustment mechanism. During the data transmission process, previous methods were unable to optimize the transmission strategy according to factors such as the network status of the video content, thereby affecting the stability and accuracy of the transmission; in response to these challenges, corresponding solutions and technical means need to be adopted to improve the signal data transmission efficiency in video streaming services. Summary of the invention

[0004] The purpose of the present invention is to solve the problems in the background technology and to propose a comprehensive signal data transmission method based on a large language model.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A comprehensive signal data transmission method based on a large language model, comprising:

[0007] Step 1: Obtain a fusion feature matrix by fusing multimodal information in the video streaming service;

[0008] Step 2: construct a high-level language model based on the acquired fusion feature matrix and generate a transmission signal matrix;

[0009] Step 3: Perform calculation and analysis based on the acquired fusion feature matrix and the generated transmission signal matrix, and dynamically adjust the signal transmission parameters;

[0010] Step 4: Create an adaptive feedback loop mechanism to adjust and optimize the error signal matrix during data signal transmission, so as to determine the optimal transmission strategy.

[0011] It should be noted that the comprehensive signal data transmission method based on a large language model proposed in the present invention can be widely applied to video streaming media services in the field of information technology and media communication. The method can be used to optimize data transmission. Specifically, it can be used to comprehensively analyze and accurately evaluate the multimodal data of video content through technologies such as tensor decomposition, high-level language model construction, similarity measurement, and adaptive feedback loop mechanism. By dynamically adjusting signal transmission parameters and minimizing error signal matrices and other links, the accuracy and stability of data transmission are ensured, and the efficient operation of video streaming media services is promoted. At the same time, the method meets the current video streaming media services' urgent demand for improving data transmission quality and efficiency, and realizes fast and accurate transmission of video content.

[0012] Furthermore, by fusing multimodal information in the video streaming service, the process of obtaining the fusion feature matrix includes:

[0013] Acquire multimodal information in a video streaming service; wherein the multimodal information specifically includes visual information, auditory information, and text information;

[0014] Labeling visual information as a matrix ;in, Indicates the height of the video image. Indicates the width of the video image. Indicates the number of time frames; It does not directly represent a specific meaning or variable, but is used to indicate that the element type of a matrix or tensor belongs to the real number set;

[0015] Labeling auditory information as a matrix ;in, Indicates the number of video and audio signals. Indicates the sampling rate of the video and audio signals. Indicates the number of time frames;

[0016] Marking text information as a matrix ;in, The dimension representing the characteristics of video text information, Indicates the amount of video text data, i.e. the number of subtitle frames or comments;

[0017] Use tensor decomposition to fuse multimodal information, expand each modal information into a vector, and concatenate them: convert the visual information matrix , Auditory Information Matrix And the text information matrix Expand into visual vectors , auditory vector and the text vector ;in, Represents the operation of expanding a matrix into a vector;

[0018] The expanded visual vector , auditory vector and the text vector Splice to form a fusion vector , and perform SVD decomposition on it: , where represents the fusion feature matrix, is an orthogonal matrix, represents the transpose of a matrix, represents a diagonal matrix containing The singular values ​​of The key features in the fusion feature matrix are effectively integrated with the key information from different modalities.

[0019] It is understandable that in video streaming services, video content contains information in multiple modes, such as vision (image frames), hearing (audio), and text (such as subtitles, comments); by comprehensively processing multimodal information, it is possible to process visual, auditory and text information in the video at the same time, thereby improving the comprehensiveness and accuracy of information processing; these information are marked in matrix form, wherein visual information is represented by a three-dimensional matrix consisting of image height, width and time frame number; auditory information is represented by a three-dimensional matrix consisting of the number of audio signals, sampling rate and time frame number; text information is represented by a two-dimensional matrix consisting of the dimension of text information features and the number of video text data; using tensor decomposition technology, different modal information is effectively fused together to form a unified feature matrix for subsequent processing and analysis; SVD decomposition can extract key features (singular values) in the fused feature matrix, which are of great significance for understanding and analyzing video content; through key feature extraction, the data transmission process can be optimized, redundant information can be reduced, and transmission efficiency can be improved; in summary, this step can be applied to various video streaming services, is not restricted by specific platforms or formats, and has wide applicability.

[0020] Furthermore, the process of constructing a high-level language model based on the acquired fusion feature matrix and generating a transmission signal matrix includes:

[0021] Get the fusion feature matrix ;

[0022] Define the weight matrix for the high-level language model ; Among them, the weight matrix of the high-level language model It is used to establish a linear relationship between the fusion feature matrix and the transmission signal matrix, and its dimension is , represents the dimension of the weight matrix, Respectively represent the number of visual features, the number of auditory features, and the number of text features;

[0023] Combine the fusion feature matrix and weight matrix to build an advanced language model:

[0024] , where The transmission signal matrix representing the output is the final transmission signal matrix after multi-layer linear transformation and nonlinear activation function processing. In video streaming services, the transmission signal matrix represents the encoding or feature representation of the video content for subsequent transmission or processing; Represents a nonlinear activation function. When building an advanced language model framework, nonlinear characteristics are introduced through nonlinear activation functions to enhance the expressiveness of the model. represents the hyperbolic tangent function, which is used to map the input value to between 0 and 1, that is, to normalize the final transmission signal matrix of the high-level language model to a preset range; The adjustable coefficients representing the transmission signal matrix are used to control the contribution of different modal features in the final transmission signal matrix. The importance of different modes is balanced by adjusting these coefficients, thereby optimizing the performance of the model. represents the linear rectification function, which is used to set all negative values ​​to 0 while keeping positive values ​​unchanged, so that high-level language models can learn complex patterns; Represents an additional feature matrix, which contains other features related to video streaming services, such as user behavior data, network status information, etc. These features serve as supplementary information to help the model better understand the video content and generate a more accurate transmission signal matrix; Represents the bias vector, which is the same as the weight matrix A vector with matching dimensions, used to introduce an additional offset between the fused feature matrix and the transmitted signal matrix, that is, to add a constant offset to the final output;

[0025] It is understandable that by building a high-level language model, it is possible to learn and extract deep and abstract features of the video content, which is beneficial to subsequent transmission or processing tasks.

[0026] Furthermore, the process of performing calculation and analysis based on the acquired fusion feature matrix and the generated transmission signal matrix and dynamically adjusting the signal transmission parameters includes:

[0027] Get the fusion feature matrix set And the corresponding high-level language model transmission signal matrix set ;in, Represent the fusion feature matrix and index transmission signal matrix respectively, Indicates the number of matrices;

[0028] Compute the similarity matrix between the fused feature matrix and the transmitted signal matrix using advanced similarity metrics :

[0029] , where Represents the similarity value between the fusion feature matrix and the transmission signal matrix; Indicates A fusion feature matrix; Indicates A transmission signal matrix; represents the characteristic distance between the fusion feature matrix and the transmission signal matrix; Represents the attenuation factor, which is used to adjust the effect of feature distance on similarity calculation. Values ​​make the similarity calculation more sensitive to changes in feature distances; The index representing the feature distance between the fused feature matrix and the transmitted signal matrix; Respectively represent the angle or direction features of the fused feature matrix and the transmitted signal matrix (e.g., motion direction, color distribution, etc. in the video frame, converting these features into angle values); Represents the angle difference impact factor, which is used to adjust the impact of angle difference on similarity calculation; Represent the timestamps or time-related features of the fusion feature matrix and the transmission signal matrix respectively (e.g., the timestamp of the video frame, the sending time of the transmission signal, etc.); Indicates the time difference impact factor, which is used to adjust the impact of time difference on similarity calculation; is a constant term, used to avoid the situation where the denominator is zero; Indicates angle difference calculation The reference value in ;

[0030] Among them, by calculating A value between 0 and 1 can be obtained, which reflects the relative size of the angle difference between the fusion feature matrix and the transmission signal matrix: If the value of is close to 1, it means that the similarity between the fusion feature matrix and the transmission signal matrix is ​​high; if If the value of is close to 0, it means that the similarity between the fusion feature matrix and the transmission signal matrix is ​​low; similarly, by calculating A value between 0 and 1 is obtained, which reflects the relative size of the time difference between the fusion feature matrix and the transmission signal matrix: when the time characteristics between the fusion feature matrix and the transmission signal matrix are very close, then When the value of is close to 1, it means that the time difference between the fusion feature matrix and the transmission signal matrix is ​​small and the similarity is high; when the time difference between the two is large, then The value of is close to 0, indicating that the similarity between the fusion feature matrix and the transmission signal matrix is ​​low;

[0031] According to the similarity matrix , dynamically adjust the signal transmission parameters, that is, the transmission rate and encoding method :

[0032] , where Indicates the index of the transmission process, represents the function for adjusting the transmission rate and encoding method based on the similarity matrix, Respectively The standard transmission parameters corresponding to the function, namely the standard transmission rate and standard encoding method;

[0033] It is understandable that the signal transmission parameters are dynamically adjusted according to the similarity matrix between the fusion feature matrix and the transmission signal matrix. This dynamic adjustment mechanism can optimize the performance of signal transmission and improve transmission efficiency and quality according to factors such as the network conditions of the video content.

[0034] Furthermore, an adaptive feedback loop mechanism is created to adjust and optimize the error signal matrix during data signal transmission, thereby determining the optimal transmission strategy, including:

[0035] Get the output transmission signal matrix ;

[0036] Set the expected transmission signal matrix to ;

[0037] Calculate the error signal matrix using the error metric method:

[0038] , where represents the error signal matrix; represents mean square error; represents the structural similarity index; Represents the weighting coefficient, used for balancing and Contribution to the overall error;

[0039] Get the signal transmission parameters and mark them as , including transmission rate and encoding method;

[0040] The objective function is used to minimize the overall error signal matrix :

[0041] , where Represents the result value of minimizing the objective function; represents the output function; represents the weight matrix;

[0042] Among them, the iterative process of the objective function is as follows:

[0043] B1, is the weight matrix and signal transmission parameters Select the initial value;

[0044] B2. Calculation About the weight matrix and signal transmission parameters The gradient is:

[0045] , where Represent the weight matrix and signal transmission parameters The gradient value of is the symbol of partial derivative;

[0046] B3. Update the weight matrix based on the calculated gradient and signal transmission parameters Values:

[0047] , where represent the updated weight matrix and signal transmission parameters respectively; Represents the learning rate and determines the weight matrix and signal transmission parameters The update step size of

[0048] B4, check whether the stop condition is met: if it is met, stop the iteration; otherwise, return to step B2 to continue the iteration; wherein the stop condition is the maximum number of iterations or the change threshold of the objective function;

[0049] It can be understood that by creating an adaptive feedback loop mechanism, the method can continuously iteratively update the values ​​of the weight matrix and signal transmission parameters to minimize the error signal matrix, thereby continuously optimizing the transmission strategy and improving the accuracy and stability of the transmission.

[0050] Compared with the prior art, the advantages of the comprehensive signal data transmission method based on the large language model provided by the present invention are:

[0051] 1. The present invention obtains a fusion feature matrix by fusing multimodal information in video streaming services, which can more comprehensively reflect the video content and avoid information loss or one-sided interpretation. At the same time, the mutual complementation and verification of multiple modal information can improve the accuracy of information processing and provide a reliable basis for subsequent transmission or processing tasks;

[0052] 2. The present invention constructs a high-level language model based on the acquired fusion feature matrix and generates a transmission signal matrix, which is conducive to parsing video content and optimizing the structure and performance of the transmission signal. By performing calculation and analysis based on the acquired fusion feature matrix and the generated transmission signal matrix, and dynamically adjusting the signal transmission parameters, redundant information is reduced, transmission efficiency and quality are improved, and the fluency and clarity of the video content are guaranteed;

[0053] 3. The present invention creates an adaptive feedback loop mechanism to adjust and optimize the error signal matrix in the data signal transmission process, thereby determining the optimal transmission strategy, thereby improving the stability and accuracy of data transmission, reducing errors and distortions in the transmission process, and improving user experience.

[0054] In summary, the present invention, through the interrelationship and mutual promotion of the above steps, together constitutes a complete process of the comprehensive signal data transmission method based on a large language model, and together provides a strong guarantee for the efficiency, accuracy and stability of data transmission, ensuring the normal implementation of the subsequent comprehensive signal data transmission method based on a large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 The present invention provides a flowchart of the integrated signal data transmission method based on a large language model. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the implementation regulations described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] Reference Figure 1 , a comprehensive signal data transmission method based on a large language model, comprising:

[0058] Step 1: Obtain a fusion feature matrix by fusing multimodal information in the video streaming service;

[0059] Step 2: construct a high-level language model based on the acquired fusion feature matrix and generate a transmission signal matrix;

[0060] Step 3: Perform calculation and analysis based on the acquired fusion feature matrix and the generated transmission signal matrix, and dynamically adjust the signal transmission parameters;

[0061] Step 4: Create an adaptive feedback loop mechanism to adjust and optimize the error signal matrix during data signal transmission, so as to determine the optimal transmission strategy.

[0062] It should be noted that the comprehensive signal data transmission method based on a large language model proposed in the present invention can be widely applied to video streaming media services in the field of information technology and media communication. The method can be used to optimize data transmission. Specifically, it can be used to comprehensively analyze and accurately evaluate the multimodal data of video content through technologies such as tensor decomposition, high-level language model construction, similarity measurement, and adaptive feedback loop mechanism. By dynamically adjusting signal transmission parameters and minimizing error signal matrices and other links, the accuracy and stability of data transmission are ensured, and the efficient operation of video streaming media services is promoted. At the same time, the method meets the current video streaming media services' urgent demand for improving data transmission quality and efficiency, and realizes fast and accurate transmission of video content.

[0063] See also Figure 1 The present invention provides a comprehensive signal data transmission method based on a large language model. In the step 1, the step of obtaining a fusion feature matrix by fusing multimodal information in a video streaming service includes:

[0064] Step 101: Acquire multimodal information in a video streaming service; wherein the multimodal information specifically includes visual information, auditory information, and text information;

[0065] Step 102: Mark the visual information as a matrix ;in, Indicates the height of the video image. Indicates the width of the video image. Indicates the number of time frames; It does not directly represent a specific meaning or variable, but is used to indicate that the element type of a matrix or tensor belongs to the real number set;

[0066] Step 103: Mark the auditory information as a matrix ;in, Indicates the number of video and audio signals. Indicates the sampling rate of the video and audio signals. Indicates the number of time frames;

[0067] Step 104: Mark the text information as a matrix ;in, The dimension representing the characteristics of video text information, Indicates the amount of video text data, i.e. the number of subtitle frames or comments;

[0068] Step 105: Use tensor decomposition to fuse multi-modal information, expand each modal information into vectors, and concatenate them:

[0069] The visual information matrix , Auditory Information Matrix And the text information matrix Expand into visual vectors , auditory vector and the text vector ;in, Represents the operation of expanding a matrix into a vector;

[0070] The expanded visual vector , auditory vector and the text vector Splice to form a fusion vector , and perform SVD decomposition on it: , where represents the fusion feature matrix, is an orthogonal matrix, represents the transpose of a matrix, represents a diagonal matrix containing The singular values ​​of The key features in the fusion feature matrix are effectively integrated with the key information from different modalities.

[0071] In steps 101-105, in the video streaming service, the video content contains information of multiple modes such as vision (image frame), hearing (audio) and text (such as subtitles, comments); by comprehensively processing multi-modal information, the visual, auditory and text information in the video can be processed simultaneously, thereby improving the comprehensiveness and accuracy of information processing; the information is marked in matrix form, wherein the visual information is represented by a three-dimensional matrix consisting of image height, width and time frame number; the auditory information is represented by a three-dimensional matrix consisting of the number of audio signals, sampling rate and time frame number; the text information is represented by a two-dimensional matrix consisting of the dimension of the text information feature and the number of video text data; using tensor decomposition technology, different modal information is effectively fused together to form a unified feature matrix for subsequent processing and analysis; SVD decomposition can extract key features (singular values) in the fused feature matrix, which are of great significance for understanding and analyzing video content; through key feature extraction, the data transmission process can be optimized, redundant information can be reduced, and transmission efficiency can be improved; in summary, this step can be applied to various video streaming services, is not restricted by a specific platform or format, and has wide applicability.

[0072] See also Figure 1 The present invention provides a comprehensive signal data transmission method based on a large language model. In the step 2, the step of constructing a high-level language model according to the acquired fusion feature matrix and generating a transmission signal matrix includes:

[0073] Step 201: Obtain fusion feature matrix ;

[0074] Step 202: Define the weight matrix of the advanced language model ;

[0075] In step 202, the weight matrix of the high-level language model It is used to establish a linear relationship between the fusion feature matrix and the transmission signal matrix, and its dimension is , represents the dimension of the weight matrix, Respectively represent the number of visual features, the number of auditory features, and the number of text features;

[0076] Step 203: Combine the fusion feature matrix and the weight matrix to build a high-level language model:

[0077] , where The transmission signal matrix representing the output is the final transmission signal matrix after multi-layer linear transformation and nonlinear activation function processing. In video streaming services, the transmission signal matrix represents the encoding or feature representation of the video content for subsequent transmission or processing; Represents a nonlinear activation function. When building an advanced language model framework, nonlinear characteristics are introduced through nonlinear activation functions to enhance the expressiveness of the model. represents the hyperbolic tangent function, which is used to map the input value to between 0 and 1, that is, to normalize the final transmission signal matrix of the high-level language model to a preset range; The adjustable coefficients representing the transmission signal matrix are used to control the contribution of different modal features in the final transmission signal matrix. The importance of different modes is balanced by adjusting these coefficients, thereby optimizing the performance of the model. represents the linear rectification function, which is used to set all negative values ​​to 0 while keeping positive values ​​unchanged, so that high-level language models can learn complex patterns; Represents an additional feature matrix, which contains other features related to video streaming services, such as user behavior data, network status information, etc. These features serve as supplementary information to help the model better understand the video content and generate a more accurate transmission signal matrix; Represents the bias vector, which is the same as the weight matrix A vector with matching dimensions, used to introduce an additional offset between the fused feature matrix and the transmitted signal matrix, that is, to add a constant offset to the final output;

[0078] In steps 201-203, by constructing a high-level language model, deep and abstract features of the video content can be learned and extracted, which is beneficial to subsequent transmission or processing tasks.

[0079] See also Figure 1The present invention provides a comprehensive signal data transmission method based on a large language model. In step three, the steps of performing calculation and analysis based on the acquired fusion feature matrix and the generated transmission signal matrix and dynamically adjusting the signal transmission parameters include:

[0080] Step 301: Obtain fusion feature matrix set And the corresponding high-level language model transmission signal matrix set ;in, Represent the fusion feature matrix and index transmission signal matrix respectively, Indicates the number of matrices;

[0081] Step 302: Calculate the similarity matrix between the fusion feature matrix and the transmission signal matrix using an advanced similarity measurement method :

[0082] , where Represents the similarity value between the fusion feature matrix and the transmission signal matrix; Indicates A fusion feature matrix; Indicates A transmission signal matrix; represents the characteristic distance between the fusion feature matrix and the transmission signal matrix; Represents the attenuation factor, which is used to adjust the effect of feature distance on similarity calculation. Values ​​make the similarity calculation more sensitive to changes in feature distances; The index representing the feature distance between the fused feature matrix and the transmitted signal matrix; Respectively represent the angle or direction features of the fused feature matrix and the transmitted signal matrix (e.g., motion direction, color distribution, etc. in the video frame, converting these features into angle values); Represents the angle difference impact factor, which is used to adjust the impact of angle difference on similarity calculation; Represent the timestamps or time-related features of the fusion feature matrix and the transmission signal matrix respectively (e.g., the timestamp of the video frame, the sending time of the transmission signal, etc.); Indicates the time difference impact factor, which is used to adjust the impact of time difference on similarity calculation; is a constant term, used to avoid the situation where the denominator is zero; Indicates angle difference calculation The reference value in ;

[0083] In step 302, by calculating A value between 0 and 1 can be obtained, which reflects the relative size of the angle difference between the fusion feature matrix and the transmission signal matrix: If the value of is close to 1, it means that the similarity between the fusion feature matrix and the transmission signal matrix is ​​high; if If the value of is close to 0, it means that the similarity between the fusion feature matrix and the transmission signal matrix is ​​low; similarly, by calculating A value between 0 and 1 is obtained, which reflects the relative size of the time difference between the fusion feature matrix and the transmission signal matrix: when the time characteristics between the fusion feature matrix and the transmission signal matrix are very close, then When the value of is close to 1, it means that the time difference between the fusion feature matrix and the transmission signal matrix is ​​small and the similarity is high; when the time difference between the two is large, then The value of is close to 0, indicating that the similarity between the fusion feature matrix and the transmission signal matrix is ​​low;

[0084] Step 303: Based on the similarity matrix , dynamically adjust the signal transmission parameters, that is, the transmission rate and encoding method :

[0085] , where Indicates the index of the transmission process, represents the function for adjusting the transmission rate and encoding method based on the similarity matrix, Respectively The standard transmission parameters corresponding to the function, namely the standard transmission rate and standard encoding method;

[0086] In steps 301-303, the signal transmission parameters are dynamically adjusted according to the similarity matrix between the fusion feature matrix and the transmission signal matrix. This dynamic adjustment mechanism can optimize the performance of signal transmission and improve the transmission efficiency and quality according to factors such as the network status of the video content;

[0087] See also Figure 1 The present invention provides a comprehensive signal data transmission method based on a large language model. In step 4, the step of creating an adaptive feedback loop mechanism for adjusting and optimizing the error signal matrix in the data signal transmission process, thereby determining the optimal transmission strategy includes:

[0088] Step 401: Obtain output transmission signal matrix ;

[0089] Step 402: Set the expected transmission signal matrix to ;

[0090] Step 403: Calculate the error signal matrix using the error measurement method:

[0091] , where represents the error signal matrix; represents mean square error; represents the structural similarity index; Represents the weighting coefficient, used for balancing and Contribution to the overall error;

[0092] Step 404: Obtain signal transmission parameters and mark them as , including transmission rate and encoding method;

[0093] Step 405: Minimize the overall error signal matrix using the objective function :

[0094] , where Represents the result value of minimizing the objective function; represents the output function; represents the weight matrix;

[0095] In step 405, the iteration process of the objective function is as follows:

[0096] B1, is the weight matrix and signal transmission parameters Select the initial value;

[0097] B2. Calculation About the weight matrix and signal transmission parameters The gradient is:

[0098] , where Represent the weight matrix and signal transmission parameters The gradient value of is the symbol of partial derivative;

[0099] B3. Update the weight matrix based on the calculated gradient and signal transmission parameters Values:

[0100] , where represent the updated weight matrix and signal transmission parameters respectively; Represents the learning rate and determines the weight matrix and signal transmission parameters The update step size of

[0101] B4, check whether the stop condition is met: if it is met, stop the iteration; otherwise, return to step B2 to continue the iteration; wherein the stop condition is the maximum number of iterations or the change threshold of the objective function;

[0102] In steps 401-405, by creating an adaptive feedback loop mechanism, the method can continuously iteratively update the values ​​of the weight matrix and the signal transmission parameters to minimize the error signal matrix, thereby continuously optimizing the transmission strategy and improving the accuracy and stability of the transmission.

[0103] In the embodiment of the present invention, by fusing information of multiple modalities such as vision, hearing and text, the characteristics of the video content can be captured more comprehensively, and the accuracy of information processing can be improved. By using techniques such as tensor decomposition, key features (such as singular values) in the fused feature matrix can be extracted, redundant information can be reduced, the data transmission process can be optimized, and the transmission efficiency can be improved. By constructing a high-level language model, the deep and abstract features of the video content can be learned and extracted for subsequent transmission or processing tasks. By introducing nonlinear activation functions, etc., the expression ability of the model is enhanced, so that the model can better adapt to complex video content. By dynamically adjusting the signal transmission parameters (such as transmission rate and encoding method) according to the similarity matrix between the fused feature matrix and the transmission signal matrix, the performance of signal transmission can be optimized according to factors such as the network status of the video content. By dynamically adjusting the transmission parameters, the quality and efficiency of the video content in the transmission process can be ensured, and the user experience can be improved. By creating an adaptive feedback loop mechanism, the values ​​of the weight matrix and the signal transmission parameters can be continuously iterated and updated to minimize the error signal matrix, thereby continuously optimizing the transmission strategy and ensuring the effectiveness and applicability of the transmission strategy. In summary, the examples of the present invention involve data processing, comprehensive analysis and intelligent optimization decisions to solve the technical problems of low signal transmission efficiency and insufficient accuracy of video streaming services in existing solutions. In actual situations, more data and contextual information may be needed to make specific decisions and optimization solutions.

[0104] In addition, the formulas involved in the above are all calculated by removing dimensions and taking their numerical values. They are a formula that is closest to the actual situation obtained by collecting a large amount of data and performing software simulation. The proportional coefficient in the formula and the various preset thresholds in the analysis process are set by technical personnel in this field according to actual conditions or obtained by simulating a large amount of data; the size of the proportional coefficient is to quantify each parameter to obtain a specific value for subsequent comparison. The size of the proportional coefficient depends on the amount of sample data and the preliminary setting of the corresponding processing coefficient for each group of sample data by technical personnel in this field; as long as it does not affect the proportional relationship between the parameter and the quantized value.

[0105] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically based on the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0106] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0107] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0108] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0109] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0111] Secondly: In the drawings of the embodiments disclosed in the present invention, only the structures related to the embodiments disclosed in the present invention are involved, and other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of the present invention can be combined with each other;

[0112] Finally: The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A comprehensive signal data transmission method based on a large language model, characterized in that: Step 1: Obtain a fusion feature matrix by fusing multimodal information in the video streaming service; Step 2: construct a high-level language model based on the acquired fusion feature matrix and generate a transmission signal matrix; The process of step 2, constructing a high-level language model based on the acquired fusion feature matrix and generating a transmission signal matrix, includes: Get the fusion feature matrix ; Define the weight matrix for the high-level language model ; Among them, the weight matrix of the high-level language model It is used to establish a linear relationship between the fusion feature matrix and the transmission signal matrix, and its dimension is , represents the dimension of the weight matrix, Respectively represent the number of visual features, the number of auditory features, and the number of text features; Combine the fusion feature matrix and weight matrix to build an advanced language model: , where Represents the output transmission signal matrix; represents a nonlinear activation function; Represents the hyperbolic tangent function, which is used to map the input value to between 0 and 1; Represents the adjustable coefficients of the transmission signal matrix; represents the linear rectification function; represents the additional feature matrix; Represents the bias vector, which is the same as the weight matrix A vector with matching dimensions, used to introduce an additional offset between the fused feature matrix and the transmitted signal matrix, that is, to add a constant offset to the final output; Step 3: Perform calculation and analysis based on the acquired fusion feature matrix and the generated transmission signal matrix, and dynamically adjust the signal transmission parameters; Step 4: Create an adaptive feedback loop mechanism to adjust and optimize the error signal matrix during data signal transmission, so as to determine the optimal transmission strategy; The process of obtaining the error signal matrix includes: Get the output transmission signal matrix ; Set the expected transmission signal matrix to ; Calculate the error signal matrix using the error metric method: , where represents the error signal matrix; represents mean square error; represents the structural similarity index; Represents the weighting coefficient, used for balancing and contribution to the overall error.

2. The method for transmitting integrated signal data based on a large language model according to claim 1, characterized in that: In the step 1, the process of obtaining a fusion feature matrix by fusing multimodal information in the video streaming service includes: Acquire multimodal information in a video streaming service; wherein the multimodal information specifically includes visual information, auditory information, and text information; Labeling visual information as a matrix ;in, Indicates the height of the video image. Indicates the width of the video image. Indicates the number of time frames; Labeling auditory information as a matrix ;in, Indicates the number of video and audio signals. Indicates the sampling rate of the video and audio signals. Indicates the number of time frames; Marking text information as a matrix ;in, The dimension representing the characteristics of video text information, Indicates the amount of video text data, i.e. the number of subtitle frames or comments; Use tensor decomposition to fuse multimodal information, expand each modal information into a vector, and concatenate them: convert the visual information matrix , Auditory Information Matrix And the text information matrix Expand into visual vectors , auditory vector and the text vector ;in, Represents the operation of expanding a matrix into a vector; The expanded visual vector , auditory vector and the text vector Splice to form a fusion vector , and perform SVD decomposition on it: , where represents the fusion feature matrix, is an orthogonal matrix, represents the transpose of a matrix, represents a diagonal matrix containing The singular values ​​of The key features of .

3. The method for transmitting integrated signal data based on a large language model according to claim 1, characterized in that: In step 3, the process of performing calculation and analysis based on the acquired fusion feature matrix and the generated transmission signal matrix and dynamically adjusting the signal transmission parameters includes: Get the fusion feature matrix set And the corresponding high-level language model transmission signal matrix set ;in, represent the fusion feature matrix and the index transmission signal matrix respectively, Indicates the number of matrices; Compute the similarity matrix between the fused feature matrix and the transmitted signal matrix using advanced similarity metrics : , where Represents the similarity value between the fusion feature matrix and the transmission signal matrix; Indicates A fusion feature matrix; Indicates A transmission signal matrix; represents the characteristic distance between the fusion feature matrix and the transmission signal matrix; Represents the attenuation factor, which is used to adjust the impact of feature distance on similarity calculation; The index representing the feature distance between the fused feature matrix and the transmitted signal matrix; Respectively represent the angle or direction characteristics of the fusion feature matrix and the transmission signal matrix; Represents the angle difference impact factor, which is used to adjust the impact of angle difference on similarity calculation; The timestamps or time-related features of the fused feature matrix and the transmitted signal matrix are represented respectively; Indicates the time difference impact factor, which is used to adjust the impact of time difference on similarity calculation; is a constant term, used to avoid the situation where the denominator is zero; Indicates angle difference calculation The reference value in ; According to the similarity matrix , dynamically adjust the signal transmission parameters, that is, the transmission rate and encoding method : , where Indicates the index of the transmission process, represents the function for adjusting the transmission rate and encoding method based on the similarity matrix, Respectively The standard transmission parameters corresponding to the function are standard transmission rate and standard encoding method.

4. The method for transmitting integrated signal data based on a large language model according to claim 1, characterized in that: In step 4, the process of creating an adaptive feedback loop mechanism for adjusting and optimizing the error signal matrix during data signal transmission, thereby determining the optimal transmission strategy includes: Get the signal transmission parameters and mark them as , including transmission rate and encoding method; The objective function is used to minimize the overall error signal matrix : , where Represents the result value of minimizing the objective function; represents the output function; represents the weight matrix; Among them, the iterative process of the objective function is as follows: B1, is the weight matrix and signal transmission parameters Select the initial value; B2. Calculation About the weight matrix and signal transmission parameters The gradient is: , where Represent the weight matrix and signal transmission parameters The gradient value of is the symbol of partial derivative; B3. Update the weight matrix based on the calculated gradient and signal transmission parameters Values: , where represent the updated weight matrix and signal transmission parameters respectively; Represents the learning rate and determines the weight matrix and signal transmission parameters The update step size of B4. Check whether the stopping condition is met: if it is met, stop the iteration; otherwise, return to step B2 to continue the iteration; wherein the stopping condition is the maximum number of iterations or the change threshold of the objective function.

Citation Information

Patent Citations

  • Video stream adaptive transmission method based on deep reinforcement learning under QUIC protocol

    CN115022684A

  • Network packet loss optimization method and device, storage medium and electronic equipment

    CN116996397A