Content generation method, apparatus, device, medium, and product
By transforming feature vectors and combining modules in the first machine learning model, the problems of low efficiency and high memory consumption caused by large amounts of data in content generation are solved, achieving efficient data processing and quality assurance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies suffer from low processing efficiency and high memory consumption in content generation due to the handling of large amounts of data, resulting in bottlenecks in computing and memory resources.
The first machine learning model is used to convert the first feature vector into a second feature vector. The data is then processed through the first and second modules to reduce the amount of data and optimize the data processing flow. This includes the use of linear units, data block and boundary padding units, and the combined use of the first and second modules.
While reducing data processing volume and memory usage, the quality of output content is guaranteed, solving the bottleneck problem of computing and memory resources and improving processing efficiency.
Smart Images

Figure CN122491356A_ABST
Abstract
Description
Technical Field
[0001] This relates to the field of computer processing technology, and in particular to a content generation method, apparatus, device, medium, and product. Background Technology
[0002] Content generation tasks commonly employ machine learning models. However, using machine learning models presents challenges due to the need to process large amounts of data, leading to low processing efficiency and high memory consumption. Summary of the Invention
[0003] This invention provides a content generation method, apparatus, device, medium, and product that achieves the technical effect of ensuring output quality while reducing data processing volume.
[0004] In one scenario, this paper provides a content generation method that includes: Obtain the first input information; The first input information is analyzed and processed based on a first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators. The second module is used to instruct the processing of the second feature vector. The data volume of the second feature vector is less than the data volume of the first feature vector. The first content is displayed, which reflects the result of the first machine learning model based on the second feature vector.
[0005] In one instance, this document also provides a content generation apparatus, which includes: The information receiving module is used to acquire the first input information; The information processing module is used to analyze and process the first input information based on the first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the second feature vector to be obtained from the first feature vector according to multiple first indicators. The second module is used to instruct the processing of the second feature vector. The data volume of the second feature vector is less than the data volume of the first feature vector. The content display module is used to display the first content, which reflects the result of the first machine learning model based on the second feature vector.
[0006] In one instance, this document also provides an electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the content generation method as described herein.
[0007] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform the content generation methods described herein.
[0008] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the content generation method as described herein.
[0009] The beneficial effects of the above content generation method are as follows: First, the first input information is obtained; second, the first input information is analyzed and processed based on a first machine learning model, which includes a first module and a second module. The first module can instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators. The second module can instruct the processing of the second feature vector. Since the data volume of the second feature vector is less than that of the first feature vector, a limited number of feature vectors can be processed based on the first machine learning model, reducing the amount of data processing. This solves the bottleneck problem of computational and memory resources caused by redundancy of some features when uniformly processing feature vectors, achieving the technical effect of ensuring the quality of output content while reducing the amount of data processing and memory usage. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments described herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a schematic diagram of the architecture of an exemplary system in one scenario. Figure 2 This is a schematic diagram of the local structure of a third machine learning model in one scenario; Figure 3 This is a flowchart illustrating a content generation method for one scenario. Figure 4 This is a schematic diagram of the local structure of a second machine learning model in one scenario; Figure 5 This is a flowchart illustrating a content generation method for one scenario. Figure 6 This is a schematic diagram of the local structure of the first machine learning model in one scenario; Figure 7This is a flowchart illustrating a content generation method for one scenario. Figure 8 This is a flowchart illustrating a content generation method for one scenario. Figure 9 This is a schematic diagram of a content generation device in one scenario. Figure 10 This is a schematic diagram of the structure of an electronic device under one specific scenario. Detailed Implementation
[0012] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solutions can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the technical solutions herein. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solutions.
[0013] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.
[0015] It should be noted that the concepts of "first" and "second" mentioned are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0016] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this document, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this document in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.
[0022] It is understood that the data involved in the technical solutions in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] This article will first illustrate the application scenarios of this technical solution: one scenario is its applicability in generating videos from arbitrary images, videos from videos, and videos from text. Before introducing this scenario, we can first give a brief introduction to the machine learning model used in this scenario.
[0024] In one scenario, the architecture of the third machine learning model is determined based on the sequence model architecture of the self-attention mechanism. A fourth module associated with the attention module can be introduced into the sequence model of the self-attention mechanism. The fourth module mainly corresponds to the attention unit and the pooling unit. The purpose of introducing the fourth module is to determine the optimal block-segmentation strategy by calculating the loss corresponding to the feature vector after block segmentation based on the second strategy, thereby ensuring the processing accuracy of the model.
[0025] In another scenario, the architecture of the second machine learning model is based on the optimal block-splitting strategy output by the third machine learning model. The self-attention mechanism's sequence model architecture is adjusted by replacing the attention module in the sequence model architecture with a first, second, third, and fourth module. The purpose of introducing the first module is to generate inter-block correlation attributes to instruct the second module to select key vectors. The purpose of introducing the second module is to dynamically select key blocks that can be retained and perform sparse attention computation based on the selected key blocks to reduce data processing volume. The purpose of introducing the third module is to perform dimensionality reduction and block-splitting processing on the feature vectors. The purpose of introducing the fourth module is to calculate the loss based on the feature vectors output by the fourth module and the feature vectors output by the first module, to supervise the first module in learning the full attention block-level distribution reflected by the fourth module, thereby improving the quality of sparse selection.
[0026] In another scenario, the fourth module can be removed from the second machine learning model to obtain the first machine learning model. The first machine learning model then retains only the first and second modules. The reason for removing the fourth module is that it is primarily used during the training phase, providing full attention to supervise the first module's learning of block-level relevance prediction. In practical applications, only the trained first module is needed for block-level relevance prediction, and sparse attention computation is required based on the second module. This significantly reduces data processing volume and memory usage while maintaining the quality of the generated data.
[0027] In the process of processing data based on machine learning models, the majority of data processing is concentrated in the attention module. The data processed by the attention module mainly consists of the query vector Q, key vector K, and value vector V input to it. Therefore, determining and filtering the query vector Q, key vector K, and value vector V is crucial.
[0028] In one scenario, we will first introduce how to obtain the query vector Q, key vector K, and value vector V. The first, second, and third machine learning models all include linear units and data partitioning and boundary padding units.
[0029] Linear unit: Used to perform linear transformations on the input data to generate a query vector Q, a key vector K, and a value vector V. Optionally, the linear unit may include three parallel network layers, such as fully connected layers or convolutional layers, which are used to generate the corresponding query vector Q, key vector K, and value vector V, respectively.
[0030] Data partitioning and boundary padding unit: Used to partition the query vector Q, key vector K, and value vector V output by the linear unit into blocks. The vector partitioning information for dividing the query vector Q, key vector K, and value vector V can be determined based on the method provided in this case. In some cases, the provided content generation method can be applied to... Figure 1 The content generation system shown may include a client 101 and a server 102. The client 101 may include, but is not limited to, browsers, applications (Apps), HyperText Markup Language (HTML) applications, lightweight applications (also known as mini-programs), or cloud applications. The client 101 may be deployed on an electronic device and relies on the operation of that device or certain applications on the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1 The client is primarily represented in the form of a device. Other types of applications can also be configured on the electronic device, such as media content publishing applications, session applications, etc. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.
[0031] The content generation method described herein allows interaction between client 101 and server 102, such as receiving or sending messages. For example, in this paper, server 102 can receive first input information sent by client 101 based on an information carrier and send the first content to client 101 for display on the display interface.
[0032] It should be noted that the content generation method can be executed by client 101, or by client 101 and server 102, with different functional parts of the content generation device deployed on client 101 and server 102 respectively; wherein, the information receiving module and content display module of the device are deployed on client 101, and the information processing module is deployed on server 102, and client 101 and server 102 achieve data interaction and functional collaboration through network communication. It should be understood that... Figure 1 The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.
[0033] To further clarify the content generation method presented in this paper, we can first introduce how to partition the obtained key vector K, value vector V, and query vector Q into blocks. In one scenario, the partitioning of key vector K, value vector V, and query vector Q can be implemented based on a third machine learning model. Therefore, in this case, we will first introduce the third machine learning model, and then explain how to determine the vector partitioning information used for partitioning key vector K, value vector V, and query vector Q based on the third machine learning model. In one scenario, the model architecture of the third machine learning model can be combined with... Figure 2 Let's find out.
[0034] See Figure 2 The architecture of the third machine learning model includes at least linear units, data partitioning and boundary padding units, a fourth module 21, and a fifth module 22. It should be noted that the role of the third machine learning model is to determine the vector partitioning information used to divide the key vector K, value vector V, and query vector Q into blocks. After determining the vector partitioning information based on the third machine learning model, this vector partitioning information is used to perform block partitioning on the key vector K, value vector V, and query vector Q during the subsequent training of the second machine learning model and the generation of the first content using the first machine learning model.
[0035] The linear unit is used to perform a linear transformation on the input feature vector, i.e., based on the learnable weight matrix. , as well as The key vector K, value vector V, and query vector Q are generated respectively to capture the dependencies between sequence elements from multiple perspectives.
[0036] Data partitioning and boundary padding unit: It is used to partition the key vector K, value vector V and query vector Q into multiple non-overlapping blocks based on the obtained vector partitioning information, and to perform boundary padding when necessary to ensure that each dimension can be divided by the block size.
[0037] The fourth module performs attention-weighted and pooling processing on the partitioned key vector K, value vector V, and query vector Q. It's important to note that this fourth module includes an attention unit and a pooling unit. The attention unit performs full attention calculation on the partitioned key vector K, value vector V, and query vector Q, outputting complete attention-weighted features. The pooling unit performs pooling operations on the feature vectors output by the attention unit. Optionally, the pooling operation performed on the feature vectors can be one of max pooling, average pooling, or min pooling.
[0038] The fifth module is used to perform full attention computation on the block-based key vector K, value vector V, and query vector Q.
[0039] After introducing the model structure of the third machine learning model, please refer to the following content to understand how to obtain the vector partitioning information on which data blocks and boundary padding units depend based on the third machine learning model.
[0040] Figure 3 This is a flowchart illustrating a content generation method in one scenario. A third machine learning model can be used to obtain vector partitioning information for dividing the key vector K, value vector V, and query vector Q. For a detailed explanation of the specific implementation, please refer to the detailed description of this scenario. Technical terms identical or corresponding to those in the above scenario will not be repeated here. Figure 3 As shown, the content generation method may specifically include: S310, Obtain the third input information.
[0041] The third input information can be input information adapted to the task requirements of the content generation method in this scenario. The third input information includes multiple modalities, optionally one or more of text modalities, image modalities, and video modalities. For example, the third input information may include descriptive text for video generation; the third input information may include images to generate video based on the visual scene corresponding to the images; the third input information may include multiple video frame data for expansion or video of the style to be converted, etc.
[0042] S320. The third input information is processed based on the third machine learning model to obtain the ninth feature vector, and the ninth feature vector is divided based on the second strategy to obtain multiple tenth feature vectors. The second strategy is pre-obtained information used to indicate the division of the ninth feature vector.
[0043] The third machine learning model is one that incorporates an attention mechanism. In one scenario, the ninth feature vector is the feature vector output after processing the third input information using linear units. This feature vector can be represented by the key vector K, the value vector V, and the query vector Q. It should be noted that the ninth feature vector corresponds to the third input information, and it represents the feature representation of the third input information in the computer.
[0044] In another scenario, the shapes of the key vector K, value vector V, and query vector Q in the ninth feature vector are composed of two dimensions: the total number of spatiotemporal lexical units and the feature dimension. For example, if the third input information is a video segment, and its dimensional information includes the number of frames T, the height of each frame H, and the width of each frame W, then after feature processing by the third machine learning model, the total number of spatiotemporal lexical units corresponding to the key vector K, value vector V, and query vector Q is N = T × H × W, meaning each spatiotemporal position corresponds to one lexical unit. The feature dimension is the model embedding dimension, i.e., the length uniformly used within the third machine learning model. Optionally, if the third machine learning model uses a multi-head attention structure, the feature dimension of each attention head needs to be calculated based on the model embedding dimension and the number of attention heads.
[0045] In one scenario, the number of second strategies may include one or more. Each second strategy characterizes how the ninth feature vector is divided into blocks. The vector partitioning information in the second strategy represents the grouping of the ninth feature vector in the temporal and spatial dimensions. It should be noted that the vector partitioning information may include the number of consecutive lemmas contained in each block in the temporal dimension, the number of consecutive lemmas contained in the spatial height dimension, and the number of consecutive lemmas contained in the spatial width dimension.
[0046] It should be noted that the difference between the various second strategies lies in the different proportions of the number of consecutive lexical units contained in each block in the time dimension, the spatial height dimension, and the spatial width dimension, which in turn leads to differences in the coverage of each block in the three dimensions of time and space.
[0047] The time dimension represents the frame sequence direction of the video, that is, the order from the first frame to the last frame. This order represents the dynamic information of the video content changing over time. The number of consecutive ephemerals contained in the time dimension is the number of consecutive frames contained in a block in the time direction. For example, if the size of the time dimension block in the segmentation information is 4, then each block will cover 4 consecutive frames.
[0048] The spatial height dimension represents the vertical direction in each frame of the image, i.e., the index from the top to the bottom of the image. The number of consecutive epochs in the spatial height dimension represents the number of consecutive rows contained in each block in the height direction. For example, if the block size in the spatial height dimension of the segmentation information is 8, then each block will cover 8 consecutive rows of pixels.
[0049] The spatial width dimension represents the horizontal direction in each frame of the image, from the far left to the far right. This dimension characterizes the left-right spatial relationship within a video frame. The number of consecutive lemmas contained in the spatial width dimension is the number of consecutive columns contained in a block in the horizontal direction. For example, if the third input information is a video sample, and the width of each frame of the video is 64, and the block size of the spatial width dimension in the segmentation information is 8, then each frame is divided into 8 consecutive block segments in the width direction, and each block will cover 8 consecutive columns of pixels.
[0050] For example, the partitioning information in the second strategy is as follows: ; in, In this partitioning method, l is the level index and h is the head index. This represents the number of consecutive lexical elements contained in the time dimension. This represents the number of consecutive word elements contained in the spatial height dimension. This represents the number of consecutive word elements contained in the spatial width dimension.
[0051] This can be understood as follows: the third input information is a video, and the video contains 32 video frames, each with a height of 64 and a width of 64. The vector partitioning information defined by the second strategy is that the block size in the time dimension is 4, the block size in the spatial height dimension is 8, and the block size in the spatial width dimension is 4. Then each block covers 4 consecutive frames in time, 8 consecutive rows in spatial height, and 4 consecutive columns in spatial width. That is, the block contains a total of 128 tokens.
[0052] In one scenario, multiple partitioning methods must satisfy a common constraint: the partitioning information corresponding to each second strategy must satisfy the factorization of a preset hardware block size. This can be understood as the product of the sizes of the time dimension block, the spatial height dimension block, and the spatial width dimension block in the partitioning information must be consistent with the preset hardware block size. The preset hardware block size is a fixed positive integer adapted to the hardware architecture. The advantage of setting a preset hardware block size is that all candidate second strategies generate blocks with the same token capacity, thus enabling a unified mapping to the size desired by the hardware acceleration unit, maximizing hardware computational efficiency.
[0053] For example, with a preset hardware block size of B, the set storing all candidate second strategies must satisfy: ,in, This represents the number of consecutive lexical elements contained in the time dimension. This represents the number of consecutive word elements contained in the spatial height dimension. This represents the number of consecutive word elements contained in the spatial width dimension. For a three-dimensional natural number space, B is the set of all candidate second strategies, and B is the preset hardware block size.
[0054] The tenth feature vector is the feature representation of multiple blocks obtained by partitioning the ninth feature vector according to the second strategy. Each block corresponds to the feature set of all words in a spatiotemporally continuous region. A spatiotemporally continuous region can be understood as a continuous frame interval in the time dimension, and within each frame of this interval, there is also a rectangular region with continuous height and width. Therefore, the second strategy can also be understood as the rule for how to divide the ninth feature vector into multiple tenth feature vectors. It should be noted that the second strategy is not unique, and different second strategies will lead to different block shapes and numbers, thus affecting the attention calculation accuracy. Therefore, multiple second strategies need to be obtained, and the preferred second strategy is determined based on these multiple second strategies.
[0055] Specifically, after obtaining the feature vector output from the previous layer, the input features are first processed using linear units to obtain the ninth feature vector. The ninth feature vector may include the key vector K, the value vector V, and the query vector Q. Further, multiple candidate second strategies are predefined, each corresponding to a set of partitioning information. This partitioning information may include the size of the time dimension block, the size of the spatial height dimension block, and the size of the spatial width dimension block, and the partitioning information must satisfy a preset hardware block size. Further, based on each second strategy, the ninth feature vector is divided into a tenth feature vector according to the partitioning information. All preset second strategies are traversed, and the tenth feature vector corresponding to each second strategy is determined, providing a data foundation for subsequently selecting the optimal partitioning information.
[0056] S330. Based on the fourth module, the tenth feature vector is processed to obtain the eleventh feature vector, and based on the fifth module, the tenth feature vector is processed to obtain the twelfth feature vector; wherein, the fifth module is used to indicate the full processing of the tenth feature vector, and the fourth module is used to indicate the processing of the information of the tenth feature vector under the second index.
[0057] The fourth module is an attention processing module. This module is used to calculate the attention of the input tenth feature vector.
[0058] This can be understood as the tenth feature vector including at least the segmented query vector Q and the key vector K. After performing attention calculation based on the segmented query vector and key vector, a pooling operation can be further performed on the attention calculation result based on the second metric, thereby outputting the eleventh feature vector.
[0059] The second metric can be one of multiple first metrics. These multiple first metrics include, but are not limited to, at least one of the following: maximum value metric, minimum value metric, and average value metric. Pooling operations can be performed based on any one of the first metrics.
[0060] For example, we can choose the maximum value metric as the second metric. The maximum value metric corresponds to the max pooling operation, which divides the input feature vector into multiple non-overlapping continuous windows and takes the maximum value of all elements within each window as the output. In this case, the eleventh feature vector is the feature vector obtained after performing max pooling on the attention calculation result.
[0061] Of course, the second indicator can also be the minimum or average value of the first indicator.
[0062] It should be noted that the first feature vector can be understood as the feature vector determined based on the second metric, that is, the feature vector extracted after attention calculation according to the selected pooling method.
[0063] The fifth module is a full attention processing module. This module performs full processing on the input tenth feature vector. Full processing can be understood as an attention calculation process that does not require pruning of the tenth feature vector, i.e., the query vector Q and the key vector K interact. The full attention calculation process includes: ; Where A is the twelfth feature vector, which represents the attention weight of the query vector on the key vector. This is the softmax function, which exponentially normalizes each element of the input vector. For query vector, Let d be the key vector and d be the feature dimension.
[0064] The twelfth feature vector is the output of the fifth module after processing the tenth feature vector. This result can be used as a benchmark for subsequent error calculation to quantify the degree of output deviation caused by the current second strategy.
[0065] For details, see Figure 2 After obtaining the tenth feature vector, it is input into both module 21 (fourth module) and module 22 (fifth module). Based on module 4, attention and pooling calculations are performed on the tenth feature vector under the second metric to output the eleventh feature vector. Simultaneously, module 5 performs full processing on the tenth feature vector to output the twelfth feature vector.
[0066] It should be noted that the processing of the tenth feature vector based on the fourth module is as follows: First, the tenth feature vector after being segmented is obtained, and then attention-weighted calculation is performed on the tenth feature vector based on the attention unit in the fourth module. Further, the attention-weighted feature vector is pooled based on the second metric to output the eleventh feature vector.
[0067] The processing of the tenth feature vector based on the fifth module is as follows: no filtering is performed on the key vectors in the tenth feature vector; full attention calculation is performed on the input query vector and key vectors, and then the twelfth feature vector is output. This achieves a comparison between the same input and the dual outputs, ensuring the objectivity of subsequent error calculations and reducing the overhead of redundant calculations.
[0068] S340. Based on the eleventh and twelfth eigenvectors, perform error processing to obtain the error loss of the second strategy.
[0069] Error processing involves calculating the difference between the eleventh and twelfth eigenvectors. This difference reflects the degree of deviation between the output of sparse attention computation and the output of full attention computation.
[0070] The error loss is the numerical result obtained through the error processing described above. This result is used to measure the degree of inconsistency between sparse attention and full attention outputs when the second strategy is employed.
[0071] Specifically, after obtaining the eleventh and twelfth feature vectors, error processing is applied to these feature vectors to evaluate the accuracy of the current second strategy. This involves calculating the element-wise difference between the eleventh and twelfth feature vectors at the same word positions, and then using a pre-defined error metric function to aggregate these differences into a scalar value. This scalar value is used as the error loss of the current second strategy to quantitatively evaluate its accuracy, enabling a fair comparison of multiple candidate second strategies.
[0072] S350. Based on the error loss of all second strategies, obtain a first strategy for use in the first machine learning model, and perform vector partitioning on the third feature vector based on the vector partitioning information corresponding to the first strategy.
[0073] The first strategy is the optimal second strategy among all the second strategies. This strategy is used in the subsequent model training and inference phases to guide the first machine learning model in partitioning the query vector Q, key vector K, and value vector V. It should be noted that once the first strategy is determined, it will not be changed; that is, it remains fixed during the subsequent training of the second machine learning model and the inference of the first machine learning model, and will not be updated.
[0074] The vector partitioning information describes the rules for grouping feature vectors in the first strategy. This information includes the size of the time dimension block, the size of the spatial height dimension block, and the size of the spatial width dimension block. It should be noted that when the original dimensions are not divisible, boundary padding can be used to process them.
[0075] The third feature vector is the feature to be segmented in the first machine learning model. This feature vector is the query vector Q, key vector K, and value vector V obtained after linear projection through the linear units in the first machine learning model. It should be noted that the third feature vector is the feature vector obtained after dimensionality reduction after processing by the linear units in the first machine learning model during the actual inference stage of the model.
[0076] Vector partitioning is the process of cutting the third feature vector into multiple non-overlapping and continuous sub-blocks based on the vector partitioning information in the second strategy.
[0077] Furthermore, based on the error loss of all second strategies, the process for determining the first strategy used in the first machine learning model is described below: Iterate through each second strategy and obtain the eleventh and twelfth eigenvectors. Calculate the error loss of these eigenvectors, and then determine the first strategy based on the error loss. ; in, Let argmin be the first policy corresponding to the l-th layer and the h-th attention head, and let argmin be the parameter that minimizes the expression, i.e., return the value that minimizes the expected error. . For the set of candidate second strategies, Let x be the calibration set, and let x be the third input information x randomly sampled from the calibration set. The third input information x is randomly sampled from the calibration set. This is the twelfth eigenvector. This is the eleventh eigenvector. To calculate the sum of squares of the element-wise differences between the eleventh and twelfth eigenvectors. The expected squared error is the error between the eleventh and twelfth eigenvectors averaged on the calibration set.
[0078] Specifically, after calculating the error loss for all second strategies, a comparative analysis of the error losses is further performed. Each second strategy corresponds to an error loss value. By traversing each second strategy in the set of second strategies, an error loss set is obtained, and the second strategy that minimizes the error loss is selected as the first strategy. Subsequently, in the first machine learning model deployed later, the input data is first linearly projected through linear units to obtain the first feature vector, and then dimensionality reduction is performed to obtain the third feature vector. Furthermore, the vector partitioning information corresponding to the first strategy is applied to divide the third feature vector into multiple blocks, and the partitioned results are sent to subsequent processing to maximize the preservation of the model's generation quality. The determined first strategy is then directly called in the subsequent inference stage, significantly reducing runtime overhead.
[0079] In one scenario, the third eigenvector of the linear unit output is partitioned based on the first strategy obtained above.
[0080] The beneficial effects of the above technical solution are as follows: First, it acquires third input information. Second, it processes the third input information based on a third machine learning model to obtain a ninth feature vector. Then, it divides the ninth feature vector according to a defined second strategy to obtain multiple tenth feature vectors. The second strategy is used to indicate the division information of the ninth feature vector. Further, it processes the tenth feature vector based on a fourth module to obtain an eleventh feature vector. Simultaneously, it processes the tenth feature vector based on a fifth module to obtain a twelfth feature vector. The fifth module is used to indicate the full processing of the tenth feature vector, and the fourth module is used to indicate the processing of the information of the tenth feature vector under the second criterion. This solves the problem of the contradiction between accuracy and efficiency under high compression ratios, achieving the technical effect of reducing attention computational overhead while maintaining generation quality.
[0081] Furthermore, after determining the vector partitioning information for the query vector Q, key vector K, and value vector V, the second machine learning model is trained based on the first strategy. Therefore, in this case, the second machine learning model is introduced first, followed by its training process. In one scenario, the model architecture of the second machine learning model can be combined with... Figure 4 Let's find out.
[0082] See Figure 4 The second machine learning model architecture includes at least a linear unit 43, a data partitioning and boundary padding unit, a first module 41, a second module 42, and a fourth module 21. It should be noted that after determining the vector partitioning information of the query vector Q, key vector K, and value vector V based on the third machine learning model, the second machine learning model is trained based on the determined vector partitioning information, i.e., the first strategy.
[0083] The first module in the second machine learning model is used to extract statistics from the segmented feature vectors and perform block-level attention calculations. It should be noted that the first module 41 may include a first unit, a feedforward neural network unit, and a second unit.
[0084] The first unit is a triplet pooling unit, which is used to perform max pooling, average pooling and min pooling on the block-divided query vector Q and key vector K to extract the peak response, overall distribution level and lower bound information of the features within each block.
[0085] The feedforward neural network unit is a neural network unit containing multiple fully connected layers. This unit is used to non-linearly map the second result output by the triple pooling unit, transforming it to the latent space to enhance the expressive power of intra-block statistical features. Optionally, the fully connected layers in this unit include linear layers and activation functions.
[0086] The second unit is an attention calculation module, which is used to receive the feature vector output by the feedforward neural network unit and calculate the correlation attribute between the query vector and the key vector.
[0087] The second module receives the output from the first module. This second module includes a third unit and a feature fusion unit. The third unit filters out a preset number of key blocks with the highest values for each query block. The feature fusion unit performs attention operations on the received feature vectors.
[0088] Figure 5 This is a flowchart illustrating a content generation method in one scenario. Building upon the above scenario, the training process of the second machine learning model based on the determined vector partitioning information can also be described. For a detailed explanation of the training process of the second machine learning model, please refer to this paper. Technical features that are identical or similar to those described above will not be repeated here.
[0089] like Figure 5 As shown, the content generation method may specifically include: S510, Receive the second input information.
[0090] Before training the second machine learning model, multiple training samples can be obtained. Each training sample includes input information and corresponding theoretical output information. The input information is used as the second input information.
[0091] It should be noted that the second input information can be data adapted to the task requirements of the content generation method in this scenario. The second input information includes multiple modalities, optionally one or more of the following: text modality, image modality, and video modality.
[0092] Optionally, the second input information can come from a large-scale dataset, such as a large-scale multimodal dataset containing multiple video-text pairs, video-image pairs, or pure video clips. It should be noted that the source of the second input information can be the same as or different from the source of the third input information, and they are independent of each other in their purpose.
[0093] S520. Based on the second machine learning model, process the second input information to obtain the sixth feature vector, the seventh feature vector, and the second output content.
[0094] The second machine learning model adjusts its internal parameters using supervised information. Upon final convergence, a first machine learning model for content generation can be derived from the second model.
[0095] Based on the above, the second machine learning model includes at least a linear unit, a data partitioning and boundary padding unit, a first module, a second module, and a fourth module. After processing the input feature vector based on the linear unit, the query vector Q, key vector K, and value vector V corresponding to the second input information are obtained. Further, the query vector Q, key vector K, and value vector V are input to the data partitioning and boundary padding unit, so that the data partitioning and boundary padding unit performs block partitioning according to the first strategy obtained above, and outputs the partitioned query vector Q, key vector K, and value vector V.
[0096] After obtaining the segmented query vector Q, key vector K, and value vector V, the segmented query vector Q and key vector K are input into the first unit, and a first result is generated based on multiple first indicators in the first unit. These multiple first indicators correspond to the pooling criteria used by the pooling layer to pool the fourth feature vector. Optionally, the first indicators may include one or more of the following: mean indicator, maximum indicator, and minimum indicator.
[0097] Optionally, to further improve the quality of the first feature vector, the first metric can be a combination of the mean metric, the maximum metric, and the minimum pooling metric. Based on the above first metrics, the pooling layer processes each fourth feature vector to obtain the pooling result of each fourth feature vector under multiple first metrics.
[0098] In one scenario, pooling based on the mean metric can be termed mean pooling. Mean pooling is the process of calculating the arithmetic mean pooling of each feature channel in the fourth feature vector along the word dimension for the input block. The output of this mean pooling is then used as the first result corresponding to the first metric. In another scenario, pooling based on the maximum value index can be called max pooling. Max pooling is to calculate the maximum value of each feature channel in the word dimension for the input block, and use the output of the max pooling as the first result corresponding to the first index. In another scenario, pooling based on the minimum value metric can be called minimum pooling. Minimum pooling involves calculating the minimum value of each feature channel in the word dimension for the input block, and using the output of this minimum pooling as the first result corresponding to the first metric.
[0099] Furthermore, the first results corresponding to multiple first indicators are concatenated to output a second result. The second result is a joint feature vector obtained by concatenating the first results output by multiple first indicators along the feature dimension. The advantage of obtaining the second result is that it integrates the mean, maximum, and minimum value information of the features within the vector block, providing richer feature information for subsequent processing.
[0100] After obtaining the second result, it is input into the feedforward neural network unit. After the second result is input into the feedforward neural network unit, it is mapped to the latent space through multiple fully connected layers in the feedforward neural network unit to output the third result.
[0101] The third result is the feature vector output after the second result has been mapped through multiple fully connected layers of the feedforward neural network unit. This feature vector is used for similarity calculation in subsequent block attention.
[0102] Furthermore, the third result is input into the second unit for attention processing. The second unit is an attention calculation unit. This unit receives the third result output by the feedforward neural network unit and calculates the correlation attribute between the query vector and all key vectors. This correlation attribute is represented as a numerical vector, namely the sixth feature vector.
[0103] The sixth feature vector is used to characterize the estimated similarity between the query vector and the key vector. It should be noted that each element in the sixth feature vector represents the predicted importance of the corresponding query vector and the corresponding key vector. The advantage of inputting the third result into the second unit for block attention extraction is that performing a dot product operation on the block features significantly reduces computational complexity.
[0104] In one scenario, the second module 42 is used to receive the output of the first module. The second module 42 includes a third unit and a feature fusion unit. The third unit is used to filter out a predetermined number of key blocks with the highest values from the sixth feature vector output by the second unit for each query block. This unit can generate a binary mask based on the filtered key blocks. This mask can indicate the positions of the retained key blocks and the positions of the masked key blocks.
[0105] Next, the processing procedure for the sixth feature vector based on the third unit will be introduced. For each query block in the query vector, the corresponding elements (i.e., relevance values) in the sixth feature vector corresponding to the query block are sorted in descending order. Furthermore, the indices of the first preset number of positions are retained, and the remaining positions are set as masks. After sorting and filtering the corresponding elements of each query block in the query vector, a binary mask matrix is output.
[0106] It's important to note that the binary mask matrix contains two types of content: selected content is represented by a 1, and discarded content is represented by a 0. This can be understood as follows: when the corresponding content is selected, the corresponding position in the binary mask matrix is set to 1, indicating that the key block needs to participate in the attention calculation of the corresponding query block. If the corresponding content is discarded, the corresponding position in the binary mask matrix is set to 0, indicating that the key block is masked, and its attention weight is set to 0 during subsequent attention calculations.
[0107] For example, after determining the preset number of key blocks before each query block, the remaining parts are masked, thus obtaining the mask. for: ; in, Let K be the element in the i-th row and j-th column of the mask, and K be the number of key blocks that need to be retained for each query block. This represents the total number of blocks.
[0108] Therefore, each row of the binary mask matrix determined by this rule has a predetermined number of "1"s. The advantage of determining the binary mask matrix is that it reduces the number of key blocks that the query block needs to focus on, significantly reduces computational complexity, and avoids unnecessary memory reads and calculations.
[0109] Furthermore, based on the generated binary mask matrix, key vectors and query vectors for subsequent sparse attention computation are determined, and these are then filtered and concatenated to obtain new key vectors and query vectors. Specifically, the filtering process may include, for each query block, extracting the corresponding key block and query block from the segmented key vector and query vector (i.e., the first feature vector) based on the set of column indices with values of 1 in the corresponding rows of the query block in the binary mask matrix. Further, the selected key blocks are concatenated along the word dimension to form a new key vector. Similarly, the query blocks are concatenated to form a new query vector.
[0110] After obtaining the concatenated matrix, attention is calculated based on the feature fusion unit. During the attention calculation, the full attention matrix is only applied to the tokens within the selected block. Finally, the second output is generated after sparse attention calculation.
[0111] For example, the feature fusion unit calculation process is as follows: ; in, This is the output matrix of the i-th query block. For normalization function, Let i be the query vector for the i-th query block. Let be the key vector of the i-th key block, and d be the feature dimension. Let i be the value vector of the i-th value block.
[0112] The second output is the final output generated by the second machine learning model. Optionally, after the feature fusion unit, a deblocking and deplacing process can be performed. Since the output obtained after sparse attention computation is a feature matrix in blocks, deblocking and deplacing are necessary to restore the output to the same structure as the model input. Deblocking can be understood as rearranging the block output features into a continuous sequence of words. It can also be understood as restoring the words within each block to their original spatiotemporal grid positions based on the vector partitioning information. The deplacing process removes filler words added at the spatiotemporal grid boundaries that satisfy divisibility conditions before block partitioning. By performing deblocking and deplacing on the feature vectors output by the feature fusion unit, the final output of the second machine learning model, i.e., the second output content, is determined.
[0113] Module 4 (21) is dedicated to model training. It includes a computational unit based on full attention. This module receives the query vector, key vector, and value vector obtained after feature extraction from the second input information. It's important to note that this unit does not prune any tokens or blocks; it calculates the complete attention weights and outputs the weighted features, i.e., the eighth feature vector.
[0114] The eighth feature vector is the output feature vector after processing the second input information. This feature vector has not undergone any sparse pruning or filtering.
[0115] Furthermore, the eighth feature vector is input into the pooling unit in the fourth module for processing. The pooling unit is the unit that performs the pooling operation. This unit receives the eighth feature vector, performs the pooling process within a preset pooling window, and outputs the seventh feature vector. Optionally, the pooling operation is a second metric. The second metric can be the maximum value among multiple first metrics. The maximum value metric can be understood as a max pooling operation, which divides the word dimension of the eighth feature vector into multiple blocks, and takes the maximum value of each feature channel along the word dimension in each block to obtain a block-level feature vector.
[0116] For example, the process of performing a max pooling operation is as follows: ; in, The seventh feature vector represents the true correlation between the i-th query block and the j-th key block. For maximum value operation, The attention map is generated after passing through the attention unit. (u,v) are the word-level coordinates covered by the (i,j)th block, u is the query word index, and v is the key word index.
[0117] The seventh feature vector is the feature vector output by the fourth module. This feature vector can be used to calculate the loss and adjust the model parameters in the second machine learning model.
[0118] Specifically, the second machine learning model receives the second input information. First, it obtains the query, key, and value vectors through feature extraction and linear projection, and then converts them into block-level representations through block segmentation and padding. Further, the second machine learning model executes two paths in parallel: In the sparse attention path, it passes through the first module (including the first unit, the feedforward neural network unit, and the second unit) and outputs the sixth feature vector; this sixth feature vector is then input into the second module (including the third unit and the feature fusion unit) for filtering and feature extraction, generating the second output content. In the full attention path, attention is calculated on the block-level representation corresponding to the second input information, and then processed based on the second metric to output the seventh feature vector. The advantage of this approach is that it obtains the seventh feature vector, the sixth feature vector, and the second output content simultaneously in a single forward propagation, improving training efficiency.
[0119] S530. Based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector, adjust the model parameters in the second machine learning model.
[0120] The third output is the expected output used for supervision during training. For example, in a video diffusion model, the third output can be real noise. It should be noted that the third output can be generated based on pre-configured annotations of the training dataset or through a data preprocessing procedure.
[0121] The second machine learning model is a model that includes both a sparse attention branch and a full attention branch. It should be noted that the second machine learning model may include a first module, a second module, a third module, and a fourth module. The first and second modules are responsible for calculating the sparse attention branch, the third module (the data partitioning and boundary padding unit) performs dimensionality reduction and partitioning of the input features, and the fourth module is responsible for calculating the full attention branch.
[0122] The second machine learning model includes model parameters. These parameters can be understood as the learnable weights and biases in the second machine learning model. Adjusting the model parameters in the second machine learning model can be understood as using the backpropagation algorithm to calculate the gradient based on the loss function and update the learnable weights and biases in the second machine learning model.
[0123] Specifically, after inputting the second input information into the second machine learning model, the second output content, the sixth feature vector, and the seventh feature vector are output. The loss value is calculated based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector. The gradient is calculated and the model parameters in the second machine learning model are updated through the backpropagation algorithm to achieve end-to-end joint optimization and realize closed-loop optimization between sparse execution and full attention benchmark.
[0124] Furthermore, during model training, a first loss and a second loss can be used to update the model parameters. The following section elaborates on the updating of the model parameters. Optionally, the model parameters include a first parameter and a second parameter. The first parameter indicates the learnable parameters in the first module, and the second parameter indicates the learnable parameters of the first machine learning model excluding the first module. Based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector, adjust the model parameters in the second machine learning model, including: Based on the second and third output contents, determine the first loss; based on the sixth and seventh feature vectors, determine the second loss; adjust the first parameter based on the second loss, and adjust the second parameter based on the first loss.
[0125] Here, the first parameter refers to the learnable parameters in the first module of the second machine learning model. It should be noted that the learnable parameters in the first module include the first unit, the feedforward neural network unit, and the learnable variables in the second unit.
[0126] The second parameter is all the learnable parameters in the second machine learning model other than those in the first machine learning model.
[0127] The first loss is the task loss calculated based on the second and third output contents. Optionally, the first loss can be the mean square error between the predicted noise and the actual noise. The second loss is the distillation loss calculated based on the sixth and seventh feature vectors.
[0128] Next, the process of calculating the distillation loss based on the sixth and seventh feature vectors is explained. First, the seventh feature vector output by the fourth module is obtained, and row normalization is performed on it to obtain the attention weights corresponding to the query vector. Simultaneously, a softmax operation is performed on the key block of the sixth feature vector output by the first module to obtain the attention weights corresponding to the query vector within the sixth feature vector. Further, the distillation loss is minimized based on the sixth and seventh feature vectors: ; in, For distillation losses, The Kullback-Leibler (KL) divergence is a parameter used to measure the similarity between two probability distributions. The attention weights come from the fourth module. The attention weights come from the first module.
[0129] It should be noted that when calculating the attention weights mentioned above, it is possible to choose to use only a portion of the query vectors and all key vectors for attention calculation. The advantage of this setting is that by selecting only a portion of the query vectors to participate in the calculation, the computational cost during training can be effectively reduced, while providing sufficient supervision signals for distillation.
[0130] Specifically, during training, the feature vector output from the previous layer is input into the second machine learning model. After forward propagation, the second output content, the sixth feature vector, and the seventh feature vector are output simultaneously. Further, a first loss is calculated based on the second and third output contents. A second loss is calculated based on the sixth and seventh feature vectors to measure the difference between the block importance distribution of the sparse prediction and the full attention benchmark. During the optimization and update phase, the first parameter is optimized using gradient descent with the second loss to improve the accuracy of attention prediction; the second parameter is optimized using the first loss to ensure the execution accuracy of the main task. Optionally, the two losses can be backpropagated independently or weighted and then jointly updated. This aims to approximate full attention while ensuring generation quality.
[0131] S540. In response to detecting that the second machine learning model has reached the convergence condition, the fourth module in the second machine learning model is removed to obtain the first machine learning model.
[0132] The third output content is used to reflect the expected output result of the second input content, the sixth feature vector is used to reflect the feature vector after processing by the first module, the seventh feature vector is used to reflect the feature vector after processing by the fourth module, the fourth module is used to obtain the seventh feature vector of the eighth feature vector on the second index, the eighth feature vector is used to indicate the vector after the second input information is output by the third module, and the second index is the maximum value index among multiple first indices.
[0133] The convergence condition is the criterion for determining whether the second machine learning model has been successfully trained. Optionally, the convergence condition can be that the loss function value drops below a preset threshold, or the performance no longer improves after several consecutive iterations, or the maximum preset number of iterations is reached, etc.
[0134] Specifically, once the second machine learning model has completed training and met the preset convergence criteria, the fourth module no longer needs to participate in the subsequent inference process. Therefore, the fourth module is removed from the second machine learning model, and the first machine learning model is determined to improve inference speed and reduce memory usage.
[0135] The beneficial effects of the above scheme are as follows: It receives the second input information; processes the second input information based on the second machine learning model to obtain the sixth feature vector, the seventh feature vector, and the second output content; adjusts the model parameters in the second machine learning model based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector; and, in response to detecting that the second machine learning model has reached convergence, removes the fourth module from the second machine learning model to obtain the first machine learning model, thereby achieving end-to-end joint training and improving the accuracy of sparse decision-making. Simultaneously, it decouples parameter updates, ensuring stable and efficient training and improving training efficiency.
[0136] Furthermore, after obtaining the trained second model, the fourth module can be removed from the second model to obtain a usable first model. That is, the first model and the second model have the same model structure; the difference lies in the first model not including the fourth module. Therefore, in this case, the first machine learning model is introduced first, followed by the first machine learning's processing of the first input information. In one scenario, the model structure of the first machine learning model can be combined with... Figure 6 Let's find out.
[0137] See Figure 6 The architecture of the first machine learning model includes at least a linear unit 43, a data partitioning and boundary padding unit, a first module 41, and a second module 42. It should be noted that after training the second machine learning model, the fourth module is removed, and the model without the fourth module is used as the first machine learning model for content generation.
[0138] The first module is the module in the second machine learning model used to extract statistics from the segmented feature vectors and perform block-level attention calculations. It should be noted that the first module 41 may include a first unit, a feedforward neural network unit, and a second unit.
[0139] The first unit is a triplet pooling unit, which is used to perform max pooling, average pooling and minimum pooling on the key block, value block and query block after block division, so as to extract the peak response, overall distribution level and lower bound information of the features in each block.
[0140] The feedforward neural network unit is a neural network unit containing multiple fully connected layers. This unit is used to non-linearly map the second result output by the triple pooling unit, transforming it to the latent space to enhance the expressive power of intra-block statistical features. Optionally, the fully connected layers in this unit include linear layers and activation functions.
[0141] The second unit is an attention calculation module, which receives the feature vector output by the feedforward neural network unit and calculates the correlation attribute between the query vector and all key vectors.
[0142] The second module receives the output from the first module. This second module includes a third unit and a feature fusion unit. The third unit filters out a preset number of key blocks with the highest values for each query block. The feature fusion unit performs attention operations on the received feature vectors.
[0143] Figure 7 This is a flowchart illustrating a content generation method for one scenario. This content generation method is applicable to scenarios involving arbitrary image-to-video, video-to-video, and text-to-video generation. The content generation method can be executed by a content generation device, which can be implemented in software and / or hardware, or optionally through an electronic device, such as a client.
[0144] like Figure 7 As shown, the content generation method may specifically include: S610, Obtain the first input information.
[0145] The first input information can be input information adapted to the task requirements of the content generation method in this scenario. The first input information includes multiple modalities, optionally one or more of text modalities, image modalities, and video modalities. For example, the third input information may include descriptive text for video generation; the third input information may include images to generate video based on the visual scene corresponding to the images; the third input information may include multiple video frame data for expansion or video of a style to be converted, etc.
[0146] S620. Analyze and process the first input information based on the first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the second feature vector to be obtained from the first feature vector according to multiple first indicators. The second module is used to instruct the processing of the second feature vector. The amount of data in the second feature vector is less than the amount of data in the first feature vector.
[0147] The first machine learning model includes a first module and a second module. The first module extracts a second feature vector from a first feature vector based on multiple first indicators. The first feature vector is the input feature of the first module. It should be noted that the first feature vector is a feature vector processed based on the second input information. Optionally, the process of obtaining the first feature vector can be that the second input information is input into the first machine learning model for processing. After the linear unit receives the features output from the previous layer, the linear unit processes the input features and outputs the corresponding key matrix, value matrix, and query vector. Further, based on the first strategy, the key, value, and query vector are processed by block segmentation and boundary padding to output the first feature vector.
[0148] The second feature vector is the result of the first module processing the first feature vector based on multiple indicators. It should be noted that the amount of data in the second feature vector is much smaller than that in the first feature vector.
[0149] Specifically, the first module receives the first input information and obtains a first feature vector after feature extraction. Further, the first module performs pooling processing on the first feature vector based on multiple first indicators to generate a second feature vector with reduced data volume, thereby reducing the complexity of subsequent processing and the memory usage, and ultimately accelerating overall inference.
[0150] In one scenario: the first module is used to obtain the evaluation attributes of at least one first sub-vector under multiple first indicators, wherein the at least one first sub-vector is obtained by partitioning the first feature vector based on vector partitioning information. A second feature vector is used to indicate the feature vector corresponding to the second sub-vector, wherein the second sub-vector is obtained based on the evaluation attributes of at least one first sub-vector.
[0151] In another scenario, the second sub-vector is obtained based on the evaluation properties of the at least one first sub-vector in the following manner: Obtain the position information of at least one first sub-vector in a two-dimensional space, and obtain a grid diagram composed of the at least one sub-vector, wherein the grid diagram has the same grid size and each grid corresponds to a first sub-vector; obtain a first number of second sub-vectors from high to low according to the evaluation attributes of multiple grids in each row.
[0152] That is, each row can select the first number of second sub-vectors based on the evaluation attributes from high to low.
[0153] S630. Display the first content, which reflects the result of the first machine learning model based on the second feature vector.
[0154] The first content refers to the final output of the first machine learning model. It should be noted that the first content is video data corresponding to the first input information. For example, the first input information may include descriptive text, in which case the first content is a video consistent with the text description; the first input information may include an image, in which case the first content is a video generated after dynamically expanding the image content; the first input information may include a video, in which case the first content is a video adapted to the video content. Display can be understood as showing the first content in an interactive interface. Optionally, in practical applications, if a user has a need for content generation, the interactive interface corresponding to the content generation system can be triggered. This interactive interface is a window for interacting with the content generation system. The user can input the first input information in the interactive interface, and the content generation system can provide feedback on the first content corresponding to the first input information and display the first content in the interactive interface. The advantage of having an interactive interface is that it allows for a visual representation of the interaction between the user and the content generation system, achieving a visual effect of the content generation process.
[0155] Specifically, after obtaining the first content generated by the first machine learning model, the first content is displayed on the interactive interface to intuitively present the generation effect, which facilitates real-time feedback and model iteration optimization.
[0156] It should be noted that the data format of the first input information can be various, and the data formats of the first input information will now be described. Optionally, the first input information includes one or more of the following: first audio, first video, first image, or first text; the first content includes a second video, and the video content of the second video is related to the first input information.
[0157] The type of the first input information can be first audio, first video, first image, or first text.
[0158] The first content refers to the final output of the first machine learning model. It should be noted that the first content can include the second video, meaning the content of the second video is related to the first input information. This can be understood as the generated second video maintaining consistency with or being related to the first input information in terms of theme, semantics, style, or timing. For example, if the first input information is the text "A cat is playing," then the content in the second video could include actions such as the cat running.
[0159] The beneficial effects of the above scheme are: acquiring first input information; analyzing and processing the first input information based on a first machine learning model, the first machine learning model including a first module and a second module, the first module being used to instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators, the second module being used to instruct the processing of the second feature vector, the data volume of the second feature vector being less than that of the first feature vector; and displaying first content, the first content being used to reflect the results output by the first machine learning model based on the second feature vector, so as to reduce computational complexity and memory usage, improve inference speed, and retain important information in the original features, thus ensuring output quality.
[0160] Figure 8 This is a flowchart illustrating a content generation method under one scenario. Based on the above scenario, the processing of the first input information by the first machine learning model can also be explained. The technical solution in this scenario can be combined with implementation methods in other scenarios, and the specific implementation methods can be found in the detailed description herein. Technical features that are the same as or similar to those described above will not be repeated here.
[0161] like Figure 8 As shown, the content generation method may specifically include: S710, Obtain the first input information.
[0162] S720. Based on the third module, the first feature vector of the first input information is obtained, and the first feature vector is dimensionality reduced to obtain the third feature vector. The first feature vector reflects the vector after the first input information is processed based on the weight matrix.
[0163] The third module is a feature extraction module. It receives the feature vectors output from the previous layer, performs linear mapping on the feature vectors, and generates a query vector Q, a key vector K, and a value vector V. The query vector Q, key vector K, and value vector V are then used as the first feature vector.
[0164] The weight matrix consists of learnable parameters within the third machine learning model. These parameters linearly transform the input information. It should be noted that the weight matrix is trained based on the training process of the first machine learning model.
[0165] Specifically, the features output from the previous layer are first input into a linear unit. This unit contains a learnable weight matrix, and by performing a linear transformation on the input features, it outputs a first feature vector. Furthermore, to reduce the processing complexity of the first machine learning model, the first feature vector is dimensionality-reduced, outputting a third feature vector with a dimension much smaller than the first feature vector, thereby reducing the amount of data processing while retaining necessary contextual representations.
[0166] S730. Based on the first module, obtain the fifth feature vector of the fourth feature vector under multiple first indicators. The fourth feature vector is used to represent the vector after the third feature vector is divided according to the obtained vector division information. The fifth feature vector is used to instruct the second module to obtain the second feature vector. The fourth feature vector is the vector obtained by partitioning the third feature vector based on the acquired vector partitioning information. The fifth feature vector is the result of a module processing the fourth feature vector according to multiple first indicators. The second feature vector is the output of the second module.
[0167] Furthermore, a third feature vector is obtained by dimensionality reduction based on the extracted first feature vector. After obtaining the third feature vector, it is divided into blocks according to a predetermined first strategy, that is, the third feature vector is divided into several non-overlapping continuous blocks, and each block contains a fixed number of tokens. For ease of calculation, the block-based feature vector is padded with boundaries, resulting in a fourth feature vector. The fourth feature vector is the block-based feature vector. It can be understood that the third feature vector may contain the block-based key, query, and value vectors.
[0168] Furthermore, the process of processing the fourth feature based on the first module and outputting the fifth feature is further refined. Optionally, the first module includes a first unit, a feedforward neural network unit, and a second unit. Based on the first module, the fifth feature vector of the fourth feature vector under multiple first indicators is obtained, including: Based on the first unit, the first result of the third feature vector under multiple first indicators is obtained. The first results are concatenated to obtain the second result of the third feature vector. The first indicators include at least one of the mean indicator, minimum indicator and maximum indicator. The second result is processed based on the feedforward neural network unit to obtain the third result. The third result is then subjected to attention processing based on the second unit to obtain the fifth feature vector. The fifth feature vector is used to indicate the optional attributes of the fourth feature vector.
[0169] The fifth feature vector is the feature vector obtained after block attention processing in the second unit. This feature vector indicates the optional attributes of the fourth feature vector. The optional attributes can be understood as which blocks should be retained, i.e., the key vectors and query vectors that can be used for subsequent calculations.
[0170] Specifically, the first module includes a first unit, a feedforward neural network unit, and a second unit. After receiving the third feature vector, it is divided into multiple blocks according to a preset vector partitioning information to obtain a fourth feature vector. For each block in the fourth feature vector, the first unit calculates multiple first indicators along the word dimension. The first indicators include at least a mean indicator, a minimum indicator, and a maximum indicator, i.e., the fourth feature vector is processed based on the mean indicator, minimum indicator, and maximum indicator, respectively. Each indicator outputs a vector as the first result. The three first results corresponding to the same block are concatenated along the feature dimension to obtain the second result, which integrates the overall information, lower bound information, and peak information within the block.
[0171] For example, after processing the fourth feature vector based on multiple first indicators, the feature vector can be determined by concatenation. That is... ; in, This is a triple pooling operation. For average pooling operation, For max pooling operation, "+" represents the minimum pooling operation, and "+" represents the concatenation operation.
[0172] Furthermore, the processing procedures for the feature vectors by the first unit, the feedforward neural network, and the second unit are as follows: ; in, The prediction score between the i-th query block and the j-th key block. To query the parameters of the side feedforward neural network, These are the parameters of the key-side feedforward neural network. This is a triple pooling operation. For the feature corresponding to the i-th query block, Let d be the feature corresponding to the i-th key block, and d be the feature dimension.
[0173] Furthermore, the second result is input into a feedforward neural network unit for nonlinear transformation. The feedforward neural network unit can contain two linear layers and an activation function, outputting a third result by first increasing and then decreasing the dimensionality, thus enhancing the model's expressive power. Further, the second unit receives the third result and calculates the intra-block relevance attributes using a scaled dot product. For each block, it calculates the dot product between each position in the query block and each position in the key block, and then scales it by dividing by the square root of the feature dimension. The result is input into a softmax function to obtain the attention weight matrix, i.e., the fifth feature vector. Finally, based on the optional attributes of the fourth feature vector indicated by the fifth feature vector, the key block positions to be retained in subsequent feature fusion are determined, thereby guiding the model to focus on the most relevant semantic blocks and reducing data processing volume while ensuring content generation quality.
[0174] S740: Based on the second module, the second feature vector is obtained from the first feature vector based on the fifth feature vector, and the first content is output based on the second feature vector.
[0175] The second feature vector is an intermediate feature obtained by filtering and concatenating from the first feature vector according to the indication of the fifth feature vector. The first content is the final output of the second module, that is, the content generated by the first module from the first input information.
[0176] Specifically, the second module includes at least a third unit and a feature fusion unit. In the second module, for each query block, a preset number of key blocks are retained based on the relevance attributes in the fifth feature vector. Based on the determined key block positions, the original information of the retained key blocks is determined in the first feature vector. The original information is reassembled into a new matrix, which serves as the second feature vector. Furthermore, feature extraction is performed based on the second feature vector to output the first content, thereby reducing computational complexity, retaining key information, and improving sparsity accuracy.
[0177] Furthermore, after the second module receives the fifth feature vector, the fifth feature vector needs to be processed sequentially by the third unit and the feature fusion unit to finally obtain the first content. Next, the process of determining the first content based on the second module will be described in detail. Optionally, the second module includes a third unit and a feature fusion unit, which obtains the second feature vector from the first feature vector based on the fifth feature vector, and outputs the first content based on the second feature vector, including: Based on the third unit, a fourth feature vector is obtained according to the fifth feature vector to participate in the generation of the first content. A second feature vector is selected from the first feature vector according to the fourth feature vector. The second feature vector is used to represent the mask information for generating the first content. Based on the feature fusion unit, attention processing is performed on the second feature vector, and the obtained third result is processed to obtain the first content.
[0178] The second feature vector represents the mask information used to generate the first content. It can be understood that the second feature vector is a set of features selected and concatenated from the first feature vector based on the fourth feature vector. This feature is the original lexical-level feature of the retained key block and query block.
[0179] The third result is the output after attention calculation by the feature fusion unit.
[0180] Next, the process of sparse attention calculation on the second feature vector based on the feature fusion unit is described. After filtering and rearranging the key block, value block, and query block respectively, the sparse attention calculation of the query block is as follows: ; in, This is the output vector of the i-th query block. It is a normalization function. Let i be the query matrix for the i-th query block. Let be the transpose of the concatenated key vector, and d be the feature dimension. This is the concatenated value vector. It can be understood as follows: based on the i-th query block, first calculate the dot product of its query vector and the transpose of the concatenation matrix of the selected key block, then normalize it to obtain the attention weight, and finally multiply it with the selected value block to obtain the output of that query block.
[0181] Specifically, in the second module, the third unit receives the fifth feature vector. Based on each query block in the fifth feature vector, the third unit retains a preset number of key blocks with the highest values according to the relevance attribute of that row in the fifth vector, and uses these to generate the fourth feature vector. The fourth feature vector can be understood as a binary mask matrix, which is used as a filtering condition for sparse selection in the generation of the first content.
[0182] Furthermore, based on the fourth feature vector, the selected key block and value block corresponding to each query block are extracted from the first feature vector. The query block, key block, and value block are then reassembled according to the sequence dimension to obtain the second feature vector, thereby determining the selected data in the first feature vector.
[0183] Furthermore, the feature fusion unit receives the second feature vector and performs attention calculations only on the tokens in the selected key block for each query block, outputting a third result. Finally, the third result is expanded to perform deblocking to restore the original spatiotemporal order and remove boundary padding, thus obtaining the first content with the same size as the original input. This achieves dynamic adaptive sparsity, improves content quality under high sparsity, and significantly reduces computational cost.
[0184] S750. Display the first content, which reflects the result of the first machine learning model based on the second feature vector.
[0185] The beneficial effects of the above scheme are as follows: First, it obtains the first input information; based on the third module, it obtains the first feature vector of the first input information and performs dimensionality reduction on the first feature vector to obtain the third feature vector, where the first feature vector reflects the vector after processing the first input information based on the weight matrix; further, based on the first module, it obtains the fifth feature vector of the fourth feature vector under multiple first indicators, where the fourth feature vector represents the vector after partitioning the third feature vector according to the obtained vector partitioning information, and the fifth feature vector instructs the second module to obtain the second feature vector; after obtaining the fifth feature vector, based on the second module, it obtains the second feature vector from the first feature vector according to the fifth feature vector, and outputs the first content based on the second feature vector; it displays the first content, which reflects the result output by the first machine learning model based on the second feature vector; it achieves end-to-end feature compression and recovery, significantly reducing computational and memory overhead while retaining key information and ensuring generation quality.
[0186] Figure 9 This is a schematic diagram of the structure of a content generation device in one scenario, such as... Figure 9 As shown, the device includes: an information receiving module 810, an information processing module 820, and a content display module 830.
[0187] The information receiving module 810 is used to acquire first input information; the information processing module 820 is used to analyze and process the first input information based on a first machine learning model, the first machine learning model including a first module and a second module, the first module being used to instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators, the second module being used to instruct the processing of the second feature vector, the data volume of the second feature vector being less than the data volume of the first feature vector; and the content display module 830 is used to display first content, the first content being used to reflect the result output by the first machine learning model based on the second feature vector.
[0188] Optionally, in one scenario, the first machine learning model includes a third module, the information processing module comprising: The third feature vector acquisition unit is used to obtain the first feature vector of the first input information based on the third module, and to perform dimensionality reduction processing on the first feature vector to obtain the third feature vector. The first feature vector reflects the vector after processing the first input information based on the weight matrix. The fifth feature vector acquisition unit is used to obtain the fifth feature vector of the fourth feature vector under the plurality of first indicators based on the first module. The fourth feature vector is used to represent the vector after the third feature vector is divided according to the obtained vector division information. The fifth feature vector is used to instruct the second module to obtain the second feature vector. The content output unit is used to obtain the second feature vector from the first feature vector based on the fifth feature vector based on the second module, and output the first content based on the second feature vector.
[0189] In another scenario, optionally, the first module includes a first unit, a feedforward neural network unit, and a second unit, and the fifth feature vector acquisition unit includes: The second result acquisition unit is used to acquire the first result of the third feature vector under multiple first indicators based on the first unit, and concatenate the first results to obtain the second result of the third feature vector, wherein the first indicators include at least one of the mean indicator, the minimum indicator and the maximum indicator. The third result acquisition unit is used to process the second result based on the feedforward neural network unit to obtain the third result; The fifth feature vector determination unit is used to perform attention processing on the third result based on the second unit to obtain the fifth feature vector, which is used to indicate the optional attributes of the fourth feature vector.
[0190] In another scenario, optionally, the second module includes a third unit and a feature fusion unit, and the content output unit includes: The second feature selection subunit is used to obtain a fourth feature vector for participating in the generation of the first content based on the third unit and the fifth feature vector, and to select a second feature vector from the first feature vector based on the fourth feature vector. The second feature vector is used to characterize the mask information for generating the first content. The content information acquisition subunit is used to perform attention processing on the second feature vector based on the feature fusion unit, and to expand the processing of the obtained third result to obtain the first content.
[0191] In another scenario, optionally, the first machine learning model is trained in the following manner: An input information receiving module is used to receive second input information; The feature vector acquisition module is used to process the second input information based on the second machine learning model to obtain the sixth feature vector, the seventh feature vector, and the second output content; The model parameter adjustment module is used to adjust the model parameters in the second machine learning model based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector. The first machine learning model acquisition module is used to remove the fourth module in the second machine learning model in response to detecting that the second machine learning model has reached the convergence condition, and obtain the first machine learning model. The third output content is used to reflect the expected output result of the second input information, the second output content is used to reflect the predicted output result of the second input information, the sixth feature vector is used to reflect the feature vector after processing by the first module, the seventh feature vector is used to reflect the feature vector after processing by the fourth module, the fourth module is used to obtain the seventh feature vector of the eighth feature vector on the second index, the eighth feature vector is used to indicate the vector after the second input information is output by the third module, and the second index is the maximum value index among multiple first indices.
[0192] In another scenario, optionally, the model parameters include a first parameter and a second parameter, wherein the first parameter indicates the learnable parameters in the first module, and the second parameter indicates the learnable parameters of the first machine learning model other than the first module. The model parameter adjustment module includes: The first loss determination unit is used to determine the first loss based on the second output content and the third output content; The second loss determination unit is used to determine the second loss based on the sixth feature vector and the seventh feature vector; A parameter adjustment unit is configured to adjust the first parameter based on the second loss, and to adjust the second parameter based on the first loss.
[0193] In another scenario, the method may optionally further include obtaining vector partitioning information for partitioning the feature vectors by means of: The third input information acquisition module is used to acquire third input information. The tenth feature vector determination module is used to process the third input information based on the third machine learning model to obtain the ninth feature vector, and to divide the ninth feature vector based on the second strategy to obtain multiple tenth feature vectors. The second strategy is pre-obtained information used to indicate the division of the ninth feature vector. The twelfth feature vector determination module is used to process the tenth feature vector based on the fourth module to obtain the eleventh feature vector, and to process the tenth feature vector based on the fifth module to obtain the twelfth feature vector; wherein the fifth module is used to instruct the full processing of the tenth feature vector, and the fourth module is used to instruct the processing of the information of the tenth feature vector under the second index. The error loss acquisition module is used to perform error processing based on the eleventh feature vector and the twelfth feature vector to obtain the error loss of the second strategy; The vector partitioning module is used to obtain a first strategy for the first machine learning model based on the error loss of all second strategies, and to partition the third feature vector based on the vector partitioning information corresponding to the first strategy.
[0194] In another scenario, optionally, the first input information includes one or more of the following: first audio, first video, first image, or first text, and the first content includes a second video, the video content of which is related to the first input information.
[0195] The beneficial effects of the above-mentioned device are as follows: acquiring first input information; analyzing and processing the first input information based on a first machine learning model, wherein the first machine learning model includes a first module and a second module. The first module can instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators, and the second module can instruct the processing of the second feature vector. The data volume of the second feature vector is less than that of the first feature vector. Therefore, a limited number of feature vectors can be processed based on the first machine learning model, reducing the amount of data processing. This solves the problem of computation and memory resource bottlenecks caused by redundancy of some features when uniformly processing feature vectors, and achieves the technical effect of ensuring the quality of output content while reducing the amount of data processing and memory usage.
[0196] The content generation apparatus provided herein can execute any of the content generation methods provided herein, and has the corresponding functional modules and beneficial effects of executing the methods.
[0197] It is worth noting that the various units and modules included in the above-mentioned device are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this document.
[0198] The following is for reference. Figure 10This document illustrates a schematic diagram of an electronic device (e.g., a terminal device or server) 900 suitable for implementing the above-described methods. The terminal device referred to herein may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0199] like Figure 10 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0200] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0201] In particular, according to embodiments of this document, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the technical solutions of this document include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of the embodiments of this document.
[0202] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0203] The electronic device provided in this embodiment and the content generation method provided in the above technical solutions belong to the same inventive concept. Technical details not described in detail in this document can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0204] This article provides a computer storage medium on which a computer program is stored, which, when executed by a processor, implements the content generation method provided in the above embodiments.
[0205] It should be noted that the computer-readable medium mentioned above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, also known as flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0206] Based on one or more scenarios described herein, a content generation method is provided, including: Obtain the first input information; The first input information is analyzed and processed based on a first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators. The second module is used to instruct the processing of the second feature vector. The data volume of the second feature vector is less than the data volume of the first feature vector. The first content is displayed, which reflects the result of the first machine learning model based on the second feature vector.
[0207] Based on one or more scenarios described herein, Example 2 provides a content generation method, further comprising: optionally, the first machine learning model includes a third module, wherein the analysis and processing of the first input information based on the first machine learning model includes: Based on the third module, a first feature vector of the first input information is obtained, and the first feature vector is subjected to dimensionality reduction processing to obtain a third feature vector. The first feature vector reflects the vector after the first input information is processed based on the weight matrix. Based on the first module, the fifth feature vector under the multiple first indicators is obtained from the fourth feature vector. The fourth feature vector is used to indicate the vector after the third feature vector is divided based on the obtained vector division information. The fifth feature vector is used to indicate the second module to obtain the second feature vector. Based on the second module, the second feature vector is obtained from the first feature vector based on the fifth feature vector, and the first content is output based on the second feature vector.
[0208] Based on one or more scenarios described herein, Example 3 provides a content generation method, further comprising: optionally, the first module includes a first unit, a feedforward neural network unit, and a second unit; the step of obtaining a fifth feature vector under the plurality of first indicators based on the first module includes: Based on the first unit, the first result of the third feature vector under multiple first indicators is obtained, and the first result is concatenated to obtain the second result of the third feature vector, wherein the first indicator includes at least one of the mean indicator, the minimum indicator and the maximum indicator. The second result is processed based on the feedforward neural network unit to obtain the third result; Based on the second unit, attention processing is performed on the third result to obtain the fifth feature vector, which is used to indicate the optional attributes of the fourth feature vector.
[0209] According to one or more scenarios described herein, Example 4 provides a content generation method, further comprising: optionally, the second module includes a third unit and a feature fusion unit, wherein obtaining the second feature vector from the first feature vector based on the fifth feature vector, and outputting the first content based on the second feature vector, includes: Based on the third unit and the fifth feature vector, a fourth feature vector is obtained for participating in the generation of the first content, and a second feature vector is selected from the first feature vector based on the fourth feature vector. The second feature vector is used to characterize the mask information for generating the first content. Based on the attention processing of the second feature vector by the feature fusion unit, and the processing of the obtained third result, the first content is obtained.
[0210] Based on one or more scenarios described herein, Example 5 provides a method for content generation, further comprising: optionally, the first machine learning model is trained in the following manner: Receive the second input information; The second input information is processed based on the second machine learning model to obtain the sixth feature vector, the seventh feature vector, and the second output content; Based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector, adjust the model parameters in the second machine learning model; In response to detecting that the second machine learning model has reached the convergence condition, the fourth module in the second machine learning model is removed to obtain the first machine learning model; The third output content is used to reflect the expected output result of the second input information, the second output content is used to reflect the predicted output result of the second output information, the sixth feature vector is used to reflect the feature vector after processing by the first module, the seventh feature vector is used to reflect the feature vector after processing by the fourth module, the fourth module is used to obtain the seventh feature vector of the eighth feature vector on the second index, the eighth feature vector is used to indicate the vector after the second input information is output by the third module, and the second index is the maximum value index among multiple first indices.
[0211] Based on one or more scenarios described herein, Example 6 provides a content generation method, further comprising: optionally, the model parameters including a first parameter and a second parameter, wherein the first parameter indicates learnable parameters in the first module, and the second parameter indicates learnable parameters in the first machine learning model other than the first module. The step of adjusting the model parameters in the second machine learning model based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector includes: Based on the second output content and the third output content, the first loss is determined; Based on the sixth feature vector and the seventh feature vector, determine the second loss; The first parameter is adjusted based on the second loss, and the second parameter is adjusted based on the first loss.
[0212] According to one or more scenarios described herein, Example 7 provides a content generation method, which further includes: optionally, the method further includes obtaining vector partitioning information for partitioning feature vectors by means of: Obtain third input information; The third input information is processed based on the third machine learning model to obtain the ninth feature vector, and the ninth feature vector is divided based on the second strategy to obtain multiple tenth feature vectors. The second strategy is pre-obtained information used to indicate the division of the ninth feature vector. The eleventh feature vector is obtained by processing the tenth feature vector using the fourth module, and the twelfth feature vector is obtained by processing the tenth feature vector using the fifth module; wherein the fifth module is used to instruct the full processing of the tenth feature vector, and the fourth module is used to instruct the processing of the information of the tenth feature vector under the second index. Error processing is performed based on the eleventh and twelfth feature vectors to obtain the error loss of the second strategy; Based on the error loss of all second strategies, a first strategy is obtained for use in the first machine learning model to partition the third feature vector based on the vector partitioning information corresponding to the first strategy.
[0213] According to one or more scenarios described herein, Example 8 provides a method for content generation, which further includes: optionally, the first input information includes one or more of a first audio, a first video, a first image, or a first text, and the first output information includes a second video, the video content of the second video being related to the first input information.
[0214] According to one or more scenarios described herein, Example 9 provides a content generation apparatus, comprising: The information receiving module is used to acquire the first input information; The information processing module is used to analyze and process the first input information based on the first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the second feature vector to be obtained from the first feature vector according to multiple first indicators. The second module is used to instruct the processing of the second feature vector. The data volume of the second feature vector is less than the data volume of the first feature vector. The content display module is used to display the first content, which reflects the result of the first machine learning model based on the second feature vector.
[0215] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0216] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0217] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: Obtain the first input information; The first input information is analyzed and processed based on a first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators. The second module is used to instruct the processing of the second feature vector. The data volume of the second feature vector is less than the data volume of the first feature vector. The first content is displayed, which reflects the result of the first machine learning model based on the second feature vector.
[0218] Computer program code for performing the operations described herein can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0219] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0220] The modules or units described herein can be implemented in software or hardware. The names of modules or units do not necessarily limit the functionality of the module or unit itself; for example, a second result acquisition unit can also be described as a "second result receiving unit".
[0221] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include at least one of the following: Field-Programmable Gate Array (FPGA), Application-Specific Integrated Circuit (ASIC), Application-Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.
[0222] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0223] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.
[0224] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this document. Certain features described in the context of individual implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0225] Although the subject matter has been described using a programming language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims.
Claims
1. A content generation method, comprising: Obtain the first input information; The first input information is analyzed and processed based on a first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the acquisition of a second feature vector from a first feature vector based on multiple first indicators. The second module is used to instruct the processing of the second feature vector. The data volume of the second feature vector is less than the data volume of the first feature vector. The first content is displayed, which reflects the result of the first machine learning model based on the second feature vector.
2. The content generation method according to claim 1, wherein the first machine learning model includes a third module, and the analysis and processing of the first input information based on the first machine learning model includes: Based on the third module, a first feature vector of the first input information is obtained, and the first feature vector is subjected to dimensionality reduction processing to obtain a third feature vector. The first feature vector reflects the vector after the first input information is processed based on the weight matrix. Based on the first module, the fifth feature vector under the multiple first indicators is obtained from the fourth feature vector. The fourth feature vector is used to indicate the vector after the third feature vector is divided based on the obtained vector division information. The fifth feature vector is used to indicate the second module to obtain the second feature vector. Based on the second module, the second feature vector is obtained from the first feature vector based on the fifth feature vector, and the first content is output based on the second feature vector.
3. The content generation method according to claim 2, wherein the first module includes a first unit, a feedforward neural network unit, and a second unit, and the step of obtaining the fifth feature vector under the plurality of first indicators based on the first module includes: Based on the first unit, the first result of the third feature vector under multiple first indicators is obtained, and the first result is concatenated to obtain the second result of the third feature vector, wherein the first indicator includes at least one of the mean indicator, the minimum indicator and the maximum indicator. The second result is processed based on the feedforward neural network unit to obtain the third result; Based on the second unit, attention processing is performed on the third result to obtain the fifth feature vector, which is used to indicate the optional attributes of the fourth feature vector.
4. The content generation method according to claim 2, wherein the second module includes a third unit and a feature fusion unit, and the step of obtaining the second feature vector from the first feature vector based on the fifth feature vector and outputting the first content based on the second feature vector includes: Based on the third unit and the fifth feature vector, a fourth feature vector is obtained for participating in the generation of the first content, and a second feature vector is selected from the first feature vector based on the fourth feature vector. The second feature vector is used to characterize the mask information for generating the first content. Based on the attention processing of the second feature vector by the feature fusion unit, and the processing of the obtained third result, the first content is obtained.
5. The content generation method according to claim 1, wherein the first machine learning model is trained in the following manner: Receive the second input information; The second input information is processed based on the second machine learning model to obtain the sixth feature vector, the seventh feature vector, and the second output content; Based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector, adjust the model parameters in the second machine learning model; In response to detecting that the second machine learning model has reached the convergence condition, the fourth module in the second machine learning model is removed to obtain the first machine learning model; in, The third output content is used to reflect the expected output result of the second input information, the second output content is used to reflect the predicted output result of the second input information, the sixth feature vector is used to reflect the feature vector after processing by the first module, the seventh feature vector is used to reflect the feature vector after processing by the fourth module, the fourth module is used to obtain the seventh feature vector of the eighth feature vector on the second index, the eighth feature vector is used to indicate the vector after the second input information is output by the third module, and the second index is the maximum value index among multiple first indices.
6. The content generation method according to claim 5, wherein the model parameters include a first parameter and a second parameter, the first parameter indicating the learnable parameters in the first module, and the second parameter indicating the learnable parameters of the first machine learning model other than the first module. The step of adjusting the model parameters in the second machine learning model based on the second output content, the third output content, the sixth feature vector, and the seventh feature vector includes: Based on the second output content and the third output content, the first loss is determined; Based on the sixth feature vector and the seventh feature vector, determine the second loss; The first parameter is adjusted based on the second loss, and the second parameter is adjusted based on the first loss.
7. The content generation method according to claim 1, further comprising obtaining vector partitioning information for partitioning feature vectors through the following means: Obtain third input information; The third input information is processed based on the third machine learning model to obtain the ninth feature vector, and the ninth feature vector is divided based on the second strategy to obtain multiple tenth feature vectors. The second strategy is pre-obtained information used to indicate the division of the ninth feature vector. The eleventh feature vector is obtained by processing the tenth feature vector using the fourth module, and the twelfth feature vector is obtained by processing the tenth feature vector using the fifth module; wherein, The fifth module is used to instruct the full processing of the tenth feature vector, and the fourth module is used to instruct the processing of the information of the tenth feature vector under the second index; Error processing is performed based on the eleventh and twelfth feature vectors to obtain the error loss of the second strategy; Based on the error loss of all second strategies, a first strategy is obtained for use in the first machine learning model to partition the third feature vector based on the vector partitioning information corresponding to the first strategy.
8. The content generation method according to claim 1, wherein the first input information includes one or more of the following: first audio, first video, first image, or first text, and the first content includes a second video, wherein the video content of the second video is related to the first input information.
9. A content generation apparatus, comprising: The information receiving module is used to acquire the first input information; The information processing module is used to analyze and process the first input information based on the first machine learning model. The first machine learning model includes a first module and a second module. The first module is used to instruct the second feature vector to be obtained from the first feature vector according to multiple first indicators. The second module is used to instruct the processing of the second feature vector. The data volume of the second feature vector is less than the data volume of the first feature vector. The content display module is used to display the first content, which reflects the result of the first machine learning model based on the second feature vector.
10. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the content generation method as described in any one of claims 1-8.
11. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the content generation method as described in any one of claims 1-8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the content generation method as described in any one of claims 1-8.