A video generation method and related apparatus
By coordinating the processing between communication devices and generating video data in collaboration with AI models, the limitations of computing power and power consumption of communication equipment have been solved, resulting in reduced resource consumption and latency in video generation, improved user experience, and protection of privacy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing communication devices are limited by computing power and power consumption, making it difficult to achieve smooth video generation.
Through collaborative processing between different communication devices, video data is generated collaboratively using an AI model. This includes the first communication device sending instruction information to instruct network devices to generate video data, and the second or third communication device performing partial data processing, thereby reducing the resource consumption and latency of the first communication device.
It reduces resource consumption and latency in video generation, improves user experience, and protects privacy data through pre-configured rules to prevent privacy leaks.
Smart Images

Figure CN122120564A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a video generation method and related equipment. Background Technology
[0002] To address the vision of a future of intelligent and inclusive access, intelligence will further evolve at the wireless network architecture level. Taking artificial intelligence (AI) as an example, AI may be more deeply integrated with wireless networks, enabling inherent intelligence within the network, and potentially achieving intelligent terminals as well.
[0003] Currently, with the continuous development of AI technology, AI-based video generation technology has become a field of great interest. For example, text-to-video generation is a typical application of video generation technology. In the process of text-to-video generation, the video generation requirements are described using natural language, and the video generation device combines the text content corresponding to the input natural language with other conditional information to generate the video. For instance, in future communication networks, communication devices (such as terminals) may also have the need to generate various types of videos, including text-to-video, image-to-video, and video-to-video.
[0004] However, video generation demands significant computing and power consumption, and current communication devices are limited by their power consumption and computing power, making it difficult to achieve video generation. Therefore, providing smooth video stream generation services within communication networks is a pressing issue that needs to be addressed. Summary of the Invention
[0005] This application provides a video generation method and related apparatus for achieving video generation through collaborative processing between different communication devices.
[0006] The first aspect of this application provides a video generation method applied to a first communication device. For example, the first communication device may be a communication equipment (such as a terminal device), or it may be a component of the communication equipment (e.g., a processor, circuit, chip, or chip system responsible for communication functions, including but not limited to a modem chip, a baseband chip, a system-on-a-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip, etc.). Alternatively, the first communication device may also be a logic module or software capable of implementing all or part of the functions of the communication equipment. The following description uses a first communication device as an example.
[0007] In this method, a first communication device sends first information to instruct a network device to generate first video data; the first communication device receives first data, which is obtained by the network device processing the first information based on a first model; and the first communication device determines the first video data based on the first data.
[0008] Based on the above scheme, after the first communication device sends first information instructing the generation of first video data, the recipient of the first information can process the first data based on the first model and the first information to obtain first data, and then send the first data to the first communication device, enabling the first communication device to determine the first video data based on the first data. In this way, other communication devices can use the first model to process the video data generation instruction from the first communication device, obtain and send the first data, enabling the first communication device to generate first video data based on the first data generated by the models of other communication devices. This allows video generation to be achieved through collaborative processing between different communication devices.
[0009] Furthermore, during the aforementioned video data processing, the first communication device can offload some data processing to other communication devices (such as the second communication device described below) for processing. This reduces the resource consumption (including computing power and power) of the first communication device and also reduces the processing latency of video data generation, thereby improving the user experience.
[0010] In this application, the model may include an AI model, a neural network model, an AI neural network model, a machine learning model, an AI processing model, or other names defined by future networks.
[0011] Optionally, the first information is used to instruct the network device to generate the first video data, which can be understood as: the first information is used to instruct the first model in the network device to assist / help / participate in the generation of the first video data, and / or, the first information is used to instruct the network device to assist / help / participate in the generation of the first video data.
[0012] In one possible implementation of the first aspect, the first communication device determines (e.g., generates) the first video data based on the first data, comprising: the first communication device processing the first data based on a second model to obtain second data; and the first communication device determining the first video data based on the second data. For example, the process by which the first communication device determines the first video data based on the second data can be performed using a video decoder (as described below). Figure 2 This can be implemented using methods such as a decoder (or an AI model). For example, the second data can be used as input to a video decoder to obtain the first video data as output. Alternatively, the second data can be used as input to an AI model to obtain the first video data as output. The implementation methods of the video decoder and AI model are not limited.
[0013] Based on the above scheme, the first communication device processes the first data using the second model to obtain second data, and then determines the first video data based on the second data. In this way, the first communication device can call the second model to process the first data sent by other communication devices, and can perform parallel collaborative processing through different models called by different communication devices, providing more resources for video generation and reducing the latency of video generation.
[0014] In one possible implementation of the first aspect, the first data is part or all of the input data of the first AI module included in the first model, and the second data is obtained by processing the first data through the second AI module included in the second model; wherein the first AI module is associated with the second AI module.
[0015] Based on the above scheme, any model (e.g., the first model or the second model) can contain one or more AI models. In the above process, the first data sent by the second communication device can be the input data of the first AI module (or the first data can be the output data of the previous module of the first AI module). The module of the second model that processes the first data to obtain the second data can be the second AI module, and these two modules can be related modules. In this way, during the collaborative video generation process of the first and second models, these two models can perform data processing through related AI modules, which can improve data processing performance and improve the quality of subsequent video generation, thereby enhancing the user experience.
[0016] In one possible implementation of the first aspect, the method further includes: the first communication device transmitting the second data.
[0017] Based on the above scheme, after the first communication device processes the first data using the second model to obtain the second data, the first communication device can send the second data. For example, this method can be called synchronous parallel processing or interactive parallel processing. In this way, the second communication device can perform subsequent processing based on the second data, enabling parallel collaborative processing through different models invoked by different communication devices to improve the quality of the generated video and thus enhance the user experience.
[0018] Optionally, after receiving the second data, the second communication device can process the second data through one or more AI modules included in the first model, and the second communication device can send the processing result to the first communication device so that the first communication device can determine the first video data based on the processing result. In this way, the first and second communication devices can achieve parallel collaborative processing through two or more interaction processes to further improve the quality of the generated video.
[0019] For example, when the first communication device does not send (or is determined not to send) the second data, the AI network architecture of the first model is the same as or similar to the AI network architecture of the second model, and / or, the AI modules with the same index included in the first model and the second model have the same or similar functions. For example, the first AI module is a self-attention or multi-head attention module, and / or, the second AI modules are all cross-attention modules.
[0020] In one possible implementation of the first aspect, the method further includes: the first communication device receiving second information, or the first communication device sending second information, the second information being used to instruct the first AI module and / or the second AI module.
[0021] Based on the above scheme, the recipient of the second information can determine the AI module corresponding to the first data as the first AI module and / or the second AI module based on the indication information, and can subsequently use the corresponding AI module for data processing.
[0022] In one possible implementation of the first aspect, the first information is obtained by filtering based on pre-configured rules.
[0023] Based on the above scheme, the video request information of the first communication device for the first video data may indicate multiple video generation requirements. This video request information may include portions that the first communication device does not wish to send to other communication devices (for example, the first communication device determines, based on pre-configured information or user-inputted preference information, that the data it does not wish to send includes one or more types of private or sensitive data). In this process, the first communication device can filter the video request information based on pre-configured rules to obtain the first information, thereby achieving data privacy protection and preventing privacy leaks.
[0024] In one possible implementation of the first aspect, the pre-configured rules are determined based on user operation instructions.
[0025] Based on the above scheme, the pre-configured rules used for filtering can be determined based on user operation instructions, enabling the first communication device to protect data privacy based on user operation instructions, thereby improving user experience.
[0026] In one possible implementation of the first aspect, the first information is obtained by filtering the first video demand information based on the pre-configured rules; wherein the first video demand information is obtained by the third model processing the second video demand information, and the second video demand information is described in natural language.
[0027] Based on the above scheme, the second video demand information used to indicate the demand for the first video data can be described in natural language, and the third model can process the second video demand information to obtain the first video demand information. Subsequently, the first communication device can filter the first video demand information based on the pre-configured rules to obtain the first information, so as to achieve data privacy protection.
[0028] In one possible implementation of the first aspect, the first video demand information is used for processing the second model; or, the first video demand information and the first information are used to determine third video demand information, which is then used for processing the second model.
[0029] Based on the above scheme, during the data processing process of the first communication device based on the second model, the input of the second model may include first video requirement information or third video requirement information, so that the data processed by the second model can meet the video requirements, improve the video quality corresponding to the subsequently generated first video data, and enhance the user experience.
[0030] In one possible implementation of the first aspect, the method further includes: the first communication device sending third information for requesting processing resources for the first video data.
[0031] Based on the above scheme, the first communication device can also request resources through third information, so that the recipient of the third information can reserve resources for the processing of the first video data to ensure the generation of the first video data.
[0032] A second aspect of this application provides a video generation method applied to a second communication device. For example, the second communication device may be a communication equipment (such as a network device), or it may be a component of the communication equipment (e.g., a processor, circuit, chip, or chip system responsible for communication functions, including but not limited to a modem chip, a baseband chip, a system-on-a-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip, etc.). Alternatively, the second communication device may also be a logic module or software capable of implementing all or part of the functions of the communication equipment. The following description uses a second communication device as an example.
[0033] In this method, a second communication device receives first information, which instructs a network device to generate first video data; the second communication device sends first data, which is obtained by the network device processing the first information based on the first model; wherein, the first data is used to determine the first video data.
[0034] Based on the above scheme, after receiving the first information indicating the generation of first video data, the second communication device can process the first data based on the first model and the first information to obtain first data, and then send the first data to the first communication device, enabling the first communication device to determine the first video data based on the first data. In this way, the second communication device can use the first model to process the video data generation instruction from the first communication device, obtain and send the first data, allowing the first communication device to generate the first video data based on the first data generated by the model of the second communication device. This enables video generation through collaborative processing between different communication devices.
[0035] Furthermore, during the aforementioned video data processing, the first communication device can offload some data processing to the second communication device, which can reduce the resource consumption (including computing power, power, etc.) of the first communication device and reduce the processing latency of video data generation, thereby improving the user experience.
[0036] In one possible implementation of the second aspect, the first data is used to obtain second data through processing of the second model, and the second data is used to determine the first video data.
[0037] Based on the above scheme, after the first communication device processes the first data using the second model to obtain the second data, the first communication device can determine the first video data based on the second data. In this way, the first communication device can call the second model to process the first data sent by other communication devices, and can perform parallel collaborative processing through different models called by different communication devices, providing more resources for video generation and reducing the latency of video generation.
[0038] In one possible implementation of the second aspect, the first data is part or all of the input data of the first AI module included in the first model, and the second data is obtained by processing the first data through the second AI module included in the second model; wherein the first AI module is associated with the second AI module.
[0039] Based on the above scheme, a model can contain one or more AI models. In the process described above, the first data sent by the second communication device can be the input data of the first AI module (or the first data can be the output data of the previous module of the first AI module). The module that processes the first data to obtain the second data can be the second AI module, and these two modules can be related modules. In this way, during the collaborative video generation process of the first and second models, these two models can perform data processing through related AI modules, which can improve data processing performance and enhance the quality of subsequent video generation, thereby improving the user experience.
[0040] In one possible implementation of the second aspect, the method further includes: the second communication device receiving the second data.
[0041] Based on the above scheme, after the first communication device processes the first data using the second model to obtain the second data, the first communication device can send the second data. For example, this method can be called synchronous parallel processing or interactive parallel processing. In this way, the second communication device can perform subsequent processing based on the second data, enabling parallel collaborative processing through different models invoked by different communication devices to improve the quality of the generated video and thus enhance the user experience.
[0042] In one possible implementation of the second aspect, the method further includes: the second communication device receiving or sending second information, the second information being used to instruct the first AI module and / or the second AI module.
[0043] Based on the above scheme, the recipient of the second information can determine the AI module corresponding to the first data as the first AI module and / or the second AI module based on the indication information, and can subsequently use the corresponding AI module for data processing.
[0044] In one possible implementation of the second aspect, the first information is obtained by filtering based on pre-configured rules.
[0045] Based on the above scheme, the video request information of the first communication device for the first video data may indicate multiple video generation requirements. This video request information may include portions that the first communication device does not wish to send to other communication devices (such as one or more pieces of privacy data or sensitive data). In this process, the first communication device can filter the video request information based on pre-configured rules to obtain the first information, thereby achieving data privacy protection and preventing privacy leaks.
[0046] In one possible implementation of the second aspect, the pre-configured rules are determined based on user operation instructions.
[0047] Based on the above scheme, the pre-configured rules used for filtering can be determined based on user operation instructions, enabling the first communication device to protect data privacy based on user operation instructions, thereby improving user experience.
[0048] In one possible implementation of the second aspect, the first information is obtained by filtering the first video demand information based on the pre-configured rules; wherein the first video demand information is obtained by the third model processing the second video demand information, and the second video demand information is described in natural language.
[0049] Based on the above scheme, the second video demand information used to indicate the demand for the first video data can be described in natural language, and the third model can process the second video demand information to obtain the first video demand information. Subsequently, the first communication device can filter the first video demand information based on the pre-configured rules to obtain the first information, so as to achieve data privacy protection.
[0050] In one possible implementation of the second aspect, the method further includes: the second communication device receiving third information for requesting processing resources for the first video data.
[0051] Based on the above scheme, the first communication device can also request resources through third information, so that the second communication device can reserve resources for the processing of the first video data to ensure the generation of the first video data.
[0052] A third aspect of this application provides a video generation method applied to a third communication device. For example, the third communication device may be a communication equipment (such as a terminal device), or it may be a component of a communication equipment (such as a processor, circuit, chip, or chip system responsible for communication functions), or it may be a logic module or software capable of implementing all or part of the functions of the communication equipment. The following description uses a third communication device as an example.
[0053] In this method, a third communication device acquires third video demand information, which is used to indicate the generation of first video data; the third communication device processes the third video demand information based on a second model to obtain second data; the third communication device sends the second data; wherein the second data is used to determine the first video data.
[0054] Based on the above scheme, after acquiring third video demand information indicating the generation of first video data, the third communication device can process the third video demand information based on the second model to obtain and send second data, enabling the recipient of the second data to determine the first video data based on the second data. In this way, the third communication device can generate video data based on video demand to achieve video generation.
[0055] In one possible implementation of the third aspect, the method further includes: the third communication device receiving first data, the first data being obtained by processing the first information based on the first model, the first information being used to indicate the generation of the first video data; the third communication device processing the third video demand information based on the second model to obtain second data, including: the third communication device processing the third video demand information and the first data based on the second model to obtain the second data.
[0056] Based on the above scheme, the first data received by the third communication device can be obtained by processing the first information using the first model. Furthermore, the third communication device can process the third video demand information and the first data using the second model to obtain the second data. In this way, parallel collaborative processing using different models invoked by different communication devices provides more resources for video generation, thereby reducing video generation latency.
[0057] Furthermore, during the aforementioned video data processing, the third communication device can offload some data processing to other communication devices, thereby reducing the resource consumption (including computing power and power) of the third communication device and reducing the processing latency of video data generation, thus improving the user experience.
[0058] In one possible implementation of the third aspect, the method further includes: the third communication device receiving the third video demand information, which is determined based on the first video demand information and the first information. Optionally, the first information is used to instruct the generation of first video data (e.g., instructing a first model to assist / help / participate in the generation of the first video data).
[0059] Based on the above scheme, the third communication device can also receive third video demand information to obtain the third video demand information, and then generate corresponding data based on the third video demand information to facilitate the subsequent generation of video data that meets the requirements.
[0060] In one possible implementation of the third aspect, the first video demand information is obtained by the third model based on the second video demand information, which is described in natural language.
[0061] Based on the above scheme, the second video demand information used to indicate the demand for the first video data can be described in natural language, and the third model can process the second video demand information to obtain the first video demand information. Subsequently, the third communication device can filter the first video demand information based on the pre-configured rules to obtain the first information, so as to achieve data privacy protection.
[0062] It should be noted that the implementation process of the third aspect can also refer to the first aspect and related descriptions above. Optionally, the third communication device can be an internal module (such as a chip module, software module, hardware module, etc.) included in the first communication device that is capable of running the second model, or the third communication device can be an external module, device, or apparatus connected to the first communication device that is capable of running the second model.
[0063] A fourth aspect of this application provides a communication device that performs the functions described in the first aspect. For example, the communication device includes modules, units, or means corresponding to the operations involved in the first aspect. These modules, units, or means can be implemented in software, hardware, or a combination of both. For instance, the device includes a processing unit and a transceiver unit. The transceiver unit is used to send first information indicating the generation of first video data. The transceiver unit is also used to receive first data obtained by processing the first information using a first model. The processing unit is used to determine the first video data based on the first data.
[0064] In the fourth aspect of this application, the constituent modules of the communication device can also be used to perform the steps executed in various possible implementations of the first aspect and achieve the corresponding technical effects. For details, please refer to the first aspect, which will not be repeated here.
[0065] A fifth aspect of this application provides a communication device that performs the functions described in the second aspect above. For example, the communication device includes modules, units, or means corresponding to the operations involved in the second aspect. These modules, units, or means can be implemented in software, hardware, or a combination of both. For instance, the device includes a processing unit and a transceiver unit. The transceiver unit receives first information, which instructs a network device to generate first video data. The processing unit determines the first data. The transceiver unit also transmits the first data, which is obtained by the network device processing the first information based on a first model. The first data is used to determine the first video data.
[0066] In the fifth aspect of this application, the constituent modules of the communication device can also be used to perform the steps executed in various possible implementations of the second aspect and achieve the corresponding technical effects. For details, please refer to the second aspect, which will not be repeated here.
[0067] A sixth aspect of this application provides a communication device that performs the functions described in the third aspect above. For example, the communication device includes modules, units, or means corresponding to the operations involved in the third aspect. These modules, units, or means can be implemented in software, hardware, or a combination of both. For instance, the device includes a processing unit and a transceiver unit. The processing unit acquires third video demand information, which instructs the generation of first video data. The processing unit further processes the third video demand information based on a second model to obtain second data. The transceiver unit transmits the second data, wherein the second data is used to determine the first video data.
[0068] In the sixth aspect of this application, the constituent modules of the communication device can also be used to perform the steps executed in various possible implementations of the third aspect and achieve the corresponding technical effects. For details, please refer to the third aspect, which will not be repeated here.
[0069] A seventh aspect of this application provides a communication device including at least one processor for executing computer programs or instructions to enable the device to implement any one of the first to third aspects and any possible implementation thereof.
[0070] Optionally, the at least one processor is coupled to a memory for storing computer programs or instructions.
[0071] Optionally, the communication device includes the memory. Optionally, the memory is integrated with at least one processor.
[0072] The eighth aspect of this application provides a communication device including at least one logic circuit and an input / output interface; the logic circuit is used to perform the method as described in any one of the possible implementations of the first to third aspects.
[0073] In one possible implementation, the communication device is a chip or chip system.
[0074] A ninth aspect of this application provides a communication system including the first and second communication devices described above. Optionally, the communication system further includes the third communication device described above.
[0075] The tenth aspect of this application provides a computer-readable storage medium for storing one or more computer-executable instructions, which, when executed by a processor, perform a method as described in any possible implementation of any of the first to third aspects above.
[0076] The eleventh aspect of this application provides a computer program product (or computer program) in which, when the computer program in the computer program product is executed by the processor, the processor executes any possible implementation of any of the first to third aspects of the method described above.
[0077] The twelfth aspect of this application provides a chip or chip system including at least one processor for supporting a communication device in implementing any possible implementation of any of the first to third aspects described above. For example, the chip may be a baseband chip, a modem chip, a system-on-a-chip (SoC) chip containing a modem core, a system-in-package (SIP) chip, or a communication module, etc.
[0078] In one possible design, the chip or chip system may further include a memory for storing program instructions and data necessary for the communication device. The chip system may consist of chips or may include chips and other discrete devices. Optionally, the chip system may also include interface circuitry that provides program instructions and / or data to the at least one processor.
[0079] The technical effects of any of the design methods in aspects four through twelfth can be found in the technical effects of the different design methods in aspects one through three above, and will not be repeated here. Attached Figure Description
[0080] Figure 1 A schematic diagram of the communication system provided in this application;
[0081] Figure 2 This is a schematic diagram of the AI processing involved in this application;
[0082] Figures 3a to 3c An interactive schematic diagram of the video generation method provided in this application;
[0083] Figures 4a to 4b This is a schematic diagram of the model processing procedure involved in this application;
[0084] Figure 5 This is a schematic diagram of the model processing procedure involved in this application;
[0085] Figure 6 An interactive schematic diagram of the video generation method provided in this application;
[0086] Figures 7 to 11 A schematic diagram of the communication device provided in this application. Detailed Implementation
[0087] First, some terms used in the embodiments of this application will be explained to facilitate understanding by those skilled in the art.
[0088] (1) Terminal device: can be a wireless terminal device that can receive network device scheduling and instruction information. The wireless terminal device can be a device that provides voice and / or data connectivity to the user, or a handheld device with wireless connection function, or other processing device connected to a wireless modem.
[0089] Terminal devices can communicate with one or more core networks or the Internet via a radio access network (RAN). Terminal devices can be mobile terminal devices, such as mobile phones (or "cellular" phones), computers, and data cards. For example, they can be portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile devices that exchange voice and / or data with the RAN. Examples include personal communication service (PCS) phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), tablets, and computers with wireless transceiver capabilities. Wireless terminal equipment can also be referred to as a system, subscriber unit, subscriber station, mobile station, mobile station (MS), remote station, access point (AP), remote terminal, access terminal, user terminal, user agent, subscriber station (SS), customer premises equipment (CPE), terminal, user equipment (UE), mobile terminal (MT), etc.
[0090] By way of example and not limitation, in this embodiment, the terminal device can also be a wearable device. Wearable devices, also known as wearable smart devices or smart wearable devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets, smart helmets, and smart jewelry for vital sign monitoring.
[0091] Terminals can also be drones, robots, devices in device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical care, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc.
[0092] Furthermore, terminal devices can also be terminal devices in future communication systems evolving from fifth-generation (5G) communication systems, or terminal devices in future public land mobile networks (PLMNs). For example, future communication networks can further expand the form and function of 5G communication terminals; future communication system terminals include, but are not limited to, vehicles, cellular network terminals (integrating satellite terminal functions), drones, and Internet of Things (IoT) devices.
[0093] In this embodiment, the terminal device can also obtain AI services provided by the network device. Optionally, the terminal device can also have AI processing capabilities.
[0094] (2) Network equipment: This can be equipment in a wireless network. For example, network equipment can be a RAN node (or device) that connects terminal devices to the wireless network, and can also be called a base station. Currently, some examples of RAN equipment include: base station, evolved NodeB (eNodeB), gNB (gNodeB) in 5G communication systems, transmission reception point (TRP), evolved Node B (eNB), radio network controller (RNC), Node B (NB), home base station (e.g., home-evolved Node B, or home Node B, HNB), base band unit (BBU), or wireless fidelity (Wi-Fi) access point (AP), etc. In addition, in a network structure, network equipment can include centralized unit (CU) nodes, distributed unit (DU) nodes, or RAN equipment including CU nodes and DU nodes.
[0095] Optionally, RAN nodes can also be macro base stations, micro base stations, indoor stations, relay nodes, donor nodes, or radio controllers in cloud radio access network (CRAN) scenarios. RAN nodes can also be servers, wearable devices, vehicles, or in-vehicle equipment. For example, the access network equipment in V2X technology can be a roadside unit (RSU).
[0096] In another possible scenario, multiple RAN nodes collaborate to assist the terminal in achieving wireless access, with different RAN nodes each implementing a portion of the base station's functions. For example, RAN nodes can be central units (CUs), distributed units (DUs), CU-control plane (CPs), CU-user plane (UPs), or radio units (RUs), etc. CUs and DUs can be set up separately or included in the same network element, such as a baseband unit (BBU). RUs can be included in radio frequency equipment or radio frequency units, such as remote radio units (RRUs), active antenna units (AAUs), or remote radio heads (RRHs).
[0097] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an open radio access network (O-RAN or ORAN) system, CU can also be called O-CU (open CU), DU can also be called O-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CU-UP, and RU can also be called O-RU. For ease of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples. Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software modules and hardware modules.
[0098] Network devices can be other devices that provide wireless communication functions for terminal devices. The embodiments of this application do not limit the specific technology or form of the network device. For ease of description, the embodiments of this application are not limited.
[0099] In this embodiment of the application, the network device may also have network nodes with AI capabilities, which can provide AI services to terminals or other network devices. For example, it may be an AI node, computing power node, RAN node with AI capabilities, core network element with AI capabilities, etc. on the network side (access network or core network).
[0100] In this application embodiment, the device for implementing the function of the network device can be the network device itself, or it can be a device capable of supporting the network device in implementing that function, such as a chip system, which can be installed in the network device. In the technical solutions provided in this application embodiment, the example of a network device being used to implement the function of the network device is used to describe the technical solutions provided in this application embodiment.
[0101] (3) Configuration and Pre-configuration: In this application, both configuration and pre-configuration are used. Configuration refers to the network device / server sending configuration information or parameter values to the terminal via messages or signaling, so that the terminal can determine communication parameters or resources for transmission based on these values or information. Pre-configuration is similar to configuration; it can be parameter information or parameter values pre-negotiated between the network device / server and the terminal device, parameter information or parameter values specified by standard protocols for use by the base station / network device or terminal device, or parameter information or parameter values pre-stored in the base station / server or terminal device. This application does not limit this.
[0102] Furthermore, these values and parameters can be changed or updated.
[0103] (4) The terms "system" and "network" in the embodiments of this application can be used interchangeably. "Multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, "at least one of A, B and C" includes A, B, C, AB, AC, BC or ABC. And, unless otherwise specified, the ordinal numbers such as "first" and "second" mentioned in the embodiments of this application are used to distinguish multiple objects and are not used to limit the order, sequence, priority or importance of multiple objects.
[0104] (5) In the embodiments of this application, "send" and "receive" indicate the direction of signal transmission. For example, "send information to XX" can be understood as the destination of the information being XX, which may include sending directly through the air interface or sending indirectly through the air interface by other units or modules. "Receive information from YY" can be understood as the source of the information being YY, which may include receiving directly from YY through the air interface or receiving indirectly from YY through the air interface by other units or modules. "Send" can also be understood as the "output" of the chip interface, and "receive" can also be understood as the "input" of the chip interface.
[0105] In other words, sending and receiving can occur between devices, such as between network devices and terminal devices, or within a device, such as between components, modules, chips, software modules, or hardware modules within the device via buses, wiring, or interfaces.
[0106] It is understandable that information may undergo necessary processing, such as encoding and modulation, between the source and destination, but the destination can understand the valid information from the source. Similar statements in this application can be interpreted in a similar way and will not be elaborated further.
[0107] (6) In the embodiments of this application, "instruction" may include direct instruction and indirect instruction, as well as explicit instruction and implicit instruction. The information indicated by a certain piece of information (as described below, the instruction information) is called the information to be instructed. In the specific implementation process, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can only indicate a part of the information to be instructed, while the other parts of the information to be instructed are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement order of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed; for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.
[0108] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, and the various methods / designs / implementations within each embodiment, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments and between the various methods / designs / implementations within each embodiment are consistent and can be mutually referenced. The technical features in different embodiments and the various methods / designs / implementations within each embodiment can be combined to form new embodiments, methods, or implementations based on their inherent logical relationships. The following descriptions of the embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0109] This application can be applied to long-term evolution (LTE) systems, new radio (NR) systems, or communication systems evolving beyond 5G. These communication systems include at least one network device and / or at least one terminal device.
[0110] Please see Figure 1 This is a schematic diagram of the architecture of the communication system 1000 used in an embodiment of this application. Figure 1 As shown, the communication system includes RAN 100 and core network 200. Optionally, the communication system 1000 may also include Internet 300. RAN 100 includes at least one RAN node (e.g., Figure 1 110a and 110b, collectively referred to as 110, may also include at least one terminal (such as...). Figure 1 RAN 100, denoted as RAN 120a-120j, is collectively referred to as RAN 120. RAN 100 may also include other RAN nodes, such as wireless relay equipment and / or wireless backhaul equipment. Figure 1 (Not shown in the image). Terminal 120 connects wirelessly to RAN node 110, and RAN node 110 connects wirelessly or via a wired connection to core network 200. The core network equipment in core network 200 and RAN node 110 in RAN 100 can be independent physical devices, or they can be the same physical device integrating the logical functions of core network equipment and RAN nodes. Terminals can connect to each other, and RAN nodes can connect to each other, via wired or wireless connections.
[0111] RAN 100 can be an evolved universal terrestrial radio access (E-UTRA) system, an NR system, or a future radio access system as defined in the 3rd generation partnership project (3GPP). RAN 100 can also include two or more of the above-mentioned different radio access systems. RAN 100 can also be an open RAN (O-RAN).
[0112] For ease of description, the following text uses a base station as an example of a RAN node.
[0113] Base stations and terminals can be fixed or mobile. They can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; they can also be deployed on water; and they can be deployed on aircraft, balloons, and satellites. The embodiments of this application do not limit the application scenarios of the base stations and terminals.
[0114] The roles of base stations and terminals can be relative, for example, Figure 1The helicopter or drone 120i can be configured as a mobile base station. For terminals 120j accessing the wireless access network 100 via 120i, terminal 120i is a base station; however, for base station 110a, 120i is a terminal, meaning that 110a and 120i communicate via a wireless air interface protocol. Of course, 110a and 120i can also communicate via a base station-to-base station interface protocol; in this case, 120i is also a base station relative to 110a. Therefore, both base stations and terminals can be collectively referred to as communication devices. Figure 1 The 110a and 110b in the text can be referred to as communication devices with base station functions. Figure 1 The 120a-120j in the text can be referred to as communication devices with terminal functions.
[0115] In the embodiments of this application, the functions of the base station can be executed by modules (such as chips) within the base station, or by a control subsystem that includes base station functions. This control subsystem, including base station functions, can be a control center in the aforementioned application scenarios such as smart grids, industrial control, intelligent transportation, and smart cities. Similarly, the functions of the terminal can be executed by modules (such as chips or modems) within the terminal, or by a device that includes terminal functions.
[0116] The technical solution provided in this application can be applied to wireless communication systems (e.g.) Figure 1 The system shown in this application, for example, the communication system provided in this application, can incorporate AI network elements to implement some or all AI-related operations. AI network elements can also be called AI nodes, AI devices, AI entities, AI modules, AI models, or AI units, etc. The AI network element can be built into a network element within the communication system. For example, an AI network element can be an AI module built into: terminal equipment, access network equipment, core network equipment, cloud server, or operation, administration and maintenance (OAM) management system, used to implement AI-related functions. The OAM can be the management system for core network equipment and / or the management system for access network equipment. Alternatively, the AI network element can also be a network element independently set up in the communication system. Optionally, the terminal or its built-in chip can also include an AI entity to implement AI-related functions.
[0117] The following is a brief introduction to the artificial intelligence (AI) that may be involved in this application.
[0118] AI (Artificial Intelligence) enables machines to possess human-like intelligence, such as allowing them to use computer hardware and software to simulate certain intelligent human behaviors. To achieve artificial intelligence, machine learning methods can be employed. In machine learning, machines learn (or train) models using training data. These models represent the mapping between inputs and outputs. The learned model can be used for reasoning (or prediction), that is, it can be used to predict the output corresponding to a given input. This output can also be called the reasoning result (or prediction result).
[0119] Machine learning can include supervised learning, unsupervised learning, and reinforcement learning.
[0120] Neural networks (NNs) are a specific model in machine learning techniques. According to the general approximation theorem, neural networks can theoretically approximate any continuous function, thus enabling them to learn arbitrary mappings. Traditional communication systems rely on extensive expert knowledge to design communication modules, while deep learning communication systems based on neural networks can automatically discover hidden pattern structures from large datasets, establish mapping relationships between data, and achieve performance superior to traditional modeling methods.
[0121] The idea behind neural networks comes from the neuronal structure of the brain. For example, each neuron performs a weighted summation of its input values and outputs the result through an activation function.
[0122] Neural networks typically consist of multiple layers, each containing one or more neurons. Increasing the depth and / or width of a neural network enhances its expressive power, providing more robust information extraction and abstract modeling capabilities for complex systems. The depth of a neural network refers to the number of layers, while the number of neurons in each layer is called its width. In one implementation, a neural network includes an input layer and an output layer. The input layer processes the received input information through neurons and passes the results to the output layer, which then provides the network's output. In another implementation, a neural network includes an input layer, hidden layers, and an output layer. The input layer processes the received input information through neurons and passes the results to the hidden layers. The hidden layers perform calculations on the received results and pass them to the output layer or the next adjacent hidden layer, ultimately providing the network's output. A neural network may contain one hidden layer or multiple sequentially connected hidden layers; there is no limitation on this.
[0123] Models (such as the first model, second model, third model, etc. mentioned below) can also be called AI models, rules, or other names. An AI model can be considered a specific method for implementing AI functions. An AI model represents the mapping relationship or function between the model's input and output. AI functions can include one or more of the following: data collection, model training (or model learning), model information dissemination, model inference (or model reasoning, inference, or prediction, etc.), model monitoring or model validation, or inference result publication, etc. AI functions can also be called AI (related) operations or AI-related functions.
[0124] The technical solution provided in this application can be applied to wireless communication systems (e.g.) Figure 1 The system shown illustrates this. In wireless communication systems, communication nodes typically possess both signal transmission and reception capabilities and computational capabilities. Taking a network device with computational capabilities as an example, its computational power primarily provides support for signal transmission and reception (e.g., processing signals for transmission and reception) to enable communication tasks between the network device and other communication nodes. In a communication network, in addition to providing computational support for the aforementioned communication tasks, communication nodes may also possess surplus computational power, which can be applied to intelligent processing. To address the vision of future intelligent and inclusive accessibility, intelligentization may further evolve at the wireless network architecture level. Taking artificial intelligence (AI) as an example, AI may be further deeply integrated with wireless networks to achieve network-native intelligence, and may also include terminal intelligence.
[0125] Currently, with the continuous development of AI technology, AI-based video generation technology has become a field of great interest. Video generation technologies include, but are not limited to, text-to-video, image-to-video, and video-to-video conversion. For example, text-to-video conversion is a typical application of video generation technology. In this process, the video generation requirements are described using natural language, and the video generation device combines the input natural language text content with other conditional information to convert it into video.
[0126] like Figure 2 The image shown is a schematic diagram illustrating one implementation of text-based video. Figure 2 In this model, the video generator is the core processing module. For example, the video generator can incorporate one or more of the following models: DiT model, diffusion model, or transformer model. The input to the video generator can include the two types of data illustrated in the diagram:
[0127] One type of data is data where the encoder can encode video data① from the pixel space into the latent space, obtaining latent space data A (optionally, data A can be obtained after adding noise). Data A can then be divided into blocks.
[0128] (patchify) can generate multiple patches.
[0129] Another type of data involves processing human instructions (e.g., intent recognition processing) using large models (such as GPT / LLM). These human instructions can be described in natural language, and can be referred to as the conditions for video generation.
[0130] Subsequently, the video generator can process the two types of input data mentioned above to obtain data B in the latent space. Then, the decoder can decode data B to obtain video data ② that meets the intent, thus completing the video generation process.
[0131] Optionally, video data ① can be referred to as the initial video, and video data ② can be referred to as the target video. For example, video data ① can be referred to as the initial video of video data ②, and video data ② can be referred to as the target video of video data ①.
[0132] In the above process, video generation requirements can be reflected through two types of input data to the video generator. For example, the patch mentioned above can indicate the video generation requirements corresponding to the initial video. Similarly, the intent mentioned above can indicate the video generation requirements from humans or users. Generally, to improve video processing effects, video generators contain a large number of model parameters; for example, a video generator may contain approximately 30 billion (B represents billions). Taking a DiT model in a video generator as an example, the DiT model can contain multiple DiT blocks. During video generation, different DiT blocks can undergo multiple data processing steps (e.g., 20-50 iterations of cyclic denoising) through looping, serial processing, or parallel processing. Each DiT block can include an AI module with learning capabilities and / or attention mechanisms, such as a self-attention module or a cross-attention module. Optionally, the self-attention module can include a multi-head self-attention module. Optionally, each DiT Block may also include other modules such as layer normalization, feedforward neural network layers, scaling, shifting, or MLP.
[0133] A typical application of the aforementioned video generation process is the large-scale text-based video model. This model includes the LLM, video compression encoder and decoder, and the DiT model for video generation, with the entire model having billions of parameters. For example, based on current computing power and model capabilities, considering that a 1-minute high-definition video patch is approximately 1MB and requires multiple iterations through the DiT model for video generation, the computing power requirement for generating a 1-minute high-definition video is approximately 10... 6 Approximately one trillion floating-point operations per second (TFLOPs).
[0134] In future communication networks, terminals may also require video generation. However, due to the high computational and power consumption requirements of video generation, current communication devices are limited by their power consumption and computational capabilities, making video generation difficult to achieve. Therefore, how to implement video generation on communication devices within a communication network is a pressing issue that needs to be addressed.
[0135] Figure 3a This is a schematic diagram illustrating an implementation of the video generation method provided in this application. Figure 3aThis application illustrates the method using a first communication device and other communication devices (such as a second communication device) as examples of the execution subjects of the interaction, but it does not limit the execution subjects of the interaction. For example, the communication device can be a communication equipment (such as a terminal device or a network device), or a chip or baseband chip in the communication equipment.
[0136] (baseband) chip, modem chip, SoC chip (such as an SoC chip containing a modem core), SIP chip, communication module, chip system, processor, logic module or software, etc.
[0137] As an example, the first communication device can be a terminal device and the second communication device can be a network device.
[0138] As another example, both the first and second communication devices are terminal devices.
[0139] Optionally, the aforementioned network equipment may be access network equipment or ORAN equipment (including at least one of O-CU, O-DU, and O-RU).
[0140] S301. The first communication device sends first information, and correspondingly, the second communication device receives the first information. The first information is used to instruct the network device to generate first video data.
[0141] Optionally, the first information is used to instruct the network device to generate first video data. This can be understood as: the first information is used to instruct the network device to assist / help / participate in the generation of the first video data using a first model, and / or, the first information is used to instruct the network device to assist / help / participate in the generation of the first video data. For example, the first information may indicate video generation requirements, enabling the network device to generate video-related data that meets those requirements based on the first information (e.g., the network device may refer to...). Figure 2 The video generator in the middle processes the first information to obtain the first data (which will be described later), so that the first communication device can generate the first video data based on the video-related data.
[0142] exist Figure 3aIn one possible implementation of the method shown, the first information is obtained by filtering based on pre-configured rules. For example, the first video demand information of the first communication device for the first video data may indicate multiple video generation demands, wherein the first video demand information may include portions that the first communication device does not wish to send to other communication devices (e.g., one or more types of privacy data, sensitive data, etc.). In the above process, the first communication device can filter the video demand information based on pre-configured rules to obtain the first information (i.e., the first information can indicate the filtered video generation demands), thereby achieving data privacy protection and preventing privacy leaks.
[0143] Optionally, the first communication device can obtain the first video request information in various ways. For example, the first communication device obtains user operation instructions through a user interface and determines the first video request information based on the user operation instructions. In this case, the request indicated by the first video request information can be understood as the user's request. Alternatively, the first communication device can receive the first video request information sent by other devices through a wired or wireless communication interface. In this case, the request indicated by the first video request information can be understood as the request of the other devices. For example, the request can be instructing video generation requirements through text (including text corresponding to natural language), voice, images, video, etc.
[0144] Optionally, the pre-configured rules used for filtering can be determined based on user operation instructions, enabling the first communication device to protect data privacy based on user operation instructions, thereby improving user experience.
[0145] Optionally, the pre-configured rules are either factory-configured by the device or configured by the network device. For example, the pre-configured rules may instruct the filtering of data involving privacy (including but not limited to facial image data, medical image data, etc.). In this case, the first communication device can filter the video requirement information of the first video data based on the pre-configured rules, and the resulting first information includes data that does not involve privacy. As another example, the pre-configured rules may instruct the filtering of user-specified data categories (such as non-background image data). In this case, the first communication device can filter the video requirement information of the first video data based on the pre-configured rules, and the resulting first information includes background image data.
[0146] As an implementation example, the first information is obtained by filtering the first video demand information based on the pre-configured rules; wherein, the first video demand information is obtained by the third model processing the second video demand information, which is described in natural language. Specifically, the second video demand information used to indicate the demand for the first video data can be described in natural language, and the third model can process the second video demand information to obtain the first video demand information. Subsequently, the first communication device can filter the first video demand information based on the pre-configured rules to obtain the first information, thereby achieving data privacy protection.
[0147] Optionally, the third model can be used to process natural language. For example, the third model can be a large language model (LLM), a generative pre-trained transformer (GPT), or other large models that can be used for natural language processing.
[0148] Optionally, the first video requirement information is used for processing the second model (e.g., the first communication device can process the first video requirement information based on the second model to obtain video-related data); or, the first video requirement information and the first information are used to determine third video requirement information, which is then used for processing the second model (e.g., the first communication device can process the third video requirement information based on the second model to obtain video-related data). Based on this, during data processing based on the second model, the input to the second model can include either the first video requirement information or the third video requirement information, ensuring that the data processed by the second model meets video requirements and improves the video quality corresponding to the subsequently generated first video data, thereby enhancing the user experience.
[0149] For example, the multiple demands indicated by the first video demand information may include the demands indicated by the third video demand information and the demands indicated by the first information, and the demands indicated by the third video demand information and the demands indicated by the first information may be different. For instance, the first communication device may filter out the portion that it does not want to send to other communication devices based on pre-configured rules, including the third video demand information, and correspondingly, the portion that is allowed to be sent (or expected to be sent) includes the first information.
[0150] Optionally, the first information includes at least one of the following (requirements): the intent corresponding to the first video data; second video data, which is part or all of the initial video data of the first video data; or, feature data corresponding to the first video data. Similarly, any one of the first video requirement information, the second video requirement information, and the third video requirement information may also include at least one of the above (requirements).
[0151] S302. The second communication device sends first data, and correspondingly, the first communication device receives the first data. The first data is obtained by the first model processing the first information.
[0152] In this application, the model may include an AI model, a neural network model, an AI neural network model, a machine learning model, an AI processing model, or other names defined by future networks.
[0153] It should be noted that the first model (or the second model below) can be a model used to generate video-related data. For example, the first model (or the second model) can be a video generator, a video generation model, a large video generation model, or other names defined by the network in the future.
[0154] As an example, the first model (or second model) may include a diffusion transformer (DiT), which can process video-related data in the latent space.
[0155] As another example, the first model (or second model) may include a visual encoder / decoder. The visual encoder / decoder can be used for the conversion between pixel-space data and latent-space data; for example, a visual encoder can process pixel-space data into latent-space data; similarly, a visual decoder can process latent-space data into pixel-space data.
[0156] As another example, the first model (or second model) may include a large language model (LLM). For instance, an LLM can be used to perform intent recognition processing on video generation requirement information (such as the second video generation requirement information described later) to obtain processed video generation requirement information.
[0157] Optionally, video-related data may include one or more of the following: image and / or video patches, patch vectors, tokens, token vectors, embeddings, or other information. For example, a patch is the basic unit of image processing. A token is the smallest unit of meaning that the model can understand and generate; it is the basic unit of the model. For example, a fragment of a word in the user's input intent; or a vector representation mapped from an image patch, can be considered an abstract representation of a patch. Embedding refers to a relatively low-dimensional vector space into which high-dimensional vectors can be transformed for easier machine learning processing, including but not limited to patch embeddings and token embeddings.
[0158] S303. The first communication device determines the first video data based on the first data.
[0159] As an example, the first communication device may include (or be able to invoke) a second model (e.g., the second model may be the one mentioned above). Figure 2 (As illustrated in the example video generator and decoder, etc.), the first communication device can process the first data based on the second model to determine the first video data.
[0160] As an example, the first communication device may include (or be able to invoke) a video decoder (e.g., as mentioned above). Figure 2 (The decoder shown in the example) Accordingly, the first communication device can process the first data based on the video decoder to determine the first video data.
[0161] exist Figure 3a In one possible implementation of the method shown, the first communication device processes the first data based on the second model to obtain second data; the first communication device determines the first video data based on the second data.
[0162] In other words, after the first communication device processes the first data based on the second model to obtain second data, it determines the first video data based on the second data. For example, the process by which the first communication device determines the first video data based on the second data can be performed using a video decoder (as described above). Figure 2 This can be achieved either through a decoder (in the context of video processing) or through an AI model (e.g., using the second data as input to an AI model to output the first video data). In this way, the first communication device can call the second model to process the first data sent by other communication devices. Parallel collaborative processing can be performed through different models called by different communication devices, providing more resources for video generation and reducing the latency of video generation.
[0163] exist Figure 3a In one possible implementation of the method shown, the first communication device processes the first data and the third data based on the second model to obtain the second data; wherein the third data is intermediate data obtained by the processing of the second model. In this way, the first video data can be obtained through the processing of the first model and the second model, and the quality of the subsequently generated video can be improved through the cooperation of these two models.
[0164] Optionally, the third data is obtained by processing the third video demand information using the second model. For example, the first information is obtained by filtering the first video demand information based on the pre-configured rules; wherein the third video demand information is the same as the first video demand information, or the third video demand information is different from the first information. The third video demand information can be referred to later. Figure 5 And related descriptions.
[0165] As an example (hereinafter referred to as Example A), the second model can be a model deployed on the first communication device, in which the first communication device can use its own deployed model to participate in video generation, thereby reducing processing latency.
[0166] Optionally, in Example A, the second model may be the same module (e.g., software and / or hardware module) integrated with other data processing functions and deployed in the first communication device; or, the second model may be a different module (e.g., software and / or hardware module) integrated with other data processing functions and deployed in the first communication device. The latter will be described exemplarily below through some examples.
[0167] like Figure 3b As shown, this is one implementation example of Example A. The first communication device may include a third communication device and a fourth communication device. The third communication device is used to deploy the second model (i.e., the third communication device has the model function of the second model), and the fourth communication device is used to deploy other data processing functions, including but not limited to the first information transmission in step S301, the first data reception in step S302, and the process of determining the first video data based on the first data and the second data in step S303.
[0168] Optionally, in Figure 3b In addition, the fourth communication device also includes other functions, such as the ability to deploy the third model described later (i.e., the fourth communication device has the modeling function of the third model).
[0169] For example, in Figure 3bIn this process, the third and fourth communication devices can be internal modules of the first communication device. The interaction process between the two includes: Step A, the fourth communication device sends third video demand information to the third communication device, so that the third communication device processes the third video demand information based on the second model in step C to obtain second data, and sends the second data to the fourth communication device in step D.
[0170] Optionally, the fourth communication device may also send the first data to the third communication device in step B, so that the third communication device can process the third video demand information and the first data based on the second model to obtain the second data.
[0171] Optionally, the fourth communication device may also send second data to the second communication device in step E, so that the second communication device can process the second data based on the first model.
[0172] As another example (hereinafter referred to as Example B), the second model can be a model deployed on other communication devices (e.g., a third communication device). For example, the first communication device can receive second data from the third communication device. In this way, the first communication device can utilize models deployed on other communication devices to participate in video generation, thereby reducing the computing power and power consumption of the first communication device.
[0173] like Figure 3c As shown, this is one implementation example of Example B. The first communication device can establish a communication connection with the third communication device via wired or wireless means. Furthermore, the third communication device is used to deploy the second model (i.e., the third communication device possesses the model functionality of the second model), and the first communication device is used to deploy other data processing functions, including but not limited to the first information transmission in step S301, the first data reception in step S302, and the process of determining the first video data based on the first and second data in step S303.
[0174] Optionally, in Figure 3c In addition, the first communication device also includes other functions, such as the ability to deploy the third model described below (i.e., the first communication device has the modeling function of the third model).
[0175] For example, in Figure 3c In this process, the third communication device and the first communication device can be different communication devices. The interaction process between the two includes: Step A, the first communication device sends third video demand information to the third communication device, so that the third communication device processes the third video demand information based on the second model in step C to obtain second data, and sends the second data to the first communication device in step D.
[0176] Optionally, the first communication device may also send first data to the third communication device in step B, so that the third communication device can process the third video demand information and the first data based on the second model to obtain the second data.
[0177] Optionally, the first communication device may also send second data to the second communication device in step E, so that the second communication device can process the second data based on the first model.
[0178] It should be understood that deploying a model on a communication device (e.g., the second model described above can be deployed on the first communication device, and the first model described above can be deployed on the second communication device, etc.) can be described as a model being deployed on a communication device. For example, after a communication device obtains the model parameters of a model, it can obtain / generate / construct the model based on the model parameters, and subsequently, the communication device can process the model. Optionally, the model parameters may include one or more of the following: model hyperparameters, model dataset (including the model's input data and the corresponding label data), and model structural parameters.
[0179] Optionally, the first model is associated with the second model. For example, the first communication device may deploy or invoke one or more models containing the second model, and the second communication device may also deploy or invoke one or more models containing the first model. In the above process, the association of the first model with the second model can be understood as follows: the first model is a model that matches the second model; or, the first model and the second model may be models used to perform the same task (e.g., the task is a video generation task); or, the first model and the second model may be models with the same or similar functions (e.g., the function is a video generation function).
[0180] For example, the first communication device can determine, through configuration or pre-configuration, to use the second model in the generation of the first video data. Similarly, the second communication device can also determine, through configuration or pre-configuration, to use the second model in the generation of the first video data.
[0181] For example, the first communication device may send instruction information to the second communication device, which is used to instruct the first model and / or the second model, so that the second communication device determines to use the second model to determine the first data based on the instruction information.
[0182] For example, the second communication device can send instruction information to the first communication device, which is used to instruct the first model and / or the second model, so that the first communication device determines to use the first model to determine the second data based on the instruction information.
[0183] exist Figure 3aIn one possible implementation of the method shown, the first data is part or all of the input data of the first AI module included in the first model, and the second data is obtained by processing the first data through the second AI module included in the second model; wherein the first AI module is associated with the second AI module. Specifically, a model (e.g., a first model or a second model) may include one or more AI models. In the above process, the first data sent by the second communication device may be the input data of the first AI module (or the first data may be the output data of the previous module of the first AI module), and the module of the second model that processes the first data to obtain the second data may be the second AI module. Furthermore, these two modules may be associated modules. In this way, during the collaborative video generation process of the first model and the second model, the two models can perform data processing through the associated AI module, which can improve data processing performance and the quality of subsequent video generation, thereby enhancing the user experience.
[0184] In this application, the term "AI module" may be replaced with other terms, such as neural network module, AI processing module, processing module, or other names defined in the future network definition.
[0185] Optionally, the association between the first AI module and the second AI module may include: the functions of the first AI module and the second AI module being the same or similar; and / or, the first index of the first AI module in the multiple AI modules included in the first model being associated with the second index of the second AI module in the multiple AI modules included in the second model. In this way, since the data processed by AI modules with associated indices in different models can be strongly correlated (for example, the first index corresponding to the first AI model and the second index corresponding to the second AI module are correlated, meaning the data processed by the first AI module and the data processed by the second AI module can be strongly correlated, for example, these two data may be used to generate video data for the same frame or adjacent frames), the second AI module can process the data output by the first AI module, thereby improving the quality of the subsequently generated video.
[0186] More examples will be provided below for further description.
[0187] As an implementation example, the first index of the first AI module in the multiple AI modules included in the first model differs from the second index of the second AI module in the multiple AI modules included in the second model by 1 (e.g., the first index is less than the second index). In this way, the input of the second AI module corresponding to a certain index in the second model can include the input of the first AI module corresponding to the previous index in the first model, so that the data features corresponding to the input of the first AI module can be used as one of the processing bases of the second AI module, thereby improving the processing performance of the AI module.
[0188] For example, when the first communication device sends the second data, the AI network architecture of the first model is the same as or similar to that of the second model, and / or the AI modules with the same index included in the first and second models have the same or similar functions. For example, both the first AI module and the second AI module are self-attention modules.
[0189] like Figure 4a The image shows an application example of the above solution. Figure 4a In this example, the first AI model contains at least N AI modules as shown in the diagram, and the second AI model contains at least N AI modules as shown in the diagram. Figure 4a In this model, the first data can be the output of the i-th AI module (or the input of the (i+1)-th AI module, where i ranges from 1 to N), and the second data can be the output of the (i+1)-th AI module (or part or all of the input of the (i+2)-th AI module). In this way, the first and second AI models can achieve parallel processing through synchronous parallelism, thereby improving the quality of the generated video.
[0190] Optionally, the first communication device may choose not to send (or decide not to send) the second data. For example, this approach could be called permutation parallelism or asynchronous parallelism. By doing so, compared to different nodes interacting in parallel to generate video, interaction overhead can be reduced, thereby reducing the processing latency of video generation by reducing transmission latency. Furthermore, if the second data is data that the first communication device does not want to send to other communication devices (e.g., one or more types of private or sensitive data), the first communication device's decision not to send the second data also achieves data privacy protection, preventing privacy leaks and improving user experience.
[0191] In one possible implementation, Figure 3aThe method also includes the first communication device sending the second data. Specifically, after the first communication device processes the first data based on the second model to obtain the second data, the first communication device can send the second data. For example, this method can be called synchronous parallel processing or interactive parallel processing. In this way, the second communication device can perform subsequent processing based on the second data, and can perform parallel collaborative processing through different models called by different communication devices to improve the quality of the generated video and thus improve the user experience.
[0192] Optionally, after receiving the second data, the second communication device can use the second data as the first AI module in the first model (e.g., Figure 4a The first model contains part or all of the input of the next AI module (i.e., AI module_i+2), which is the (i+2)th AI module, so that the second communication device can be used without executing the second AI module (e.g., Figure 4a The AI processing procedure corresponding to the (i+1)th AI module in the first model can reduce the processing latency of the first model.
[0193] Optionally, after receiving the second data, the second communication device can process the second data through one or more AI modules included in the first model, and the second communication device can send the processing result to the first communication device so that the first communication device can determine the first video data based on the processing result. In this way, the first and second communication devices can achieve parallel collaborative processing through two or more interaction processes to further improve the quality of the generated video.
[0194] For example, if the first communication device does not send (or determines not to send) the second data, the AI network architecture of the first model is the same as or similar to the AI network architecture of the second model, and / or, the AI modules with the same index included in the first model and the second model have the same or similar functions. For example, the first AI module is a self-attention module, and / or, the second AI modules are all cross-attention modules.
[0195] Optionally, if the first communication device may not send (or is determined not to send) the second data, this method may be called permutation parallelism or asynchronous parallelism, etc.
[0196] like Figure 4b The image shows an application example of the above solution. Figure 4b In this example, the first AI model contains at least N AI modules as shown in the diagram, and the second AI model contains at least N AI modules as shown in the diagram. Figure 4bIn this context, the first data can be the output of the i-th AI module of the first model (or the input of the (i+1)-th AI module, where i takes values from 1 to N). In this way, the first AI model and the second AI model can achieve parallel processing through permutation parallelism, thereby improving the quality of the generated video.
[0197] exist Figure 3a In one possible implementation of the method, the method further includes: the first communication device receiving or sending second information, which is used to instruct the first AI module and / or the second AI module. For example, if the first communication device sends the second information, it can be understood that the first communication device determines the AI module to be executed in parallel. Similarly, if the first communication device receives the second information, it can be understood that the second communication device determines the AI module to be executed in parallel. Thus, the recipient of the second information can determine, based on the instruction information, that the AI module corresponding to the first data is the first AI module and / or the second AI module, and can subsequently use the corresponding AI module for data processing.
[0198] The following will be through Figure 5 The example shown illustrates the above models and video requirements information.
[0199] like Figure 5 As shown, in the processing involved in the first communication device, the third model (e.g., the GPT / LLM illustrated) can process the second video demand information described by natural language (e.g., intent recognition processing) to obtain the first video demand information, which can be filtered to obtain the first information.
[0200] For example, the first communication device can filter the first video demand information, and the obtained first information includes the parts of the first video demand information that are not expected to be sent (e.g., parts that require privacy protection), and the first information can then be used for the processing of the first model.
[0201] For example, a portion of the input of the second model in the first communication device (e.g., input ① in the figure) may include third video demand information, which may be the same as the first video demand information, or the third video demand information may include the part of the first video demand information that is not expected to be sent (e.g., the part that involves privacy protection), that is, the third video demand information may be different from the first information.
[0202] Optionally, the filtering process may also involve second video data, i.e., the video data of the initial video corresponding to the first video data. Similarly, the filtering process for the second video data can refer to the filtering process for the first video requirement information described above.
[0203] For example, the first information may also include the portion of the second video data that is allowed to be sent (or expected to be sent) after filtering out portions that are not intended to be sent (e.g., portions that require privacy protection).
[0204] For example, the input ① of the second model (i.e., the third video requirement information) may also include the second video data, or the input ① of the second model (i.e., the third video requirement information) may also include the parts of the second video data that are not expected to be sent (e.g., parts that require privacy protection).
[0205] Optionally, the second video data can be implemented in other ways. For example, the second video data can be determined, generated, or provided by the second communication device. Alternatively, the second video data can be a portion determined independently by the first and second communication devices to reduce transmission overhead. Or, the second video data does not need to be generated; for example, the processing of the first model and / or the second model does not require initial video.
[0206] Furthermore, the processing of the first model based on the first information can be performed in parallel with the second model (e.g., synchronous parallelism and / or permutation parallelism as described above) to improve the quality of the subsequently generated video. For example, the data provided by the first model (e.g., the first data mentioned above) can be used as input ② in the figure, enabling the first communication device to offload some data processing to the second communication device. This reduces the resource consumption (including computing power and power) of the first communication device while also reducing the processing latency of video data generation, thereby improving the user experience.
[0207] Optionally, the above process may involve a third model, as shown in Examples A and B above. There may be multiple deployment methods for the third model, and correspondingly, there may be corresponding interaction processes for the above video requirement information. Some examples will be used to illustrate this below.
[0208] In example A, such as Figure 3b As shown, the first communication device includes a third communication device and a fourth communication device. The fourth communication device can deploy a third model, and in the above process, the fourth communication device can process the second video demand information based on the third model to obtain the first video demand information. Furthermore, the fourth communication device can send the first video demand information or the third video demand information to the third communication device in step A, so that the third communication device can process the first video demand information or the third video demand information in step C to obtain the second data. Subsequently, the third communication device can send the second data to the first communication device in step D, enabling the fourth communication device to further determine the first video data. For example, the basis for determining the first video data may include the first data, and optionally, it may also include the second data, etc.
[0209] Optionally, in Figure 3b In step B, the fourth communication device can also send the first data to the third communication device so that the third communication device can process the first video demand information or the third video demand information and the first data through the second model to obtain the second data, thereby improving the quality of the subsequently generated video through collaborative processing.
[0210] Optionally, in Figure 3b In step E, the first or fourth communication device may also send second data to the second communication device. The processing of the second data can be referred to the above description.
[0211] In example B, such as Figure 3c As shown, the first communication device and the third communication device may be connected via wired or wireless means. The first communication device may deploy a third model, and in the above process, the first communication device can process the second video demand information based on the third model to obtain the first video demand information. Furthermore, the first communication device may send the first video demand information or the third video demand information to the third communication device in step A, so that the third communication device can process the first video demand information or the third video demand information in step C to obtain the second data. Subsequently, the third communication device may send the second data to the first communication device in step D, enabling the first communication device to further determine the first video data. For example, the determination of the first video data may include the first data, and optionally, it may also include the second data, etc.
[0212] Optionally, in Figure 3b In step B, the first communication device can also send first data to the third communication device so that the third communication device can process the first video demand information or the third video demand information and the first data through the second model to obtain second data, thereby improving the quality of the subsequently generated video through collaborative processing.
[0213] Optionally, in Figure 3b In step E, the first communication device may also send second data to the second communication device, and the processing of the second data can be referred to the above description.
[0214] In one possible implementation, Figure 3a The method further includes: the first communication device sending third information, which requests processing resources for the first video data. Specifically, the first communication device can also request resources through the third information, enabling the recipient of the third information to reserve resources for the processing of the first video data, thereby ensuring the generation of the first video data.
[0215] Optionally, the resources requested by the third information may include one or more of the computing resources, transmission resources, data storage or caching resources of the related data of the first video data (for example, the related data may include one or more of the first data, the second video data or other data mentioned above).
[0216] For example, taking the second video data as an example, the third information may indicate the size of the second video data, and / or the third information may indicate the number of patches contained in the second video data.
[0217] Optionally, the third information may also carry indication information indicating the first video data. For example, the indication information may indicate the index or identifier of the first video data, and / or the task index or task identifier of the task corresponding to the first video data (e.g., video generation task).
[0218] Optionally, the first communication device may also receive a fourth message, which is a response to the request indicated by the third message. For example, the fourth message may be an acknowledgement (ACK) of the request, indicating acceptance of the request. Alternatively, the fourth message may be a negative acknowledgement (NACK) of the request, indicating rejection of the request. Furthermore, the fourth message may also indicate modification of the requested resources, including but not limited to situations where the second communication device determines that the resources requested by the third message are inappropriate or that current processing resources are insufficient.
[0219] based on Figure 3a In the illustrated scheme, after the first communication device sends first information instructing the generation of first video data to the second communication device in step S301, the second communication device can process the first data based on the first model and the first information, and then send the first data to the first communication device in step S302. This allows the first communication device to determine the first video data based on the first data in step S303. In this way, the second communication device can use the first model to process the video data generation instruction from the first communication device, obtain and send the first data, enabling the first communication device to generate the first video data based on the first data generated by the model of the second communication device. This allows video generation to be achieved through collaborative processing between different communication devices.
[0220] Furthermore, during the aforementioned video data processing, the first communication device can offload some data processing to other communication devices (such as the second communication device described below) for processing. This reduces the resource consumption (including computing power and power) of the first communication device and also reduces the processing latency of video data generation, thereby improving the user experience.
[0221] Please see Figure 6 This is a schematic diagram illustrating another implementation of the video generation method provided in this application. Figure 6 In this example, the first communication device is taken as the terminal equipment, and the second communication device is taken as the network equipment. Figure 6 As shown, the terminal device may include a model inference interface (e.g., a user interface (UI) that can be used to receive user operation commands, play audio and video, etc.), a terminal compression module (taking the generation of the initial video on the terminal as an example), a terminal large model (such as the third model mentioned above), a terminal DiT module (such as the second model mentioned above), and a terminal decoding module.
[0222] Step 1. The model inference interface inputs a request to the large model on the edge based on user commands. This request can indicate video generation requirements. Optionally, the request can be described in natural language.
[0223] Step 2. The client-side compression module filters based on the generated initial video (e.g., filtering initial video data to determine if privacy protection is required), and the client-side large model filters based on the request from Step 1 (e.g., filtering to determine if privacy protection is required). Step 2 is optional.
[0224] Step 3. The end-side compression module estimates the amount of resources (including transmission resources, processing resources, etc.) required to generate the video based on the initial video it generates, and sends a resource request to the network device.
[0225] Step 4. If the network device determines that it is allocating resources to the terminal device, it can send a resource request response to the terminal device.
[0226] Step 5. The end-side compression module sends the first information. This first information can be either filtered in step 2 or not.
[0227] It should be understood that step 5 is an implementation example of step S301 mentioned above.
[0228] Step 6. The network device processes data based on the first information.
[0229] Step 7. Instructions for parallel processing are exchanged between network devices and terminal devices. For example, the instructions for parallel processing may be used to instruct network devices and / or terminal devices to perform parallel processing of AI module information (e.g., the index, number, etc. of AI modules), and / or the instructions for parallel processing may be used to instruct network devices and / or terminal devices to perform parallel processing of AI module information (e.g., the index, number, etc. of AI modules).
[0230] Step 8. The terminal device and network device perform parallel permutation based on the AI module information determined in Step 7. The implementation process of parallel permutation can be referred to the previous description, for example, the previous... Figure 4b And related examples.
[0231] Step 9. The terminal device and network device perform synchronous parallel operation based on the AI module information determined in Step 7. The implementation process of synchronous parallel operation can be referred to the previous description, for example, the previous... Figure 4a And related examples.
[0232] Step 10a. The network device sends video-related data to the end-side decoding module.
[0233] Step 10b. The end-side DiT module sends video-related data to the end-side decoding module.
[0234] It should be understood that the first data described above (e.g., the first data in step S302) may include one or more of the data sent by the network device during the replacement parallel process in step 8, the data sent by the network device during the synchronization parallel process in step 9, and the video-related data sent by the network device in step 10a.
[0235] Step 11. The end-side decoding module performs video decoding based on the received data to obtain the decoded video data (such as the first video data described above).
[0236] It should be understood that step 11 is an implementation example of step S303 mentioned above.
[0237] Step 12. The edge decoding module sends the decoded video data to the model inference interface to enable video playback.
[0238] pass Figure 6 As shown in the implementation example, on the one hand, some data from the device side can be offloaded to the network side for processing, reducing the computing power overhead on the device side; on the other hand, it satisfies the requirement that user privacy data is processed locally on the device side and does not need to be disclosed. This can both protect important local data on the device side from being leaked and reduce model inference latency.
[0239] Please see Figure 7 This application provides a communication device 700, which can realize the functions of the second or first communication device in the above method embodiments, and thus also achieve the beneficial effects of the above method embodiments. In this application embodiment, the communication device 700 can be the first communication device (or the second communication device), or it can be an integrated circuit or component inside the first communication device (or the second communication device), such as a chip.
[0240] It should be noted that the transceiver unit 702 may include a transmitting unit and a receiving unit, which are used to perform transmitting and receiving respectively.
[0241] In one possible implementation, when the device 700 is used to execute the method performed by the first communication device in the foregoing embodiments, the device 700 includes a processing unit 701 and a transceiver unit 702; the transceiver unit 702 is used to send first information, which instructs a network device to generate first video data; the transceiver unit 702 is also used to receive first data, which is obtained by the network device processing the first information based on a first model; the processing unit 701 is used to determine the first video data based on the first data.
[0242] In one possible implementation, when the device 700 is used to execute the method performed by the second communication device in the foregoing embodiments, the device 700 includes a processing unit 701 and a transceiver unit 702; the transceiver unit 702 is used to receive first information, which instructs a network device to generate first video data; the processing unit 701 is used to determine the first data; the transceiver unit 702 is also used to send the first data, which is obtained by the network device processing the first information based on a first model; wherein, the first data is used to determine the first video data.
[0243] In one possible implementation, when the device 700 is used to execute the method performed by the third communication device in the foregoing embodiments, the device 700 includes a processing unit 701 and a transceiver unit 702; the processing unit 701 is used to acquire third video demand information, which is used to indicate the generation of first video data; the processing unit 701 is also used to process the third video demand information based on a second model to obtain second data; the transceiver unit 702 is used to send the second data; wherein, the second data is used to determine the first video data.
[0244] It should be noted that the information execution process of the unit of the above-mentioned communication device 700 can be specifically described in the method embodiment shown above in this application, and will not be repeated here.
[0245] Please see Figure 8 This is another schematic structural diagram of the communication device 800 provided in this application. The communication device 800 includes a logic circuit 801 and an input / output interface 802. The communication device 800 can be a chip or an integrated circuit.
[0246] in, Figure 7 The transceiver unit 702 shown can be a communication interface, which can be... Figure 8The input / output interface 802 may include an input interface and an output interface. Alternatively, the communication interface may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0247] Optionally, the input / output interface 802 sends first information, which is used to instruct the network device to generate first video data; the input / output interface 802 is also used to receive first data, which is obtained by the network device processing the first information based on the first model; the logic unit 801 determines the first video data based on the first data.
[0248] Optionally, the input / output interface 802 receives first information, which is used to instruct the network device to generate first video data; the logic unit 801 determines the first data; the input / output interface 802 sends the first data, which is obtained by the network device processing the first information based on the first model; wherein, the first data is used to determine the first video data.
[0249] Optionally, the logic unit 801 acquires third video demand information, which is used to indicate the generation of first video data; the logic unit 801 processes the third video demand information based on the second model to obtain second data; the input / output interface 802 is used to send the second data; wherein, the second data is used to determine the first video data.
[0250] The logic circuit 801 and the input / output interface 802 can also perform other steps performed by the first or second communication device in any embodiment and achieve corresponding beneficial effects, which will not be elaborated here.
[0251] In one possible implementation, Figure 7 The processing unit 701 shown can be Figure 8 The logic circuit 801 in the middle.
[0252] Optionally, the logic circuit 801 can be a processing device, the functions of which can be partially or entirely implemented in software.
[0253] Optionally, the processing apparatus may include a memory and a processor, wherein the memory is used to store a computer program, and the processor reads and executes the computer program stored in the memory to perform the corresponding processing and / or steps in any of the method embodiments.
[0254] Optionally, the processing device may consist of only a processor. A memory for storing computer programs is located outside the processing device, and the processor is connected to the memory via circuitry / wires to read and execute the computer programs stored in the memory. The memory and processor may be integrated together or physically independent of each other.
[0255] Optionally, the processing device may be one or more chips, or one or more integrated circuits. For example, the processing device may be one or more field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), system on-chips (SoCs), central processors (CPUs), network processors (NPs), digital signal processors (DSPs), microcontroller units (MCUs), programmable logic devices (PLDs), or other integrated chips, or any combination of the above chips or processors.
[0256] Please see Figure 9 The communication device 900 provided in the above embodiments of this application can specifically be the communication device that serves as a terminal device in the above embodiments. Figure 9 The example shown illustrates how a terminal device can be implemented through a terminal device (or a component within a terminal device).
[0257] The present invention provides a possible logical structure diagram of the communication device 900, which may include, but is not limited to, at least one processor 901 and a communication port 902.
[0258] in, Figure 7 The transceiver unit 702 shown can be a communication interface, which can be... Figure 9 The communication port 902 in the diagram may include an input interface and an output interface. Alternatively, the communication port 902 may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0259] Further optionally, the device may also include at least one of a memory 903 and a bus 904. In the embodiments of this application, the at least one processor 901 is used to control the operation of the communication device 900.
[0260] Furthermore, the processor 901 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, etc. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0261] It should be noted that, Figure 9 The communication device 900 shown can be used to implement the steps implemented by the terminal device in the aforementioned method embodiments, and to achieve the corresponding technical effects of the terminal device. Figure 9 The specific implementation of the communication device shown can be referred to the description in the foregoing method embodiments, and will not be repeated here.
[0262] Please see Figure 10 The above-described embodiments of the communication device 1000 provided as an example of the present application are schematic diagrams of its structure. Specifically, the communication device 1000 can be a network device as described in the above embodiments. Figure 10 The example shown illustrates a network device implemented through a network device (or a component within a network device). The structure of this communication device can be referenced. Figure 10 The structure shown.
[0263] The communication device 1000 includes at least one processor 1011 and at least one network interface 1014. Optionally, the communication device further includes at least one memory 1012, at least one transceiver 1013, and one or more antennas 1015. The processor 1011, memory 1012, transceiver 1013, and network interface 1014 are connected, for example, via a bus. In this embodiment, the connection may include various interfaces, transmission lines, or buses, etc., and this embodiment is not limited thereto. The antenna 1015 is connected to the transceiver 1013. The network interface 1014 enables the communication device to communicate with other communication devices through a communication link. For example, the network interface 1014 may include a network interface between the communication device and a core network device, such as an S1 interface; the network interface may also include a network interface between the communication device and other communication devices (e.g., other network devices or core network devices), such as an X2 or Xn interface.
[0264] in, Figure 7The transceiver unit 702 shown can be a communication interface, which can be... Figure 10 The network interface 1014 may include an input interface and an output interface. Alternatively, the network interface 1014 may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0265] The processor 1011 is primarily used to process communication protocols and communication data, control the entire communication device, execute software programs, and process data from the software programs, for example, to support the communication device in performing the actions described in the embodiments. The communication device may include a baseband processor and a central processing unit (CPU). The baseband processor is primarily used to process communication protocols and communication data, while the CPU is primarily used to control the entire terminal device, execute software programs, and process data from the software programs. Figure 10 The processor 1011 can integrate the functions of a baseband processor and a central processing unit. Those skilled in the art will understand that the baseband processor and the central processing unit can also be independent processors interconnected via technologies such as buses. Those skilled in the art will understand that a terminal device can include multiple baseband processors to adapt to different network standards, and a terminal device can include multiple central processing units to enhance its processing capabilities. The various components of the terminal device can be connected via various buses. The baseband processor can also be described as a baseband processing circuit or a baseband processing chip. The central processing unit can also be described as a central processing circuit or a central processing chip. The function of processing communication protocols and communication data can be built into the processor or stored in memory as a software program, with the processor executing the software program to implement the baseband processing function.
[0266] The memory is primarily used to store software programs and data. The memory 1012 can exist independently or be connected to the processor 1011. Optionally, the memory 1012 can be integrated with the processor 1011, for example, integrated within a single chip. The memory 1012 can store program code that executes the technical solutions of the embodiments of this application, and its execution is controlled by the processor 1011. The various types of computer program code being executed can also be considered as drivers for the processor 1011.
[0267] Figure 10 Only one memory and one processor are shown. In actual terminal devices, there may be multiple processors and multiple memories. Memory can also be called storage medium or storage device, etc. Memory can be a storage element on the same chip as the processor, i.e., an on-chip storage element, or it can be a separate storage element; this application does not limit this.
[0268] Transceiver 1013 can be used to support the reception or transmission of radio frequency (RF) signals between a communication device and a terminal. Transceiver 1013 can be connected to antenna 1015. Transceiver 1013 includes a transmitter Tx and a receiver Rx. Specifically, one or more antennas 1015 can receive RF signals. The receiver Rx of transceiver 1013 is used to receive the RF signals from the antennas, convert the RF signals into digital baseband signals or digital intermediate frequency (IF) signals, and provide the digital baseband signals or IF signals to processor 1011 so that processor 1011 can perform further processing on the digital baseband signals or IF signals, such as demodulation and decoding. In addition, the transmitter Tx in transceiver 1013 is also used to receive modulated digital baseband signals or IF signals from processor 1011, convert the modulated digital baseband signals or IF signals into RF signals, and transmit the RF signals through one or more antennas 1015. Specifically, the receiver Rx can selectively perform one or more stages of downmixing and analog-to-digital conversion on the radio frequency signal to obtain a digital baseband signal or a digital intermediate frequency (IF) signal. The order of these downmixing and IF conversion processes is adjustable. The transmitter Tx can selectively perform one or more stages of upmixing and digital-to-analog conversion on the modulated digital baseband signal or digital IF signal to obtain a radio frequency signal. The order of these upmixing and IF conversion processes is also adjustable. The digital baseband signal and the digital IF signal can be collectively referred to as digital signals.
[0269] The transceiver 1013 can also be called a transceiver unit, transceiver, transceiver device, etc. Optionally, the device in the transceiver unit that performs the receiving function can be regarded as the receiving unit, and the device in the transceiver unit that performs the transmitting function can be regarded as the transmitting unit. That is, the transceiver unit includes a receiving unit and a transmitting unit. The receiving unit can also be called a receiver, input port, receiving circuit, etc., and the transmitting unit can be called a transmitter, transmitter, or transmitting circuit, etc.
[0270] It should be noted that, Figure 10 The communication device 1000 shown can be used to implement the steps implemented by the network device in the aforementioned method embodiments, and to achieve the corresponding technical effects of the network device. Figure 10 The specific implementation of the communication device 1000 shown can be referred to the description in the foregoing method embodiments, and will not be repeated here.
[0271] Please see Figure 11 The above-described embodiments of the communication device provided in this application are schematic diagrams of the structure of the communication device.
[0272] It is understood that the communication device 110 includes, for example, modules, units, elements, circuits, or interfaces, which are appropriately configured together to execute the technical solutions provided in this application. The communication device 110 may be the terminal device or network device described above, or a component (e.g., a chip) within these devices, used to implement the methods described in the following method embodiments. The communication device 110 includes one or more processors 111. The processor 111 may be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, and the central processing unit can be used to control the communication device (e.g., a RAN node, terminal, or chip), execute software programs, and process data from the software programs.
[0273] Optionally, in one design, the processor 111 may include a program 113 (sometimes also referred to as code or instructions), which can be executed on the processor 111 to cause the communication device 110 to perform the methods described in the embodiments below. In yet another possible design, the communication device 110 includes circuitry (…). Figure 11 (Not shown).
[0274] Optionally, the communication device 110 may include one or more memories 112 storing a program 114 (sometimes referred to as code or instructions), which can be run on the processor 111 to cause the communication device 110 to perform the methods described in the above method embodiments.
[0275] Optionally, the processor 111 and / or memory 112 may include AI modules 117 and 118, which are used to implement AI-related functions. The AI modules can be implemented through software, hardware, or a combination of both. For example, the AI module may include a radio intelligence control (RIC) module. For example, the AI module may be a near real-time RIC or a non-real-time RIC.
[0276] Optionally, the processor 111 and / or memory 112 may also store data. The processor and memory may be configured separately or integrated together.
[0277] Optionally, the communication device 110 may further include a transceiver 115 and / or an antenna 116. The processor 111, sometimes referred to as a processing unit, controls the communication device (e.g., a RAN node or terminal). The transceiver 115, sometimes referred to as a transceiver unit, transceiver, transceiver circuit, or transceiver, is used to realize the transmission and reception functions of the communication device through the antenna 116.
[0278] in, Figure 7 The processing unit 701 shown may be a processor 111. Figure 7 The transceiver unit 702 shown can be a communication interface, which can be... Figure 11 The transceiver 115 may include an input interface and an output interface. Alternatively, the transceiver 115 may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0279] This application also provides a computer-readable storage medium for storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor performs the method described in the possible implementations of the first or second communication device in the foregoing embodiments.
[0280] This application also provides a computer program product (or computer program) containing programs or instructions. When the computer program product is executed by the processor, the processor executes the method of the first communication device or the second communication device that may be implemented as described above.
[0281] This application also provides a chip system including at least one processor for supporting a communication device in implementing the functions involved in the possible implementations of the communication device described above. Optionally, the chip system further includes an interface circuit that provides program instructions and / or data to the at least one processor. In one possible design, the chip system may further include a memory for storing the program instructions and data necessary for the communication device. The chip system may be composed of chips or may include chips and other discrete devices, wherein the communication device may specifically be the first communication device or the second communication device in the aforementioned method embodiments.
[0282] This application also provides a communication system, the network system architecture of which includes a first communication device and a second communication device in any of the above embodiments.
[0283] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0284] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0285] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A video generation method, characterized in that, include: Send first information to the network device, the first information being used to instruct the network device to generate first video data; Receive first data from the network device, wherein the first data is obtained by the network device processing the first information based on a first model; The first video data is determined based on the first data.
2. The method according to claim 1, characterized in that, Determining the first video data based on the first data includes: The first data is processed based on the second model to obtain the second data; The first video data is determined based on the second data.
3. The method according to claim 2, characterized in that, The process of processing the first data based on the second model to obtain the second data includes: The second data is obtained by processing the first data and the third data based on the second model; wherein the third data is intermediate data obtained by processing the second model.
4. The method according to claim 2 or 3, characterized in that, The first data is part or all of the output data of the first artificial intelligence (AI) module included in the first model, and the second data is obtained by processing the first data through the second AI module included in the second model; The first AI module is associated with the second AI module.
5. The method according to claim 4, characterized in that, The first index of the first AI module in the plurality of AI modules included in the first model is associated with the second index of the second AI module in the plurality of AI modules included in the second model.
6. The method according to any one of claims 2 to 5, characterized in that, The method further includes: The second data is sent to the network device.
7. The method according to any one of claims 4 to 6, characterized in that, The method further includes: Receive or send second information, the second information being used to instruct the first AI module and / or the second AI module.
8. The method according to any one of claims 1 to 7, characterized in that, The first piece of information is obtained by filtering based on pre-configured rules.
9. The method according to claim 8, characterized in that, The pre-configured rules are determined based on user operation commands.
10. The method according to claim 8 or 9, characterized in that, The first information is obtained by filtering the first video demand information based on the pre-configured rules; The first video demand information is obtained by the third model based on the second video demand information, which is described in natural language.
11. The method according to any one of claims 1 to 10, characterized in that, The first information includes at least one of the following: The intent corresponding to the first video data; Second video data, wherein the second video data is part or all of the initial video data of the first video data; or The feature data corresponding to the first video data.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: A third message is sent to the network device, the third message being used to request processing resources for the first video data.
13. A video generation method, characterized in that, include: Receive first information, the first information being used to instruct the network device to generate first video data; Send first data, which is obtained by the network device processing the first information based on the first model; wherein, the first data is used to determine the first video data.
14. The method according to claim 13, characterized in that, The first data is used to obtain the second data through processing by the second model, and the second data is used to determine the first video data.
15. The method according to claim 14, characterized in that, The first data is used to obtain the second data through processing by the second model, including: The first data and the third data are used to obtain the second data through the processing of the second model; wherein, the third data is intermediate data obtained by the processing of the second model.
16. The method according to any one of claims 13 to 15, characterized in that, The first data is part or all of the input data of the first artificial intelligence (AI) module included in the first model, and the second data is obtained by processing the first data through the second AI module included in the second model; The first AI module is associated with the second AI module.
17. The method according to claim 16, characterized in that, The first index of the first AI module in the plurality of AI modules included in the first model is associated with the second index of the second AI module in the plurality of AI modules included in the second model.
18. The method according to any one of claims 14 to 17, characterized in that, The method further includes: Receive the second data.
19. The method according to any one of claims 16 to 18, characterized in that, The method further includes: Receive or send second information, the second information being used to instruct the first AI module and / or the second AI module.
20. The method according to any one of claims 13 to 19, characterized in that, The first piece of information is obtained by filtering based on pre-configured rules.
21. The method according to claim 20, characterized in that, The pre-configured rules are determined based on user operation commands.
22. The method according to claim 20 or 21, characterized in that, The first information is obtained by filtering the first video demand information based on the pre-configured rules; The first video demand information is obtained by the third model based on the second video demand information, which is described in natural language.
23. The method according to any one of claims 13 to 22, characterized in that, The first information includes at least one of the following: The intent corresponding to the first video data; Second video data, wherein the second video data is part or all of the initial video data of the first video data; or The feature data corresponding to the first video data.
24. The method according to any one of claims 13 to 23, characterized in that, The method further includes: Receive third information, which is used to request processing resources for the first video data.
25. A communication device, characterized in that, Includes a module for performing the method as described in any one of claims 1 to 24.
26. A communication device, characterized in that, It includes at least one processor, said at least one processor being used to perform the method as described in any one of claims 1 to 24.
27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when executed by a communication device, implement the method as described in any one of claims 1 to 24.
28. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a computer, implement the method as described in any one of claims 1 to 24.
29. A chip or chip system, characterized in that, It includes at least one processor, said at least one processor being used to implement the method as claimed in any one of claims 1 to 24.