Chip system and terminal device
By deploying AIGC chips and switching modules in terminal devices, dynamically adjusting data transmission priorities, solving the privacy and delay problems in interaction between terminal devices and remote service devices, and achieving efficient and secure content generation task execution.
Patent Information
- Application Number
- CN202510263851.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-06
AI Technical Summary
In the prior art, when the terminal device interacts with the remote service device to realize the content generation task, there are problems of poor privacy and high delay.
Deploy chip systems in terminal devices, including application processors and artificial intelligence content generation (AIGC) chips. The AIGC chip has built-in content generation model, connects to application processors and other chips through switching modules, dynamically adjusts data transmission priority and resource allocation, and reduces dependence on remote servers.
It improves the privacy of terminal devices and task execution efficiency, reduces network needs and delays, and improves user experience.
Smart Images

Figure CN119807118B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of terminal devices, and particularly to a chip system and a terminal device. Background Art
[0002] Artificial intelligence generated content (AIGC) refers to the process and result of automatically creating or generating various types of content using artificial intelligence technology, and these contents can cover various forms such as text, images, audio, and video. Currently, the implementation of content generation tasks mainly relies on the interaction between a terminal device and a remote service device. However, there are problems of poor privacy and high latency. Summary of the Invention
[0003] The embodiments of the present application provide a chip system and a terminal device, which are used to improve the problems of poor privacy, the need to access the network, and high latency existing in the implementation of content generation tasks relying on the interaction between a terminal device and a remote service device.
[0004] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0005] In a first aspect, the embodiments of the present application provide a chip system, including an application processor (AP) and an artificial intelligence generated content AIGC chip. A content generation model is deployed inside the AIGC chip. The AP is configured to: send first data to the AIGC chip. The AIGC chip is configured to: perform a corresponding content generation task according to the first data and the content generation model.
[0006] The chip system provided by the embodiments of the present application can complete content generation tasks on a terminal device, reduce dependence on a remote server, and thus reduce network requirements. Since data does not need to be frequently uploaded to the cloud for calculation, the user's input data is always processed locally, avoiding the risk of data leakage during transmission and significantly improving privacy. At the same time, the AIGC chip can select and call a corresponding content generation model according to the type of the first data sent by the AP, enabling various content generation tasks to be directly and efficiently executed locally, thus avoiding the latency caused by remote calculation and data transmission back, and improving task execution efficiency and user experience.
[0007] In an embodiment of the first aspect of the present application, the chip system further includes a switching module and a first chip. The switching module includes a first PCIe interface, a second PCIe interface, and a third PCIe interface. The switching module is connected to the AP through the first PCIe interface, connected to the AIGC chip through the second PCIe interface, and connected to the first chip through the third PCIe interface. The AP is configured to: send first data and second data to the switching module. The switching module is configured to: forward the first data to the AIGC chip, forward the second data to the first chip, or receive third data sent by the AIGC chip and fourth data sent by the first chip and send them to the AP.
[0008] In the chip system provided by the embodiment of the present application, by introducing a switching module, the AIGC chip can be efficiently deployed in the terminal device. When the PCIe interface on the AP side is limited, the bandwidth of the PCIe interface can be fully utilized. The switching module can be respectively connected to the AP, the AIGC chip, and the first chip, so as to build an efficient data interaction channel. The AP can manage the content generation task calculation and other calculation tasks of the first chip through the switching module at the same time. Therefore, the resources of the existing PCIe interface can be fully utilized, the additional PCIe interface overhead introduced by independently deploying the AIGC chip can be avoided, and the hardware resource utilization rate of the system can be improved.
[0009] In an embodiment of the first aspect of the present application, the AP is configured to: obtain the current application scenario of the terminal device, and send a time slice length parameter, a data transmission weight parameter of the third data, and a data transmission weight parameter of the fourth data to the switching module according to the current application scenario. The time slice length parameter is used to represent the total time of a single data transmission, and the data transmission weight parameter is used to represent the time proportion of data transmission.
[0010] The chip system provided by the embodiment of the present application can dynamically adjust the priority of the data stream according to the current application scenario of the terminal device, ensure that the content generation task and other calculation tasks can obtain reasonable resource allocation in different application scenarios, and thus improve the computing performance and response speed of the terminal device.
[0011] In an embodiment of the first aspect of the present application, the AP is specifically configured to: determine the time slice length parameter according to the data access requirement quantities of the first chip and the AIGC chip and the load status of the current application scenario. The time slice length parameter is negatively correlated with the data access requirement quantities of the first chip and the AIGC chip. The higher the load, the smaller the value of the time slice length parameter, and the lower the load, the larger the value of the time slice length parameter.
[0012] In an embodiment of the first aspect of the present application, the AP is specifically configured to: determine the data transmission weight parameter of the third data according to the matching relationship between the current application scenario and the data transmission weight parameter of the preset third data.
[0013] In an embodiment of the first aspect of the present application, the AP is specifically configured to: determine the fourth data transmission weight parameter according to the matching relationship between the current application scenario and the data transmission weight parameter of the preset fourth data.
[0014] In an embodiment of the first aspect of the present application, the switching module includes an arbiter, and the arbiter is configured to: determine the time transmission ratio of the third data and the fourth data according to the proportional relationship between the data transmission weight parameter of the third data and the data transmission weight parameter of the fourth data, determine the time for the second PCIe interface to receive the third data according to the time slice length parameter and the time transmission ratio of the third data, and determine the time for the third PCIe interface to receive the fourth data according to the time slice length parameter and the time transmission ratio of the fourth data.
[0015] The chip system provided by the embodiments of the present application can ensure that high-priority data is preferentially transmitted during business data transmission, improving the response speed and data transmission efficiency of high-priority data.
[0016] In an embodiment of the first aspect of the present application, the switching module further includes a buffer, and the buffer is configured to: buffer the first data, the second data, the third data, and the fourth data.
[0017] In an embodiment of the first aspect of the present application, the AIGC chip and the switching module are integrated on the same chip die.
[0018] The chip system provided by the embodiments of the present application can reduce the packaging complexity, make the integration of the AIGC chip more flexible, and reduce the manufacturing cost of the terminal device.
[0019] In a second aspect, the present application provides a terminal device, including the chip system according to any one of the first aspect.
[0020] Among them, the technical effects of the second aspect refer to the technical effects of the first aspect and any of its embodiments, and will not be repeated here. Description of the Drawings
[0021] Figure 1 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application;
[0022] Figure 2 It is an interaction schematic diagram of a terminal device and a remote service device provided by an embodiment of the present application;
[0023] Figure 3An architecture schematic of a chip system provided by an embodiment of the present application Figure 1 ;
[0024] Figure 4 An architecture schematic of a chip system provided by an embodiment of the present application Figure 2 ;
[0025] Figure 5 An architecture schematic of a chip system provided by an embodiment of the present application Figure 3 ;
[0026] Figure 6 A schematic diagram of time slice configuration provided by an embodiment of the present application;
[0027] Figure 7 An architecture schematic of a chip system provided by an embodiment of the present application Figure 4 ;
[0028] Figure 8 An architecture schematic of a chip system provided by an embodiment of the present application Figure 5 ;
[0029] Figure 9 An architecture schematic of a chip system provided by an embodiment of the present application Figure 6 。 Detailed implementation manners
[0030] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the above", "the", and "this" are also intended to include, for example, the expression "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The character " / " generally indicates an "or" relationship between the associated objects before and after.
[0031] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0032] Hereinafter, the terms "first", "second", etc. are only for convenience of description and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. For example, a plurality of processing units means two or more than two processing units.
[0033] In addition, in the embodiments of the present application, the terms "upper", "lower", "left", and "right" are not limited to the orientations defined relative to the schematic placement of components in the drawings. It should be understood that these directional terms can be relative concepts, which are used for relative description and clarification, and can change accordingly with the change of the orientation of the components in the drawings. In the drawings, for clarity, the thickness of layers and regions is exaggerated, and the dimensional proportional relationships between the various parts in the illustrations do not reflect the actual dimensional proportional relationships.
[0034] In the embodiments of the present application, unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral one; it can be directly connected or indirectly connected through an intermediate medium. In addition, the term "electrical connection" can be a direct electrical connection or an indirect electrical connection through an intermediate medium.
[0035] In the embodiments of the present application, the term "module" is usually a functional structure divided according to logic, and this "module" can be implemented by pure hardware, or by a combination of software and hardware. In the embodiments of the present application, "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, B exists alone, and both A and B exist simultaneously.
[0036] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0037] Currently, when a terminal device executes a content generation task, it first obtains the information input by the user, and this information can include various forms of information such as text, voice, image, or video. The user submits requirements or instructions through an interaction interface on the terminal device, and the terminal device performs preliminary preprocessing on this data, and then transmits the request to a remote service device through the network.
[0038] The remote service device can be located in the cloud or a data center, has powerful computing capabilities and storage resources, can run complex AIGC models, and can implement various content generation tasks such as text generation, image processing, and speech synthesis. After receiving the request sent by the terminal device, the remote service device can call the corresponding model or algorithm according to the request content, perform in-depth calculations and inferences, and generate the required content.
[0039] After completing the content generation task, the remote service device will package, compress the data and return it to the terminal device through a secure network channel. The terminal device then performs necessary post-processing on the returned data, such as formatting, rendering, decoding, etc., and finally presents the generated content to the user. This interaction architecture between the cloud service and the terminal device can not only utilize the powerful computing power of the remote service device to process large-scale data and complex algorithms, but also enable the terminal device to execute high-quality content generation tasks even with limited hardware resources.
[0040] As Figure 1 shown, an embodiment of the present application provides a terminal device 100, which is an electronic device with wireless communication function. The electronic device can be mobile or fixed. The electronic device can be deployed on land (such as indoors or outdoors, handheld or vehicle-mounted, etc.), can also be deployed on water (such as a ship, etc.), and can also be deployed in the air (such as an airplane, a balloon, a satellite, etc.). The electronic device can be referred to as a user equipment (UE), an access terminal, a terminal unit, a subscriber unit, a terminal station, a mobile station (MS), a mobile phone, a terminal agent, or a terminal device, etc. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a smart bracelet, a smart screen, a smart watch, a virtual reality (VR) device, an augmented reality (AR) device, a terminal in industrial control, a terminal in self-driving, a terminal in remote medical, a terminal in smart grid, a terminal in transportation safety, a terminal in smart city, a terminal in smart home, etc. The embodiment of the present application does not limit the specific type and structure of the electronic device. A possible structure of the electronic device will be described below.
[0041] As Figure 1 shown in A, the terminal device 100 may include a front camera 2931 and a display screen 294. As Figure 1 shown in B, the terminal device 100 may include a rear camera 2932. The front camera 2931 and the rear camera 2932 are used to capture static images or dynamic videos (which can be collectively referred to as images). The display screen 294 is used to display images or receive touch operations from the user.
[0042] Figure 2The interactive schematic diagram of the terminal device 100 and the remote service device 200 is shown. After the terminal device 100 obtains the information input by the user, it sends the corresponding request information to the remote service device 200. After receiving the request sent by the terminal device 100, the remote service device 200 can call the corresponding model or algorithm according to the request content, perform in-depth calculation and reasoning, generate the required content and send the processing result to the terminal device 100.
[0043] However, in the above embodiment, there are problems such as high network requirements, poor privacy, and high latency.
[0044] The reason for high network requirements is that continuous data interaction is required between the remote service device and the terminal device, including the sending of task requests, the return of calculation results, and the transmission of necessary post-processing data. The frequency and amount of these data transmissions will increase with the complexity of the content generation task, thus posing higher requirements for network bandwidth and stability. Especially in scenarios such as high-definition image generation and large-scale text processing, the data transmission volume is larger, further exacerbating the dependence on the network.
[0045] The reason for poor privacy is that data needs to be transmitted between the terminal device and the remote service device, that is, the user's input information, task requests, and some intermediate calculation results may pass through external servers, posing risks such as data leakage, unauthorized access, or abuse.
[0046] The reason for high latency lies in that the execution process of the content generation task involves multiple links, including the terminal device sending a request to the remote service device, the remote service device performing calculation processing, the data being packaged and compressed and then transmitted back to the terminal device, and the terminal device performing post-processing and final presentation. Any limitation in network latency, calculation processing time, or data transmission speed in these links will result in the execution latency of the overall task. Especially when the load of the remote service device is high or the network environment is poor, the latency problem is more obvious, affecting the user experience.
[0047] To improve the above problems, the present application provides a chip system that can be deployed in the terminal device 100 in the above embodiment. Refer to Figure 3, the chip system may include an AP 301 and an AIGC chip 302. A content generation model is deployed in the AIGC chip 302. The content generation model deployed in the AIGC chip 302 may include various different types for performing different content generation tasks. As an example, a language model for generating text content may be deployed in the AIGC chip 302, which can automatically generate text according to the input prompt information. A model for generating image content may also be deployed in the AIGC chip 302, enabling intelligent control of various aspects such as image style and details. A model for generating speech synthesis may also be deployed in the AIGC chip 302, capable of converting text into high-quality voice output to meet the voice requirements in multimedia interaction scenarios. A model for video generation or editing may also be deployed in the AIGC chip 302, which can complete the generation or dynamic adjustment of video content to meet the requirements of real-time video applications.
[0048] When performing a content generation task, the AP 301 first sends first data to the AIGC chip 302. The first data may be initial data matching the current content generation task, such as text, voice instructions, images, or video clips, etc. After receiving the first data sent by the AP 301, the AIGC chip 302 will select and call the corresponding content generation model deployed internally according to the type of the first data to perform the corresponding content generation task.
[0049] As an example, in the text-to-image generation task, the AP 301 may send a piece of instruction information text to the AIGC chip 302. After receiving the instruction information text, the AIGC chip 302 will call the internal text-to-image generation model to generate corresponding image data according to the text and send the generated image data back to the AP 301 for display or editing by the terminal device.
[0050] In the text-to-speech synthesis task, the AP 301 may send a piece of instruction information text to the AIGC chip 302. After receiving the instruction information text, the AIGC chip 302 will call the internal speech synthesis model to convert the text into corresponding voice data and return the synthesized voice data to the AP 301, enabling the device to play or store it through the speaker.
[0051] In the text-to-video task, the AP 301 may send a piece of instruction information text to the AIGC chip 302. After receiving the instruction information text, the AIGC chip 302 will call the internal text-to-video model, parse the text into corresponding visual elements, generate video data that conforms to the text, and finally return the video data to the AP 301, enabling the terminal device to play or store it.
[0052] Through the above process, the collaborative work between the AP 301 and the AIGC chip 302 enables the content generation task to be executed efficiently. A variety of content generation models inside the AIGC chip 302 can be flexibly invoked according to the task requirements, achieving high-quality and diverse content generation capabilities, thereby enhancing the intelligence level and user experience of the terminal device.
[0053] The chip system provided by the embodiments of this application can complete content generation tasks on terminal devices, reducing the dependence on remote servers, thereby reducing network requirements. Since data does not need to be frequently uploaded to the cloud for calculation, the user's input data is always processed locally, avoiding the risk of data leakage during transmission and significantly enhancing privacy. At the same time, the AIGC chip 302 can select and invoke the corresponding content generation model according to the first data type sent by the AP 301, enabling various content generation tasks to be directly and efficiently executed locally, thereby avoiding the latency caused by remote computing and data backhaul and improving task execution efficiency and user experience.
[0054] When deploying the AIGC chip 302 in a terminal device, since the number of PCIE interfaces of the AP 301 is limited, it is necessary to reuse the existing PCIE interfaces with unoccupied bandwidth. And there is a large redundancy in the PCIe channels connecting the first chip and the AP 301. Therefore, the connection link between the first chip and the AP 301 can be extended by using the switching module to realize the deployment of the AIGC chip 302.
[0055] Refer to Figure 4 The shown chip system not only includes the AP 301 and the AIGC chip 302, but also includes a switching module 303 and a first chip 304. The switching module 303 includes a first PCIe interface, a second PCIe interface, and a third PCIe interface. The switching module 303 is connected to the AP 301 through the first PCIe interface, the switching module 303 is connected to the AIGC chip 302 through the second PCIe interface, and the switching module 303 is connected to the first chip 304 through the third PCIe interface.
[0056] As an example, the AP 301 can be the SM8750 platform, the first chip 304 can be a WLAN chip, a Bluetooth chip, or other chips with a large PCIe channel redundancy. The SM8750 platform can be connected to the WLAN chip through a group of PCIe Gen3x2 interfaces, and the overall bandwidth of the PCIe Gen3x2 interfaces is 16 Gbps, while the WLAN chip only requires 5.8 Gbps of bandwidth. Therefore, when there is still a large redundancy in the bandwidth utilization of the PCIe interface, a part of the bandwidth can be divided on the same physical channel for the AIGC chip 302 to communicate with the AP 301. In this way, the AIGC chip 302 can share the remaining PCIe bandwidth, thus avoiding the additional addition of an independent high-speed data bus, reducing the complexity of the hardware design, and still providing sufficient bandwidth to support the large data throughput required for AIGC model inference.
[0057] Under the above architecture, the first chip 304 and the AIGC chip 302 can perform corresponding tasks simultaneously without interfering with each other. Specifically, the AP 301 can send the first data and the second data to the switching module 303, and the second data can be the service data corresponding to the task executed by the first chip 304. After receiving the first data and the second data, the switching module 303 can forward the first data to the AIGC chip 302 through the second PCIe interface according to the address identification information in the first data, and forward the second data to the first chip 304 through the third PCIe interface according to the address identification information in the second data.
[0058] After receiving the first data, the AIGC chip 302 performs the corresponding content generation task. After receiving the second data, the first chip 304 performs the corresponding data processing task. After the first chip 304 and the AIGC chip 302 complete the corresponding data processing tasks, they need to send the corresponding data processing results to the AP 301, that is, the AIGC chip 302 needs to send the processing result corresponding to the first data to the AP 301, and the first chip 304 needs to send the processing result corresponding to the second data to the AP 301. Specifically, the AIGC chip 302 needs to send the processing result of the first data (i.e., the third data) to the switching module 303, and the first chip 304 needs to send the processing result of the second data (i.e., the fourth data) to the switching module 303.
[0059] The chip system provided by the embodiments of the present application enables the AIGC chip 302 to be efficiently deployed in a terminal device by introducing a switching module 303 and a first chip 304, and optimizes resources by utilizing the existing PCIe channel redundancy. The switching module 303 can be respectively connected to the AP 301, the AIGC chip 302, and the first chip 304, thereby constructing an efficient data interaction channel. The AP 301 can simultaneously manage the content generation task calculation and other calculation tasks of the first chip 304 through the switching module 303. Therefore, the existing PCIe channel resources can be fully utilized, the additional PCIe interface overhead introduced by independently deploying the AIGC chip 302 can be avoided, and the hardware resource utilization rate of the system can be improved.
[0060] However, in the case where both the first chip 304 and the AIGC chip 302 have data feedback requirements, since the PCIe interfaces through which the first chip 304 and the AIGC chip 302 are connected to the switching module 303, and the PCIe interface through which the switching module 303 is connected to the AP 301 have the same speed, it is necessary to determine the time length for the switching module 303 to receive the third data and the fourth data.
[0061] In a feasible embodiment of the present application, refer to Figure 3 , the AP 301 is configured to: obtain the current application scenario of the terminal device, and send to the switching module 303 the time slice length parameter and data transmission weight parameter of the third data, and the time slice length parameter and data transmission weight parameter of the fourth data according to the current application scenario. The time slice length parameter is used to represent the total time of a single data transmission, and the data transmission weight parameter is used to represent the time proportion of data transmission.
[0062] When determining the current application scenario of the terminal device, the application scenario can be obtained by monitoring the currently running applications on the terminal device and their foreground activity status. As an example, the application process information of the system can be read, the types of currently active applications can be analyzed, such as social media, video playback, games, or office software, etc., and combined with the user's interaction behaviors, such as touch operations, input methods, screen brightness adjustment, etc., to further determine the specific application scenario. For example, when it is detected that the user is using a video playback software and the screen remains in landscape mode, the current application scenario can be determined as the video viewing scenario; when it is detected that the user frequently types and the foreground application is a document editing software, the current application scenario can be determined as the office scenario; when it is detected that there is high-frame-rate rendering and high-frequency touch operations and the current application is a game process, the current application scenario can be determined as the game scenario.
[0063] After determining the current application scenario of the terminal device, the time slice length parameter, the data transmission weight parameter of the third data, and the data transmission weight parameter of the fourth data can be determined according to the current application scenario.
[0064] The time slice length parameter refers to the total time length for receiving the third data and the fourth data. That is, within a scheduling period, the total transmission duration of the third data and the fourth data is restricted by the time slice length parameter, and the data transmission weight parameter determines the time proportion occupied by the third data and the fourth data respectively within this time slice.
[0065] The time slice length parameter is used to characterize the total time of a single data transmission. That is, within a scheduling period, it represents the maximum duration allowed for data transmission. The time slice length parameter determines the overall execution duration of the data transmission task within this time range and directly affects the data throughput rate, transmission delay, and the utilization rate of system resources. For example, in bus data scheduling, the time slice length parameter can be set to a fixed time window, such as 5 ms, 10 ms, etc., indicating the upper limit of the duration of each data transmission cycle. If the time slice length is short, the frequency of data transmission may increase, but the amount of data transmitted each time may be limited; if the time slice length is long, the amount of data transmitted in a single transmission increases, but it may lead to a decrease in scheduling flexibility.
[0066] The data transmission weight parameter is used to characterize the time proportion of data transmission. That is, within a time slice cycle, it represents the time proportion occupied by data transmission. This parameter is usually used in multi-task scheduling or resource allocation scenarios to determine the priority and bandwidth allocation of different data streams or tasks in the total time slice. For example, within a scheduling period with a time slice length of 10 ms, if the transmission weight parameter of data stream A is set to 6 and the transmission weight parameter of data stream B is set to 4, then the transmission time of data stream A is 6 / (6 + 4), that is, data stream A occupies 6 ms for data transmission within this cycle, and the transmission time of data stream B is 4 / (6 + 4), that is, data stream B occupies 4 ms for data transmission within this cycle.
[0067] When determining the time slice length parameter, the AP 301 can determine the time slice length parameter according to the load status of the current application scenario and the number of data access requirements for the first chip and the AIGC chip.
[0068] The time slice length parameter is negatively correlated with the data access requirement quantities of the first chip and the AIGC chip. The data access requirement quantity refers to the demand status of the current application scenario for the first chip and the AIGC chip. As an example, if there are data interaction requirements for both the first chip and the AIGC chip in the current application scenario, the value of the time slice length parameter is small. If there is a data interaction requirement only for one of the first chip or the AIGC chip in the current application scenario, the value of the time slice length parameter is large.
[0069] Furthermore, in the case where there are data interaction requirements for both the first chip and the AIGC chip in the previous application scenario, the time slice length parameter can be further adjusted according to the load status of the previous application scenario. That is, the load is positively correlated with the value of the time slice length parameter. That is, the higher the load, the smaller the value of the time slice length parameter, and the lower the load, the larger the value of the time slice length parameter. The load can be measured by indicators such as CPU occupancy rate, memory usage, I / O bandwidth occupancy, network traffic, and GPU computing load. In a game scenario, hardware resources such as the CPU and GPU are highly occupied, and the load is large, which will shorten the time slice length parameter, that is, reduce the time of a single data transmission cycle, and provide the real-time performance of data transmission. Therefore, in a game scenario, the time slice length parameter can be 10 ms.
[0070] In the scenario of executing a content generation task, the content generation task has relatively high requirements for the continuous computing power of the system, but its real-time requirement is relatively low. Therefore, when executing a content generation task, the overall load of the system is relatively low. Especially when the content generation task mainly relies on the independent processing of the AIGC chip 302, the resource occupancy of the AP 301 is less, making the system in a relatively low load during the task execution. In this case, the time slice length parameter will be increased, that is, the time slice of data transmission is extended, so that data can be continuously transmitted within a longer time window. This adjustment can reduce frequent task switching, improve data throughput efficiency, and reduce the overhead of system resource scheduling. Therefore, in the scenario of executing a content generation task, the time slice length parameter can be 50 ms.
[0071] When determining the data transmission weight parameters of the third data and the fourth data, the AP 301 can determine them according to the matching relationship between the current application scenario and the preset data transmission weight parameters. As an example, when the application scenario is a gaming scenario, the demand for the WLAN chip is relatively high. Therefore, a larger weight is assigned to the data sent by the WLAN chip, that is, the data transmission weight of the fourth data is relatively high, while the demand for the AIGC chip 302 is relatively low. Therefore, the data transmission weight of the third data is relatively low to ensure the smoothness and low-latency response of the game. When the application scenario is a scenario for performing content generation tasks, the demand for the AIGC chip 302 is relatively high. Therefore, a larger weight is assigned to the data sent by the AIGC chip 302, that is, the data transmission weight of the third data is relatively high, while the demand for the WLAN chip is relatively small. Therefore, the data transmission weight of the fourth data is relatively low to ensure the smooth execution of the content generation task.
[0072] The matching relationships between the application scenario and the data transmission weight of the third data, and between the application scenario and the data transmission weight of the fourth data can be stored in the buffer 3032 of the AP 301 in advance. After determining the application scenario, the data transmission weight of the third data and the data transmission weight of the fourth data corresponding to the current application scenario can be found based on the application scenario as the retrieval basis.
[0073] After determining the time slice length parameter and the data transmission weight parameters of each data, the time when the second PCIe interface receives the third data and the time when the third PCIe interface receives the fourth data can be determined.
[0074] The time slice length parameter and the data transmission weight can be dynamically adjusted according to the switching of the scenario. As an example, refer to Figure 6 When the current application scenario of the terminal device is a video-watching scenario, the time slice length parameter can be 15 ms, and the weight ratio of the data sent by the WLAN chip and the data sent by the AIGC chip 302 is 7:3. When the application scenario switches to a gaming scenario, the time slice length parameter can be 10 ms, and the weight ratio of the data sent by the WLAN chip and the data sent by the AIGC chip 302 is 9:1. When the application scenario switches to a scenario for performing content generation tasks, the time slice length parameter can be 40 ms, and the weight ratio of the data sent by the WLAN chip and the data sent by the AIGC chip 302 is 2:8.
[0075] In a feasible embodiment of the present application, continue to refer to Figure 5, the switching module 303 includes an arbiter 3031, and the arbiter 3031 is configured to: determine the time transmission occupancy ratio of the third data and the fourth data according to the proportional relationship between the data transmission weight parameter of the third data and the data transmission weight parameter of the fourth data, determine the time when the second PCIe interface receives the third data according to the time slice length parameter and the time transmission occupancy ratio of the third data, and determine the time when the third PCIe interface receives the fourth data according to the time slice length parameter and the time transmission occupancy ratio of the fourth data.
[0076] As an example, when the application scenario is a game scenario, the time slice length parameter is 10 ms, the data transmission weight parameter of the third data is 1, and the data transmission weight parameter of the fourth data is 9. Then the time transmission occupancy ratio of the third data and the fourth data is 1:9. Therefore, the time when the second PCIe interface receives the third data can be 1 ms, and the time when the third PCIe interface receives the fourth data can be 9 ms. And because the data transmission weight parameter of the fourth data is higher, the moment when the third PCIe interface receives the fourth data is earlier than the moment when the second PCIe interface receives the third data. That is, the third PCIe interface receives the fourth data in the first 9 ms of a time slice, and the second PCIe interface receives the third data in the last 1 ms of a time slice.
[0077] The chip system provided by the embodiments of the present application can ensure that high-priority data is preferentially transmitted during business data transmission, thereby ensuring fairness and efficiency between different data streams, and improving the response speed and data transmission efficiency of the system.
[0078] In a feasible embodiment of the present application, continue to refer to Figure 5 , the switching module 303 further includes a buffer 3032. The buffer 3032 can be coupled to the arbiter 3031, and the buffer 3032 is configured to: buffer the first data, the second data, the third data, and the fourth data.
[0079] The function of the buffer 3032 is to provide a temporary storage area for storing the first data, the second data, the third data, and the fourth data. The main purpose of setting the buffer is to reduce data loss or delay during data transmission and improve data transmission efficiency. Through the buffer 3032, the data sent by the AP 301, the first chip 304, or the AIGC chip 302 can be temporarily stored, thus avoiding loss or excessive waiting time caused by the receiving end being unable to process the data in time.
[0080] In the above embodiment, the AIGC chip 302 and the switching module 303 are independently arranged. Although this architecture can provide a certain degree of flexibility, there is still room for optimization in terms of chip area, packaging complexity, and data transmission efficiency. As an example, refer to Figure 7, the interaction module includes three external PCIe interfaces, and each PCIe interface contains a corresponding controller and a physical interface.
[0081] To further reduce the chip area and packaging cost. In a feasible embodiment of the present application, a chip system is provided. Refer to Figure 8 , the AIGC module 305 and the switching module 303 in the chip system can be arranged on an integrated chip 300. The AIGC module 305 has the same function as the aforementioned AIGC chip 302. The integrated chip 300 is connected to the AP 301 through the first PCIe interface and connected to the first chip 304 through the second PCIe interface. The AIGC module 305 and the switching module 303 are connected through an embedded PCIe interface. The embedded PCIe interface has the characteristic of low latency, enabling the AIGC module 305 to quickly receive data from the switching module 303 and at the same time efficiently transmit the processed results back to the switching module 303 for further data scheduling and distribution. Arranging the AIGC module 305 and the switching module 303 on the integrated chip 300 can reduce the packaging cost of the chip.
[0082] By integrating the AIGC module 305 and the switching module 303 on the same chip die, the data transmission latency between the AIGC module 305 and the switching module 303 can be effectively reduced, the data throughput rate can be improved, and at the same time, the power consumption overhead caused by cross-chip communication can be reduced. In addition, at the physical packaging level, this integration method can reduce the number of pins required for packaging, thereby reducing the overall manufacturing cost and packaging difficulty. At the same time, since the switching module 303 and the AIGC module 305 share the same die, data interaction can be carried out through on-chip interconnection without passing through an external bus or an additional physical interface, thereby further improving the real-time performance and stability of data exchange.
[0083] Specifically, refer to Figure 9 The provided chip system, compared with Figure 7For the chip system shown, the AIGC module 305 and the switching module 303 are connected through an embedded PCIe interface. After adopting the embedded PCIe interface, the AIGC module 305 and the switching module 303 do not need to arrange multiple physical interfaces, but only retain the corresponding controllers, thus significantly reducing the package area and making the entire chip system more compact and integrated. In addition, reducing the use of physical interfaces also reduces the power consumption of the chip system. Since the signal transmission of each physical interface is accompanied by a certain amount of power consumption, especially during high-speed signal transmission, this power consumption will be more obvious. However, through data interaction using the embedded PCIe interface, the additional power consumption generated by the simultaneous operation of multiple physical interfaces can be effectively reduced, thereby optimizing the energy efficiency performance of the entire chip system. In addition, reducing the number of physical interfaces can also simplify PCB wiring, reduce production difficulty and cost, and at the same time improve signal integrity and data transmission stability.
[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0085] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0086] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0087] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0088] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods provided in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0089] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A chip system, characterized in that, Comprising: An application processor AP and an artificial intelligence generated content AIGC chip, with a content generation model deployed inside the AIGC chip; The AP is configured to: send first data to the AIGC chip; The AIGC chip is configured to: perform a corresponding content generation task according to the first data and the content generation model; The chip system further includes a switching module and a first chip, and the switching module includes a first PCIe interface, a second PCIe interface, and a third PCIe interface; The switching module is connected to the AP through the first PCIe interface, the switching module is connected to the AIGC chip through the second PCIe interface, and the switching module is connected to the first chip through the third PCIe interface; The AP is configured to: Send first data and second data to the switching module; The switching module is configured to: Forward the first data to the AIGC chip and forward the second data to the first chip; Or, receive third data sent by the AIGC chip and fourth data sent by the first chip and send them to the AP; The AP is configured to: Obtain the current application scenario of the terminal device; According to the current application scenario, send a time slice length parameter, a data transmission weight parameter of the third data, and a data transmission weight parameter of the fourth data to the switching module, where the time slice length parameter is used to represent the total time of a single data transmission, and the data transmission weight parameter is used to represent the time proportion of data transmission.
2. The chip system according to claim 1, wherein The AP is specifically configured to: Determine the time slice length parameter according to the data access requirement quantities of the first chip and the AIGC chip in the current application scenario and the load status of the current application scenario. The time slice length parameter is negatively correlated with the data access requirement quantities of the first chip and the AIGC chip. The higher the load, the smaller the value of the time slice length parameter, and the lower the load, the larger the value of the time slice length parameter.
3. The chip system according to claim 1, characterized in that, The AP is specifically configured to: Determine the data transmission weight parameter of the third data according to the matching relationship between the current application scenario and the preset data transmission weight parameter of the third data.
4. The chip system according to claim 1, wherein The AP is specifically configured to: Determine the data transmission weight parameter of the fourth data according to the matching relationship between the current application scenario and the preset data transmission weight parameter of the fourth data.
5. The chip system according to claim 1, characterized in that, The switching module includes an arbiter, and the arbiter is configured to: Determine the time transmission proportion of the third data and the fourth data according to the proportional relationship between the data transmission weight parameter of the third data and the data transmission weight parameter of the fourth data; Determine the time when the second PCIe interface receives the third data according to the time slice length parameter and the time transmission proportion of the third data; Determine the time when the third PCIe interface receives the fourth data according to the time slice length parameter and the time transmission proportion of the fourth data.
6. The chip system according to claim 1, wherein The switching module further includes a buffer, and the buffer is configured to: Cache the first data, the second data, the third data, and the fourth data.
7. The chip system according to claim 1, wherein The AIGC chip and the switching module are integrated on the same chip die.
8. The chip system according to any one of claims 1-7, characterized in that, The first chip includes a WALN chip.
9. A terminal device, characterized in that, A chip system according to any one of claims 1-8 is included.
Citation Information
Patent Citations
Using quantization in training an artificial intelligence model in a semiconductor solution
US20200320385A1
Display device for generating multimedia content, and operation method of the display device
US20210390314A1