Cloud mobile phone terminal decoding chip architecture, collaborative optimization method thereof, processing equipment and storage medium

Through the integration and optimization design of ASIC chips, low-power video decoding and touch signal processing of cloud mobile terminal devices are realized, solving the problems of high power consumption and insufficient collaborative optimization, significantly extending battery life and improving user experience.

CN120264005APending Publication Date: 2025-07-04启朔(深圳)科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510410235.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing cloud mobile terminal devices consume high power when handling complex computing tasks, which affects the device's battery life and user experience. The existing low-power chip architecture lacks the coordinated optimization of video decoding and touch signal processing, which cannot meet the needs of long-term use.

Method used

It adopts ASIC chip, integrates video decoding module, touch processing module, dynamic power management module and communication protocol switching module, and realizes low-power processing of video decoding and touch signal through module cascade optimization, including parallel processing units of video decoding module and Kalman filter of touch processing module, and dynamic power gate of dynamic power management module.

Benefits of technology

Significantly reduce power consumption to 1/5 of ordinary mobile phones, extend device battery life, improve video decoding efficiency and touch response speed, reduce costs, and meet the needs of low-cost and high-performance cloud mobile terminal devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264005A_ABST
    Figure CN120264005A_ABST
Patent Text Reader

Abstract

The invention relates to a cloud mobile phone terminal decoding chip architecture and a collaborative optimization method thereof, processing equipment and a storage medium, the chip architecture adopts an ASIC (Application Specific Integrated Circuit) chip and comprises a video decoding module, a communication interface module, a touch processing module, a dynamic power management module and a communication protocol switching module; the video decoding module is used for performing motion compensation processing and entropy decoding on input video data and determining a power gating driving signal of the dynamic power management module; the touch control processing module is used for carrying out wavelet processing and Kalman filtering processing on an input touch control signal; the dynamic power management module is used for dynamically adjusting the power supply voltage and frequency according to the power consumption information and cutting off power supply according to the power gating driving signal and / or the detected temperature; the communication protocol switching module is used for controlling switching of communication protocols; the communication interface module is used for determining power consumption information; and sending the decoded video data and the filtered touch signal to the cloud. The method and the device can be widely applied to the technical field of cloud mobile phones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud mobile phones, and particularly to a decoding chip architecture for a cloud mobile phone terminal, a collaborative optimization method thereof, a processing device, and a storage medium. Background Art

[0002] In recent years, the mobile Internet has developed vigorously, the performance of smart phones has been continuously improved, and the application scenarios have become increasingly rich. However, due to factors such as volume and power consumption, smart phones still have bottlenecks in computing power, storage space, battery life, etc., and it is difficult to meet the growing application requirements of users for high-definition videos, large-scale games, virtual reality, etc. The maturity and popularization of cloud computing technology provide a new idea for breaking through the terminal bottleneck. Cloud computing centralizes resources such as computing and storage in the cloud and provides services through the network. Users can enjoy powerful computing capabilities without local high-performance devices. A cloud mobile phone is a product of applying cloud computing technology to the mobile terminal field. It runs the operating system and application programs of a smart phone on a cloud server, and users remotely access and control through the network to achieve the same functions as a local smart phone. With the development of cloud mobile phone technology, users have higher and higher requirements for the performance and power consumption of cloud mobile phone terminal devices.

[0003] However, existing cloud mobile phone terminal devices usually need to process complex computing tasks, resulting in high power consumption, which affects the battery life and user experience of the devices. Most cloud mobile phone terminal devices on the market currently adopt a general chip architecture. Although the functions are relatively comprehensive, there are deficiencies in power consumption control and it is difficult to meet the needs of long-term use. In addition, existing low-power chip architectures mostly focus on the optimization of specific functions, such as separate video decoding or touch signal processing, lacking the collaborative optimization of video decoding and touch signal processing, and unable to fully utilize the advantages of low power consumption. Summary of the Invention

[0004] Aiming at the above problems, the purpose of the present invention is to provide a decoding chip architecture for a cloud mobile phone terminal, a collaborative optimization method thereof, a processing device, and a storage medium, which can achieve low-power processing of video decoding and touch signal uploading.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, a decoding chip architecture for a cloud mobile phone terminal is provided, which uses an ASIC chip and includes a video decoding module, a communication interface module, a touch processing module, a dynamic power management module, and a communication protocol switching module;

[0006] The video decoding module is used to perform motion compensation processing and entropy decoding on the input video data to generate decoded video data and output it to the communication interface module, and to detect the frame processing state of the video data and determine the power gating drive signal of the dynamic power management module and output it to the dynamic power management module;

[0007] The touch processing module is used to perform wavelet processing and Kalman filtering on the input touch signal, and output the processed touch signal to the communication interface module;

[0008] The dynamic power management module is used to dynamically adjust the supply voltage and frequency of the video decoding module, touch processing module and communication interface module according to the power consumption information, detect the temperature of the ASIC chip, and cut off the power supply according to the power gating drive signal and / or the detected temperature;

[0009] The communication protocol switching module is used to control the switching of communication protocols based on the bandwidth requirement;

[0010] The communication interface module is used to determine the power consumption information based on the decoded video data and the processed touch signal, and output it to the dynamic power management module; and send the decoded video data and the filtered touch signal to the cloud based on the current communication protocol.

[0011] Further, the video decoding module includes a plurality of parallel processing units, and each parallel processing unit includes:

[0012] A pipeline control logic unit, which is used to determine syntax elements based on the compressed bitstream of video data, and detect the frame processing state of video data, including frame start flag, frame end flag and frame gap;

[0013] A motion compensation unit, which is used to perform motion compensation processing on video data based on syntax elements to generate a predicted frame of video data;

[0014] An entropy decoding unit, which is used to perform entropy decoding on the predicted frame of video data to generate decoded video data;

[0015] A first output unit, which outputs the decoded video data to the communication interface module;

[0016] A local register, which is used to store the decoded video data.

[0017] Further, each pipeline control logic unit adopts a four-stage pipeline control logic, where:

[0018] The first stage of the pipeline is binary conversion, which is used to convert the compressed bitstream of the input video data into binary symbols;

[0019] The second stage of the pipeline is context model selection, which is used to dynamically calculate the context model index and select the corresponding probability model for each binary symbol;

[0020] The third stage of the pipeline is probability state update, which is used to update the probability state of the context model in real time according to the decoded symbols;

[0021] The fourth - stage pipeline is for bit - stream generation, which is used to recombine the decoded binary symbols into syntax elements and output them to the motion compensation unit through the Crossbar architecture.

[0022] Furthermore, the touch - processing module includes:

[0023] A wavelet - transform unit, which is used to receive the input touch signal and perform noise reduction and feature extraction;

[0024] A Kalman filter, which is used to perform Kalman filtering on the touch signal after noise reduction and feature extraction to obtain the processed touch signal;

[0025] A second output unit, which outputs the filtered touch signal to the communication interface module.

[0026] Furthermore, the Kalman filter is configured with:

[0027] A programmable window register group, which is used to determine the window size of Kalman filtering based on the real - time data input by the touch - speed sensor according to the dynamic adjustment logic, update the register value according to the window size; and provide the size and weight parameters of the sliding window for Kalman filtering;

[0028] A parallel - processing channel, which is used to perform Kalman filtering on the touch signal after noise reduction and feature extraction based on the parameters provided by the programmable window register group to obtain the processed touch signal.

[0029] Furthermore, the dynamic power - management module includes:

[0030] An annular oscillator array, which is used to generate a reference clock signal for full - chip clock synchronization;

[0031] An SAR - type digital calibration circuit, which is used to calibrate the power - gating turn - off or power - restoration timing of dynamic power gating based on the power - gating drive signal to obtain the calibrated power - gating drive signal;

[0032] An over - temperature protection circuit, which is used to detect the temperature of the ASIC chip;

[0033] The dynamic power gating is used to dynamically adjust the supply voltage and frequency of the video - decoding module, touch - processing module and communication interface module according to the power - consumption information, and cut off the power supply based on the calibrated power - gating drive signal and / or the detected temperature.

[0034] In a second aspect, a collaborative optimization method based on a cloud - mobile - terminal decoding - chip architecture is provided, including:

[0035] Detect the frame processing status of video data, and output a power gating drive signal when the frame gap in the frame processing status exceeds a preset frame gap threshold;

[0036] Receive the input touch signal and perform wavelet processing and Kalman filtering to obtain the processed touch signal;

[0037] Control the switching of communication protocols based on bandwidth requirements;

[0038] Determine power consumption information based on the decoded video data and the processed touch signal, and send the decoded video data and the filtered touch signal to the cloud based on the current communication protocol;

[0039] Dynamically adjust the supply voltage and frequency according to the power consumption information to control power consumption;

[0040] Detect the temperature of the ASIC chip and cut off the power supply according to the power gating drive signal and / or the detected temperature.

[0041] In a third aspect, a cloud mobile phone terminal device is provided, including the above-mentioned cloud mobile phone terminal decoding chip architecture.

[0042] In a fourth aspect, a processing device is provided, including computer program instructions, wherein when the computer program instructions are executed by the processing device, they are used to implement the steps corresponding to the above-mentioned collaborative optimization method.

[0043] In a fifth aspect, a computer-readable storage medium is provided, on which computer program instructions are stored, wherein when the computer program instructions are executed by a processor, they are used to implement the steps corresponding to the above-mentioned collaborative optimization method.

[0044] Due to the above technical solutions adopted by the present invention, it has the following advantages:

[0045] 1. Through the streamlined design and optimized processing of the dedicated ASIC chip, the present invention can significantly reduce power consumption, reduce the power consumption to 1 / 5 of that of ordinary mobile phones, effectively extend the battery life of the device, and improve the user experience. The typical power consumption of the present invention based on 28nm process simulation is 0.5W@600MHz, the cut-off delay ≤0.8ms in the worst case (SS process corner), and the setup time margin ≥0.15ns shown by PrimeTime verification.

[0046] 2. The present invention adopts an advanced video decoding algorithm and hardware acceleration method, which can improve the efficiency and quality of video decoding, ensure that users can smoothly watch high-definition videos, and the system-level energy efficiency ratio is improved by 39.4% compared with the existing scheme.

[0047] 3. The low-power touch processing module provided by the present invention can quickly and accurately capture the user's touch operations, provide more sensitive touch responses, and enhance the user's operation experience. The false alarm rate of the touch processing module is verified by simulation to be ≤2%.

[0048] 4. The video decoding module of the present invention adopts a setting of 128 parallel processing units, which improves the throughput through large-scale parallelization and avoids the bottleneck of traditional serial processing. The total occupied area of the 128 parallel processing units is only 2.3 mm 2 (including wiring), which can meet the requirements of low-cost chips.

[0049] 5. In the video decoding module of the present invention, the pipeline control logic unit of each parallel processing unit adopts a four-stage pipeline control logic. Through precise pipeline cutting and context model update delay control, it can solve the pipeline blocking problem caused by dependency relationships during the entropy decoding process, thereby significantly improving the throughput and reducing the power consumption, ensuring the efficient connection of the decoding process inside each macroblock processing unit of the parallel processing unit.

[0050] 6. In the touch processing module of the present invention, the Kalman filter decouples the dynamic configuration of window parameters from the filtering calculation, and realizes zero-delay parameter switching through a programmable window register group, solving the performance bottleneck caused by software configuration in the traditional solution.

[0051] 7. The 32-step accuracy of the SAR (Successive Approximation Register) type digital calibration circuit in the dynamic power management module of the present invention can ensure the stability of the dynamic power gating timing, avoid timeouts caused by process deviations or temperature changes. The design of dynamic power gating in the dynamic power management module can achieve a microsecond-level response speed, complete power supply cut-off within 0.5 ms, and when the power supply is cut off, the static power consumption approaches zero, which can meet the battery life requirements of cloud mobile phone terminals.

[0052] 8. The present invention is applicable to cloud mobile phone dedicated terminals at the hundred-yuan level, which can reduce the manufacturing cost of the device and has high market competitiveness.

[0053] In summary, the present invention can be widely applied in the technical field of cloud mobile phones. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0055] Figure 1 is a schematic structural diagram of the cloud mobile phone terminal decoding chip architecture provided by an embodiment of the present invention;

[0056] Figure 2 It is a schematic flowchart of a method provided by an embodiment of the present invention. Detailed implementation manners

[0057] Hereinafter, the exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0058] It should be understood that the terms used herein are for the purpose of describing specific exemplary embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0059] Although the terms first, second, third, etc. may be used herein to describe multiple elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, or section from another. Unless the context clearly indicates otherwise, terms such as "first", "second", and other numerical terms used herein do not imply an order or sequence. Thus, the first element, component, region, layer, or section discussed below may be referred to as the second element, component, region, layer, or section without departing from the teachings of the exemplary embodiments.

[0060] At present, most cloud phone terminal devices on the market adopt a general chip architecture. Although they have comprehensive functions, they have deficiencies in power consumption control and are difficult to meet the needs of long-term use. In addition, existing low-power chip architectures mostly focus on the optimization of specific functions, such as separate video decoding or touch signal processing, lacking the collaborative optimization of video decoding and touch signal processing, and unable to fully utilize the advantages of low power consumption. Through the streamlined design and optimization of the ASIC (Application-Specific Integrated Circuit) chip, this invention removes unnecessary functional modules, focuses on video decoding and touch signal processing, integrates an efficient video decoding module, a touch processing module, and a low-power communication interface module inside the chip, which can ensure the rapid transmission and processing of data; through module-level cascade optimization, it increases the collaborative working mechanism of the "decoding-touch-communication" three modules. Therefore, this invention can achieve low-power processing of video decoding and touch signal uploading. At the same time, this invention also has a high decoding efficiency and touch response speed.

[0061] Embodiment 1

[0062] As Figure 1 shown, this embodiment provides a cloud phone terminal decoding chip architecture that uses an ASIC chip and includes a video decoding module, a communication interface module, a touch processing module, a dynamic power management (DPM) module, and a communication protocol switching module.

[0063] The video decoding module is used to perform motion compensation processing and entropy decoding on the input video data to generate decoded video data and output it to the communication interface module, and to detect the frame processing status of the video data and determine the power gating drive signal of the dynamic power management module and output it to the dynamic power management module.

[0064] The touch processing module is used to perform wavelet processing and Kalman filtering on the input touch signal to obtain the processed touch signal and output it to the communication interface module.

[0065] The dynamic power management module is used to dynamically adjust the supply voltage and frequency of the video decoding module, the touch processing module, and the communication interface module according to the power consumption information, control the power consumption of the video decoding module, the touch processing module, and the communication interface module, detect the temperature of the ASIC chip, and cut off the power supply according to the power gating drive signal and / or the detected temperature.

[0066] The communication protocol switching module is used to control the switching of the communication protocol based on the bandwidth requirement, where SNR is the current signal-to-noise ratio.

[0067] The communication interface module is used to determine power consumption information based on the decoded video data and the processed touch signals, and output it to the dynamic power management module; and send the decoded video data and the filtered touch signals to the cloud based on the current communication protocol.

[0068] Through the streamlined design and optimization of the dedicated ASIC chip, and by adopting advanced video decoding algorithms and hardware acceleration methods, the present invention can achieve the technical effects of low power consumption, high decoding efficiency, fast touch response, and low cost, has significant commercial value, can meet the market demand for low-cost and high-performance cloud mobile phone terminal devices, and has broad market prospects.

[0069] In a preferred embodiment, the video decoding module adopts the H.265 decoding algorithm and improves the decoding efficiency through a hardware acceleration method. During the decoding process, the power consumption is further reduced by lowering the clock frequency of the video decoding module. The processing time of a single-frame video data by the video decoding module is 4.17 ns, and the theoretical maximum throughput is 240 Mfps, meeting the 4K@60fps requirement. Therefore, the video decoding module adopts 128 parallel processing units, and each parallel processing unit includes a pipeline control logic unit, a motion compensation unit, an entropy decoding unit, a first output unit, and a local register. The pipeline control logic unit is used to determine syntax elements based on the compressed bitstream of the video data for reconstructing pixel blocks; and detect the frame processing status of the video data, including frame start flag, frame end flag, and frame gap. The motion compensation unit is used to perform motion compensation processing on the video data based on the syntax elements to generate a predicted frame of the video data, i.e., motion vector data. The entropy decoding unit is used to perform entropy decoding on the predicted frame of the video data to generate the decoded video data and send it to the first output unit. The first output unit is used to output the decoded video data to the communication interface module, where the bit width of the first output unit is adaptively adjusted according to the size of the syntax elements. The local register is used to store the decoded video data to reduce the access latency to the global memory.

[0070] Specifically, each parallel processing unit adopts the H.265 standard decoding of a dedicated 16×16 macroblock processing unit; the local register file is 64×32bit.

[0071] Specifically, each parallel processing unit shares the context model memory through a Crossbar interconnect architecture. Based on the ARMSC12 library, the total occupied area of 128 parallel processing units is 2.3 mm 2 (including wiring), adopts a hierarchical Crossbar structure, and every 16 parallel processing units form a sub-cluster, and the sub-clusters are connected through a secondary Crossbar.

[0072] Specifically, the pipeline control logic unit of each parallel processing unit adopts a four-stage pipeline control logic. Among them, the first stage of the pipeline is Binarization, the second stage is Context Modeling, the third stage is Probability Update, and the fourth stage is Bitstream Generation. The context model update delay is controlled within 3 clock cycles, supporting throughput metrics of 8K@30fps / 4K@60fps. More specifically, the binarization pipeline is used to select corresponding binarization rules according to the syntax elements defined by the H.265 decoding algorithm standard (such as motion vectors, residual coefficients), convert the compressed bitstream of the input video data into binary symbols; and preload the context model index to reduce latency for the next stage. The context model selection pipeline is used to dynamically calculate the context model index based on the syntax element type, adjacent block information, and decoded symbol history, and then select the corresponding probability model for each binary symbol. The probability state update pipeline is used to update the probability state of the context model in real time according to the decoded symbols. The update formula is: new probability = old probability + α × (decoded symbol - predicted value of old probability), new probability = old probability + α × (decoded symbol - predicted value of old probability), where α is the adaptive step size (hardware-fixed calculation). The bitstream generation pipeline is used to recombine the decoded binary symbols into syntax elements according to the H.265 syntax rules and output them to the motion compensation unit through the Crossbar architecture. At the same time, the pipeline control logic unit of the present invention synchronizes the entropy decoding unit and the motion compensation unit through hardware semaphores to avoid data waiting.

[0073] Specifically, a counter array is also integrated in the pipeline control logic unit, including a frame end detector, a high-precision timer, and a threshold comparator. The frame end detector is used to monitor the end signals of the entropy decoding unit and the motion compensation unit, mark the completion of the current frame decoding, and obtain the frame end flag. The high-precision timer is used to calculate the frame gap between two video frames based on the reference clock signal provided by the dynamic power management module. The threshold comparator is used to send a power gating drive signal to trigger the dynamic power management module when the frame gap exceeds a preset frame gap threshold (such as 5 ms).

[0074] Specifically, the comparison of the effects of the video decoding module of the present invention and a general-purpose chip (such as Snapdragon 480) is shown in Table 1 below:

[0075] Table 1: Comparison Table of the Effects of the Video Decoding Module and the General-Purpose Chip

[0076] Indicator The solution of the present invention General-purpose chip (such as Snapdragon 480) Parallelism 128-unit spatial parallelism + pipeline Multi-core CPU / GPU, relying on instruction-level parallelism Energy efficiency ratio 509.6 GOPS / W 42.1 GOPS / W Latency control Context update ≤ 3 cycles Relying on software scheduling, with high latency Area efficiency <![CDATA[2.3mm 2 (128 cells)]]> Large area overhead for multi-core design

[0077] In a preferred embodiment, the touch processing module adopts a low-power sensor interface and signal processing algorithm, which can accurately capture the user's touch signal and convert it into a digital signal for uploading to the cloud through the communication interface module. The touch processing module also has the functions of automatic sleep and wake-up, and can dynamically adjust the power consumption according to the user's operation state. An anti-interference design is added. The environmental noise suppression adopts wavelet transform combined with an adaptive threshold algorithm, achieving a signal-to-noise ratio improvement of more than 40 dB in the 100 - 500 kHz interference frequency band. Multi-touch supports 10-channel parallel processing, and a time-division multiplexing mechanism is adopted between channels to reduce cross-interference. The quantization relationship between the touch sampling rate and power consumption is, for example: when the sampling rate is 200 Hz, the power consumption ≤ 5 mW. Therefore, the touch processing module adopts a three-stage Kalman filtering algorithm, including a wavelet transform unit, a Kalman filter, and a second output unit. The wavelet transform unit is used to receive the touch signal input by the touch speed sensor, perform noise reduction and feature extraction, and then send it to the Kalman filter. The Kalman filter is used to perform Kalman filtering on the touch signal after noise reduction and feature extraction to obtain the processed touch signal, and send it to the second output unit. The second output unit outputs the processed touch signal to the communication interface module.

[0078] Specifically, the wavelet transform unit adopts wavelet decomposition based on the Mallat algorithm. Five-layer decomposition can cover the 100 - 500 kHz interference frequency band, and the signal-to-noise ratio improvement ΔSNR = 20log 10 (N / σ), where N is the signal amplitude and σ is the noise standard deviation. The σ is reduced to 1 / 100 of the original value through an adaptive threshold.

[0079] Specifically, the Kalman filter is configured with a programmable window register group and parallel processing channels. The programmable window register group is used to determine the window size of the Kalman filter based on the real-time data input by the touch speed sensor (such as the touch movement speed), based on the dynamic adjustment logic, and update the register value according to the window size. For example, when the touch is slow (such as 15 sampling points), the window is enlarged to improve smoothness; when the touch is fast (such as 5 sampling points), the window is reduced to reduce latency; and it provides the size and weight parameters of the sliding window for the Kalman filter. For example, the size of the sliding window defines the number of sampling points for each filtering process, and the weight parameter of the sliding window is the contribution ratio of different sampling points in the filtering calculation (such as exponential decay weight). By updating the window parameters in real time, the three-stage Kalman filtering algorithm can adapt to different touch scenarios (such as writing, swiping, clicking) and reduce the communication frequency. The parallel processing channels are used to perform Kalman filtering on the touch signal after noise reduction and feature extraction based on the parameters provided by the programmable window register group to obtain the processed touch signal.

[0080] Specifically, the address range of the programmable window register group is from 0x5A00 to 0x5A0F; the parallel processing channel adopts a 10-channel time-division multiplexing architecture.

[0081] Specifically, when there is no touch operation for ≥50 ms, the touch processing module enters the sleep state.

[0082] Specifically, the comparison of the effects of the four-stage pipeline control logic of the present invention and the traditional design scheme (such as CPU / GPU) is shown in Table 2 below:

[0083] Table 2: Comparison table of the effects of the four-stage pipeline control logic and the traditional design scheme

[0084]

[0085] In a preferred embodiment, the dynamic power management module includes a ring oscillator array, a SAR-type digital calibration circuit, an over-temperature protection circuit, and dynamic power gating. The ring oscillator array is used to generate a reference clock signal for full-chip clock synchronization to facilitate correct synchronization and sampling of input signals in digital signal processing. The SAR-type digital calibration circuit is used to calibrate the power-off or power-restoration timing of the dynamic power gating based on the power gating drive signal to obtain a calibrated power gating drive signal, ensuring that the power-off delay ≤0.8 ms. Specifically, according to the current temperature and working state, based on the reference clock signal generated by the ring oscillator array, through a 32-step capacitor array, the delay of the dynamic power gating in the dynamic power gating is adjusted, and the calibration period is 10 ms ± 0.2 ms, periodically compensating for process deviations and temperature drifts to ensure the accuracy of the power switching timing. The over-temperature protection circuit is used to detect the temperature of the ASIC chip. The dynamic power gating is used to dynamically adjust the supply voltage and frequency of the video decoding module, the touch processing module, and the communication interface module according to the power consumption information, control the power consumption of the video decoding module, the touch processing module, and the communication interface module; and cut off the power supply based on the calibrated power gating drive signal and / or the detected temperature. When the temperature is lower than the set threshold, the power supply is automatically cut off and the ASIC chip is turned off; when the temperature is not lower than the set threshold, a progressive frequency reduction mechanism is started. The chip temperature is an important factor affecting the power management strategy. Directly cutting off the power supply at high temperatures may lead to a decrease in chip performance or stability problems. Therefore, the present invention provides an over-temperature protection circuit to protect the ASIC chip.

[0086] Specifically, the ring oscillator array is composed of 128 parallel ring oscillator units, which can cancel the temperature drift of a single oscillator and achieve a temperature stability of ±25 ppm / °C; the over-temperature protection circuit triggers progressive frequency reduction when the junction temperature > 85°C.

[0087] Specifically, dynamic power gating uses a switch matrix composed of high-precision NMOS / PMOS transistors to directly control the power supply on / off of the video decoding module. Its leakage current < 1 nA, supports fast switching (rise / fall time ≤ 50 ns), and the calibration signal output by the SAR-type digital calibration circuit directly drives the switch matrix to ensure that the timing meets the requirements (such as cut-off within 0.5 ms).

[0088] In a preferred embodiment, the communication interface module is used for communication between the ASIC chip and the cloud, adopting a low-power wireless communication method, such as Bluetooth 5.0 or Wi-Fi 6, etc., to ensure stable data transmission while reducing communication power consumption.

[0089] In a preferred embodiment, the communication protocol switching logic of the communication protocol switching module is as follows: when the bandwidth demand > the bandwidth demand threshold, such as 100 Mbps, the communication protocol switching module forcibly switches the communication interface module to the Wi-Fi 6 protocol; when the bandwidth demand ≤ the bandwidth demand threshold, the communication protocol switching module maintains the Bluetooth 5.0 protocol for the communication interface module. Among them, the bandwidth demand threshold is dynamically adjusted according to the channel quality, and the adjustment formula is: bandwidth demand threshold = 100×(1 + SNR / 20) Mbps, where SNR is the current signal-to-noise ratio (dB). For example: when the detected bandwidth demand is 120 Mbps and SNR = 15 dB, the bandwidth demand threshold = 100×(1 + 15 / 20) = 175 Mbps. Since 120 Mbps < 175 Mbps, the Bluetooth 5.0 connection is maintained at this time.

[0090] The beneficial effects of the cloud mobile phone terminal decoding chip architecture of the present invention are described in detail below through specific embodiments:

[0091] The cloud mobile phone terminal decoding chip architecture of this embodiment is respectively adopted in cloud mobile phone terminal devices, mobile intelligent terminals, and smart home control terminals. Among them, in the simulation example of the cloud mobile phone terminal device, the input video data is 4K@60fps H.265, the sampled touch signal is 200 Hz, and the simulated power consumption is 0.5 W (TT process corner); in the simulation example of the mobile intelligent terminal, the input video data is 1080p@60fps H.264, the sampled touch signal is 150 Hz, and the simulated power consumption is 0.35 W (TT process corner); in the simulation example of the smart home control terminal, the input video data is 720p@30fps H.264, the sampled touch signal is 100 Hz, and the simulated power consumption is 0.28 W (TT process corner). It can be seen that by adopting the cloud mobile phone terminal decoding chip architecture of the present invention, low-power processing of video decoding and touch signal uploading can be achieved.

[0092] The decoding chip architecture of the cloud phone terminal in this embodiment is subjected to simulation tests under the same test conditions as the general-purpose chip in the prior art, as shown in Table 3 below:

[0093] Table 3: Simulation comparison table

[0094]

[0095] It can be seen that, under the same test conditions, the decoding chip architecture of the cloud phone terminal in this embodiment has lower system power consumption, faster wake-up latency, and larger energy efficiency ratio than the general-purpose chip using the prior art. At the same time, all module designs in the embodiments of the present invention comply with the IEEE 2415-2019 low-power design specification, the key timing paths are verified by PrimeTime to meet the 600MHz frequency requirement, and the power consumption simulation covers three process corners (TT / FF / SS).

[0096] Embodiment 2

[0097] As Figure 2 shown, this embodiment provides a collaborative optimization method for the decoding chip architecture of a cloud phone terminal, including the following steps:

[0098] (1) The video decoding module detects the frame processing status of video data, including the frame start flag, frame end flag, and frame gap. When the frame gap exceeds the preset frame gap threshold, a power gating drive signal is sent to trigger the dynamic power management module, where the detection accuracy is ±0.2ms. Specifically:

[0099] (1.1) The frame end detector in the pipeline control logic unit monitors the end signals of the entropy decoding unit and the motion compensation unit, marks the completion of the decoding of the current frame, and obtains the frame end flag.

[0100] (1.2) The high-precision timer in the pipeline control logic unit calculates the frame gap between two video frames based on the reference clock signal provided by the dynamic power management module.

[0101] (1.3) The threshold comparator in the pipeline control logic unit compares the frame gap with the preset frame gap threshold in real time. When the frame gap exceeds the preset frame gap threshold (such as 5ms), a power gating drive signal is sent to trigger the dynamic power management module.

[0102] Specifically, frame gap detection is usually used to identify the idle state of the video stream. When the interval between video frames is long, it indicates that the processing load of the current video stream is low, and the low-power mode can be entered. Frame gap detection is the starting point of the entire process and is used to detect the time interval between video frames of video data. In video processing, the frame gap refers to the time difference between two consecutive video frames.

[0103] (2) The touch processing module receives the touch signal input by the touch speed sensor, performs wavelet processing and Kalman filtering to obtain the processed touch signal, and outputs it to the communication interface module. Specifically:

[0104] (2.1) The wavelet transform unit receives the touch signal input by the touch speed sensor, performs noise reduction and feature extraction, and then sends it to the Kalman filter.

[0105] (2.2) The Kalman filter performs Kalman filtering on the touch signal after noise reduction and feature extraction to obtain the processed touch signal, and sends it to the second output unit:

[0106] (2.2.1) The programmable window register group determines the window size of the Kalman filter based on the real-time data input by the touch speed sensor (such as the touch movement speed) according to the dynamic adjustment logic, and updates the register value according to the window size.

[0107] (2.2.2) The programmable window register group provides the size and weight parameters of the sliding window for the Kalman filter.

[0108] (2.2.3) The parallel processing channel performs Kalman filtering on the touch signal after noise reduction and feature extraction based on the parameters provided by the programmable window register group to obtain the processed touch signal.

[0109] (2.3) The second output unit outputs the processed touch signal to the communication interface module.

[0110] (3) The communication protocol switching module controls the switching of the communication protocol based on the bandwidth requirement.

[0111] Specifically, the switching logic of the communication protocol is as follows: when the bandwidth requirement > the bandwidth requirement threshold, the communication protocol switching module forcibly switches the communication interface module to the Wi-Fi 6 protocol; when the bandwidth requirement ≤ the bandwidth requirement threshold, the communication interface module maintains the Bluetooth 5.0 protocol.

[0112] (4) The communication interface module determines the power consumption information based on the decoded video data and the processed touch signal, and outputs it to the dynamic power management module, and sends the decoded video data and the filtered touch signal to the cloud based on the current communication protocol.

[0113] (5) The dynamic power management module dynamically adjusts the supply voltage and frequency of the video decoding module, the touch processing module, and the communication interface module according to the power consumption information, and controls the power consumption of the video decoding module, the touch processing module, and the communication interface module.

[0114] (6) The dynamic power management module detects the temperature of the ASIC chip and cuts off the power supply according to the power gating drive signal and / or the detected temperature.

[0115] Specifically, when the temperature of the ASIC chip is lower than 85°C, the dynamic power management module cuts off the power supply within 0.5 ms; when the temperature of the ASIC chip is higher than or equal to 85°C, the dynamic power management module starts a progressive frequency reduction mechanism to reduce power consumption by gradually decreasing the clock frequency of the ASIC chip.

[0116] Specifically, the power cut-off mechanism is usually used for dynamic power management to reduce the static power consumption of the ASIC chip by quickly cutting off the power supply. When the frame gap reaches or exceeds the preset frame gap threshold, the system triggers the power cut-off mechanism to reduce power consumption. The purpose of this step is to quickly cut off the power supply to reduce power consumption when the video processing load is low.

[0117] Specifically, in the progressive frequency reduction mechanism, the clock frequency of the ASIC chip is gradually decreased by reducing the clock frequency by 5% every 30 seconds. This gradual frequency reduction method can ensure the stable operation of the ASI chip at high temperatures while gradually reducing power consumption.

[0118] This embodiment can achieve power consumption optimization in different working states through frame gap detection, temperature detection, and dynamic power management strategies. This mechanism is particularly suitable for the decoding chip of the cloud mobile phone terminal and can significantly reduce power consumption while ensuring performance.

[0119] Embodiment 3

[0120] This embodiment provides a cloud mobile phone terminal device, including the cloud mobile phone terminal decoding chip architecture of Embodiment 1.

[0121] Embodiment 4

[0122] This embodiment provides a processing device, which can be a processing device applicable to a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the collaborative optimization method of Embodiment 2.

[0123] The processing device includes a processor, a memory, a communication interface module, and a bus. The processor, the memory, and the communication interface module are connected through the bus to complete mutual communication. The memory stores a computer program that can run on the processing device. When the processing device runs the computer program, it executes the collaborative optimization method provided in Embodiment 2.

[0124] In some implementations, the memory can be a high-speed random access memory (RAM: Random Access Memory), and may also include non-volatile memory, such as at least one disk memory.

[0125] In some other implementations, the processor can be various types of general-purpose processors such as a central processing unit (CPU) or a digital signal processor (DSP), which are not limited herein.

[0126] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0127] Those skilled in the art can understand that the structure of the above-mentioned computing device is only a part of the structure related to the solution of the present invention and does not constitute a limitation on the computing device to which the solution of the present invention is applied. The specific computing device may include more or fewer components, or combine certain components, or have different component arrangements.

[0128] Embodiment 5

[0129] This embodiment provides a computer program product corresponding to the collaborative optimization method provided in Embodiment 2. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing the collaborative optimization method described in Embodiment 2 are loaded.

[0130] A computer-readable storage medium can be a tangible device that holds and stores instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.

[0131] The computer-readable storage medium provided in the above-mentioned embodiment has the same implementation principle and technical effects as the above method embodiment, and will not be elaborated herein.

[0132] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0133] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0135] The above embodiments are only used to illustrate the present invention. The structures, connection manners, manufacturing processes, etc. of the components can all be changed. Any equivalent transformation and improvement based on the technical solutions of the present invention should not be excluded from the protection scope of the present invention.

Claims

1. A cloud mobile phone terminal decoding chip architecture, characterized in that, An ASIC chip is adopted, which includes a video decoding module, a communication interface module, a touch processing module, a dynamic power management module, and a communication protocol switching module; The video decoding module is used to perform motion compensation processing and entropy decoding on the input video data to generate decoded video data and output it to the communication interface module, and to detect the frame processing status of the video data and determine the power gating drive signal of the dynamic power management module and output it to the dynamic power management module; The touch processing module is used to perform wavelet processing and Kalman filtering on the input touch signal to obtain the processed touch signal and output it to the communication interface module; The dynamic power management module is used to dynamically adjust the supply voltage and frequency of the video decoding module, the touch processing module, and the communication interface module according to the power consumption information, detect the temperature of the ASIC chip, and cut off the power supply according to the power gating drive signal and / or the detected temperature; The communication protocol switching module is used to control the switching of the communication protocol based on the bandwidth requirement; The communication interface module is used to determine the power consumption information based on the decoded video data and the processed touch signal and output it to the dynamic power management module; And based on the current communication protocol, send the decoded video data and the filtered touch signal to the cloud.

2. The cloud mobile phone terminal decoding chip architecture according to claim 1, wherein The video decoding module includes a number of parallel processing units, and each of the parallel processing units includes: A pipeline control logic unit, which is used to determine the syntax elements based on the compressed bitstream of the video data and detect the frame processing status of the video data, including the frame start flag, the frame end flag, and the frame gap; A motion compensation unit, which is used to perform motion compensation processing on the video data based on the syntax elements to generate a predicted frame of the video data; An entropy decoding unit, which is used to perform entropy decoding on the predicted frame of the video data to generate decoded video data; A first output unit, which outputs the decoded video data to the communication interface module; A local register, which is used to store the decoded video data.

3. The cloud mobile phone terminal decoding chip architecture according to claim 2, characterized in that, Each of the pipeline control logic units adopts a four-stage pipeline control logic, where: The first stage of the pipeline is binary conversion, which is used to convert the compressed bitstream of the input video data into binary symbols; The second stage of the pipeline is context model selection, which is used to dynamically calculate the context model index and select the corresponding probability model for each binary symbol; The third stage of the pipeline is probability state update, which is used to update the probability state of the context model in real time according to the decoded symbols; The fourth stage of the pipeline is bitstream generation, which is used to recombine the decoded binary symbols into syntax elements and output them to the motion compensation unit.

4. The cloud mobile phone terminal decoding chip architecture according to claim 1, characterized in that, The touch processing module includes: A wavelet transform unit, which is used to receive the input touch signal and perform noise reduction and feature extraction; A Kalman filter, which is used to perform Kalman filtering on the touch signal after noise reduction and feature extraction to obtain the processed touch signal; A second output unit, which outputs the filtered touch signal to the communication interface module.

5. The architecture of a decoding chip for a cloud mobile phone terminal according to claim 4, characterized in that The Kalman filter is configured with: A programmable window register bank, which is used to determine the window size of Kalman filtering based on dynamic adjustment logic according to the real-time data input by the touch speed sensor, and update the register value according to the window size; and provide the size and weight parameters of the sliding window for Kalman filtering; A parallel processing channel, which is used to perform Kalman filtering processing on the touch signal after noise reduction and feature extraction based on the parameters provided by the programmable window register bank to obtain the processed touch signal.

6. The decoding chip architecture of a cloud mobile phone terminal according to claim 1, wherein, The dynamic power management module includes: A ring oscillator array, which is used to generate a reference clock signal for full-chip clock synchronization; A SAR-type digital calibration circuit, which is used to calibrate the power gating turn-off or power restoration timing of the dynamic power gating based on the power gating drive signal to obtain the calibrated power gating drive signal; An over-temperature protection circuit, which is used to detect the temperature of the ASIC chip; The dynamic power gating is used to dynamically adjust the supply voltage and frequency of the video decoding module, touch processing module, and communication interface module according to the power consumption information, and cut off the power supply based on the calibrated power gating drive signal and / or the detected temperature.

7. A collaborative optimization method for the cloud mobile phone terminal decoding chip architecture according to any one of claims 1 to 6, characterized in that Includes: Detect the frame processing status of the video data, and output a power gating drive signal when the frame gap in the frame processing status exceeds a preset frame gap threshold; Receive the input touch signal and perform wavelet processing and Kalman filtering processing to obtain the processed touch signal; Control the switching of the communication protocol based on the bandwidth requirement; Determine the power consumption information based on the decoded video data and the processed touch signal, and send the decoded video data and the filtered touch signal to the cloud based on the current communication protocol; Dynamically adjust the supply voltage and frequency according to the power consumption information to control the power consumption; Detect the temperature of the ASIC chip and cut off the power supply based on the power gating drive signal and / or the detected temperature.

8. A cloud mobile phone terminal device, characterized in that, Includes the cloud phone terminal decoding chip architecture according to any one of claims 1 to 6.

9. A processing device, characterized in that, Includes computer program instructions, wherein when the computer program instructions are executed by a processing device, they are used to implement the steps corresponding to the collaborative optimization method described in claim 7.

10. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, wherein when the computer program instructions are executed by a processor, they are used to implement the steps corresponding to the collaborative optimization method described in claim 7.

Citation Information

Cited By

  • A node wake-up and task scheduling method for edge video stream cooperative processing

    CN122845856A