Intelligent television end cloud collaborative interaction method and system based on large model

By employing NPU+ model compression and edge-cloud collaborative interaction methods in smart TVs, the problems of high response latency, high energy consumption, and insufficient privacy protection in smart TVs have been solved, achieving efficient and secure multimodal interaction and data processing.

CN121099093APending Publication Date: 2025-12-09SHENZHEN IBD INTELLIGENT TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511184595.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing smart TVs suffer from high response latency, high energy consumption, insufficient privacy protection, and weak multimodal fusion capabilities during interaction. Furthermore, traditional solutions rely on cloud computing power, which leads to response latency and high energy consumption.

Method used

By employing an NPU+model compression approach combined with edge-cloud collaborative interaction, and through local deployment of the DeepSeek model and pruning strategies, multimodal data synchronization and processing are achieved. Furthermore, data security is ensured through a task-leveling mechanism and a federated learning framework to guarantee user privacy.

Benefits of technology

It achieves reduced response latency and energy consumption while ensuring model accuracy, improves the accuracy of multimodal instruction coordination and user privacy protection, and enhances the interaction efficiency and security of smart TVs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121099093A_ABST
    Figure CN121099093A_ABST
Patent Text Reader

Abstract

The invention provides a smart television end-cloud collaborative interaction method and system based on a large model. The method comprises the steps that a personalized service application module calls services provided by the system through a lightweight sdk api interface of an end-cloud collaborative module; the service opens a sensor of a bottom hardware module to obtain multi-modal data; inputting the multi-modal data to a multi-modal fusion analysis module for data synchronization and alignment, feature extraction and fusion analysis, and then entering an end-cloud collaboration module; and the end cloud collaboration module determines a model to be called according to the task classification, processes according to the corresponding model, and transmits a data result to the personalized service application module through an API interface of the service in a callback manner. By adopting an NPU + model compression mode, the problem of high response delay (gt; gt) caused by the fact that a traditional television chip cannot locally deploy a large model is solved; the method solves the problems of high energy consumption and large energy consumption, and achieves the purposes of reducing the memory occupancy rate and shortening the voice response delay on the premise of ensuring the model precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent television, in particular to an intelligent television terminal cloud collaborative interaction method and system based on a large model. BACKGROUND

[0002] The current intelligent television market is highly competitive, and mainstream brands have tried to access large language models, but there are problems such as rigid interaction, functional homogeneity, and insufficient privacy protection. Traditional solutions rely on cloud computing power, resulting in response delays, and weak multi-modal fusion capabilities. And the traditional television chip cannot locally deploy large models, resulting in high response delay (> 500ms) and high energy consumption.

[0003] Patent document CN119402690A discloses a television interaction method and system based on application scenario large model intelligent decision-making, the method comprising: an intelligent television receiving a user input interaction instruction, and performing intent recognition reasoning on the interaction instruction to obtain an intent set; performing intent classification on the intent set to obtain an intent label corresponding to each intent in the intent set; based on the intent label, identifying the current application scenario through an AI large model service, and making intelligent decisions according to the intent label and the current application scenario to obtain an application scenario large model intelligent decision-making instruction set; and issuing the instructions of the application scenario large model intelligent decision-making instruction set to the corresponding television system or application to control the television system or application to respond to the instructions. However, patent document CN119402690A lacks privacy protection.

[0004] Therefore, there is a need in the market for an intelligent television terminal cloud collaborative interaction method and system based on a large model that can protect user privacy during the interaction process while improving, reducing cloud computing power calling costs, and improving multi-modal instruction collaborative accuracy. SUMMARY

[0005] In view of the defects in the prior art, the purpose of the present application is to provide an intelligent television terminal cloud collaborative interaction method and system based on a large model.

[0006] According to the intelligent television terminal cloud collaborative interaction method and system based on a large model provided by the present application, the method comprises:

[0007] The personalized service application module calls the services provided by the system through the lightweight sdk api interface of the terminal cloud collaborative module;

[0008] The service opens the sensors of the underlying hardware module to obtain multi-modal data;

[0009] The multi-modal data is input into the multi-modal fusion analysis module for data synchronization and alignment, as well as feature extraction and fusion analysis, and then enters the terminal cloud collaborative module;

[0010] The end-cloud cooperation module determines the model to be called according to the task classification, and transmits the data result to the personalized service application module through the API interface callback of the service after processing according to the corresponding model.

[0011] Preferably, the model includes a local model and a cloud large model, when the cloud large model is called, the data in the cloud is first desensitized by data security; when the local model is called for processing, the DeepSeek model supported by NPU computing power is called, and the local knowledge base is combined to optimize the personalized experience.

[0012] Preferably, the end-cloud cooperation module determines the model to be called according to the task classification, and realizes the task allocation decision through double-dimensional dynamic adjustment, including:

[0013] According to the real-time optimization of the task execution path of the hardware resource load and the network state, including calculating the task classification score, and judging whether the current task classification score is less than the dynamic threshold, if yes, the local model is executed; if not, the cloud large model is executed.

[0014] Preferably, the task classification score refers to the classification score of various tasks through three dimensions of real-time, computational complexity and information privacy, including the following steps:

[0015] Step S201: Set the real-time score St, evaluate the score by measuring the sensitivity of the task to delay, the higher the delay tolerance, the higher the score;

[0016] Step S202: Set the complexity score Sc, evaluate the score by evaluating the demand of the task for computing resources, the higher the resource demand, the higher the score;

[0017] Step S203: Set the information privacy score Sp, evaluate the score by quantifying the data sensitivity, the higher the sensitivity, the lower the score;

[0018] Step S204: Calculate the classification score of each task according to the real-time score, the complexity score and the information privacy score, the formula is as follows:

[0019] Score=α*St+β*Sc+γ*Sp

[0020] α+β+γ=1

[0021] Wherein, α, β and γ are the corresponding weights of real-time, complexity and information privacy set in the end-cloud cooperation module.

[0022] Preferably, the dynamic threshold calculation formula is as follows:

[0023] Dynamic threshold=δ×(NPU available computing power / total computing power)+θ×(1-network delay / baseline delay)

[0024] wherein, δ represents the NPU computing power weight, and θ represents the network quality weight.

[0025] Preferably, the desensitization processing through data security includes camera / microphone data local desensitization processing, only uploading feature values to the cloud, and user privacy data being encrypted and stored in an end-side local knowledge base.

[0026] Preferably, the hardware module adopts a television main control chip supporting an INT8 instruction set and having an NPU computing power greater than 5 TOPS, a PCIe extensible interface, and a sensor.

[0027] The sensor includes a microphone array / image / environment sensor.

[0028] Preferably, the Deep Seek model is deployed locally by hybrid precision quantization and pruning strategy.

[0029] The quantization method includes weight quantization and activation value quantization.

[0030] The pruning strategy includes hierarchical pruning and channel-level pruning.

[0031] The weight quantization adopts an asymmetric quantization strategy to reduce the influence of quantization error on model accuracy, and the steps include:

[0032] Step S101: Collecting the minimum value and the maximum value of the weight data, determining the dynamic range of the weight, and directly using the actual range corresponding to the original weight for quantization, that is, the quantization range.

[0033] Step S102: Calculating the step size and the quantization zero point according to the quantization bit number and the quantization range.

[0034] Step S103: Mapping the original weight to the quantized value.

[0035] Preferably, the activation value quantization adopts 4-bit symmetric quantization.

[0036] The hierarchical pruning first calculates the weight contribution degree of each layer, and removes redundant layers with a weight contribution degree lower than 0.3 in the attention mechanism.

[0037] The channel-level pruning first performs importance evaluation on each channel in the fully connected layer, determines a pruning threshold, retains channels with importance higher than the threshold, removes channels with importance lower than the threshold, and retains a rate of ≥60%, further compressing the model volume.

[0038] According to the intelligent television end-cloud collaborative interaction system based on a large model provided by the application, the system includes a personalized service application module, an end-cloud collaborative module, a hardware module, and a multi-modal fusion analysis module.

[0039] The personalized service application module calls the services provided by the system through the lightweight sdk api interface of the end-cloud collaboration module; the service opens the sensor of the underlying hardware module to obtain multi-modal data; the multi-modal data is input into the multi-modal fusion analysis module for data synchronization and alignment, and feature extraction and fusion analysis, and then enters the end-cloud collaboration module; the end-cloud collaboration module determines the model to be called according to the task classification, processes the data according to the corresponding model, and then transmits the data result to the personalized service application module through the API interface of the service.

[0040] Compared with the prior art, the application has the following beneficial effects:

[0041] 1、The application solves the problems of high response delay (> 500ms) and large energy consumption caused by the inability of traditional TV chips to locally deploy large models by adopting the NPU+model compression method, thereby achieving the reduction of memory occupancy and the shortening of voice response delay under the premise of ensuring model accuracy.

[0042] 2、The application solves the problem that a single processing architecture cannot balance real-time interaction and complex computing requirements by designing an end-cloud collaboration task classification mechanism, thereby realizing the hierarchical optimization of local processing of real-time tasks (voice wake-up) and cloud-side collaboration of generated tasks (video creation), improving 4K video generation efficiency, and reducing cloud computing power calling costs.

[0043] 3、The application solves the problems of single interaction mode (only voice control) and high false trigger rate of traditional smart TVs by constructing a multi-modal context perception model, thereby improving the multi-modal instruction collaboration accuracy, emotion recognition and content recommendation matching degree.

[0044] 4、The application solves the security risk problem of user privacy data cloud transmission by designing a federated learning framework and hierarchical encryption, realizes local desensitization processing of biometric data, and the encryption strength of the end-side knowledge base reaches the AES-256 standard, which meets the GDPR BRIEF DESCRIPTION OF DRAWINGS

[0045] Other features, objects and advantages of the application will become more apparent after reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0046] Figure 1 It is a flowchart of the end-cloud collaborative interaction method of the smart TV based on a large model. DETAILED DESCRIPTION

[0047] The application will be described in detail below with specific examples. The following examples will help those skilled in the art to further understand the application, but do not limit the application in any form. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the application. These are within the scope of the present application.

[0048] The application constructs a double-layer architecture of "end-side multimodal fusion + cloud complex generation", breaking through the single interaction mode of traditional smart TVs; through the design of localized privacy protection and dynamic emotion analysis technology, the safety and personalization are doubled.

[0049] Example 1

[0050] According to the intelligent television end-cloud collaborative interaction method based on a large model provided by the application, as shown in Figure 1 The personalized service application module calls the services provided by the system through the lightweight sdk api interface of the end-cloud collaborative module; the service opens the microphone and camera sensors of the underlying hardware module to obtain multi-modal data, including voice and image data; the multi-modal data is input into the multi-modal fusion analysis module for data synchronization and alignment, as well as feature extraction and fusion analysis, and then enters the end-cloud collaborative module; the end-cloud collaborative module determines the model to be called according to the task classification, processes the data according to the corresponding model, and then transmits the data results to the personalized service application module through the API interface callback of the service. The model includes a local model and a cloud large model, when the cloud large model is called, the data of the cloud is first desensitized by data security; when the local model is called for processing, the DeepSeek model supported by NPU computing power is called, and the local knowledge base is combined to optimize the personalized experience. Among them, the lightweight SDK encapsulates the API, which supports third-party application calling to generate services (such as automatic short video editing).

[0051] The end-cloud collaborative module determines the model to be called through double-dimensional dynamic adjustment to realize task allocation decision, the core of which is to optimize the task execution path according to the real-time hardware resource load and network state, including calculating the task classification score and determining whether the current task classification score is less than the dynamic threshold, if yes, the local model is executed; if not, the cloud large model is executed. Among them, the task classification score refers to the classification score of various tasks through real-time, computational complexity, and information privacy, including the following steps:

[0052] Step S201: Set the real-time score St (score: 0-1), evaluate the score by measuring the sensitivity of the task to delay, the higher the delay tolerance, the higher the score.

[0053] Step S202: Set the complexity score Sc (score: 0-1), evaluate the score by assessing the task's demand for computing resources, the higher the resource demand, the higher the score.

[0054] Step S203: Set the information privacy score Sp (score: 0-1), evaluate the score by quantifying data sensitivity, the higher the sensitivity, the lower the score.

[0055] Step S204: Calculate the individual task ranking score according to the real-time score, complexity score and information privacy score, the formula is as follows:

[0056] Score = a * St + b * Sc + g * Sp

[0057] a + b + g = 1

[0058] Where a, b and g are the real-time, complexity and information privacy weights set in the end-cloud collaboration module.

[0059] The dynamic threshold formula is as follows:

[0060] Dynamic threshold = d x (NPU available computing power / total computing power) + 0 x (1 - network delay / baseline delay)

[0061] Where d represents the NPU computing power weight, reflecting the priority of local processing capability, and the preferred value is 0.6, 0 represents the network quality weight, reflecting the influence of cloud transmission efficiency, and the preferred value is 0.4. The baseline delay is the ideal network delay set according to historical data, the higher the NPU available computing power ratio, the higher the score of this part, the threshold is adjusted up, more tasks are processed locally, the higher the network delay ratio, the lower the score of this part, the threshold is adjusted down, more tasks are transferred to the cloud. At the same time, when the local Deep Seek model credibility return value score < 0.85, the task cloud execution is automatically triggered, and the classification error is corrected.

[0062] The desensitization processing through data security includes local desensitization processing of camera / microphone data, uploading only feature values to the cloud, and encrypting user privacy data and storing it in the local knowledge base on the terminal side.

[0063] The hardware module uses a television main control chip that supports INT8 instruction set and has computing power > 5TOPS, PCIe extensible interface and sensors. The television main control chip can reduce the computing complexity and memory occupation, and efficiently execute the local Deep Seek model inference. The PCIe extensible interface supports external power expansion dock, realizes dynamic expansion of hardware computing power, and provides hardware support for future larger scale model upgrade. The sensors include microphone array / image / environmental sensors.

[0064] The Deep Seek model is deployed locally by combining mixed precision quantization (8-bit weights + 4-bit activation values) with pruning strategies. The quantization method includes weight quantization and activation value quantization. The pruning strategy includes hierarchical pruning and channel-level pruning. The weight quantization adopts an asymmetric quantization strategy to reduce the impact of quantization error on model accuracy. The steps include:

[0065] Step S101: Collect the minimum and maximum values of the weight data to determine the dynamic range of the weights. The asymmetric quantization directly uses the actual range corresponding to the original weights for quantization, i.e., the quantization range. For example, the original weight distribution is in the range [-1.5, 2.0].

[0066] Step S102: Calculate the step size and quantization zero point based on the quantization bit number (e.g., 8 bits) and the quantization range.

[0067] Step S103: Map the original weights to the quantized values.

[0068] The activation value quantization adopts 4-bit symmetric quantization to reduce memory usage while maintaining model performance.

[0069] The hierarchical pruning first calculates the weight contribution degree of each layer, removes redundant layers in the attention mechanism with a weight contribution degree lower than 0.3, and reduces the model parameter quantity by 30%.

[0070] The channel-level pruning first evaluates the importance of each channel in the fully connected layer, determines the pruning threshold, retains channels with importance higher than the threshold, removes channels with importance lower than the threshold, and retains a rate of ≥60%, further compressing the model volume.

[0071] The multi-modal fusion analysis module includes fusion of speech, gesture, and picture understanding technology, constructs a context-aware model, and supports dialect recognition and emotion analysis in voice interaction (e.g., recommending comedy content after identifying user excitement emotions). Gesture control is linked with screen content.

[0072] The personalized service application module is based on Deep Seek real-time analysis of user behavior data, dynamically generates education content / fitness plans, and optimizes recommendation weights based on emotional feedback. It also includes a child mode that filters inappropriate content through a local model and limits viewing time through visual recognition.

[0073] Example 2

[0074] The application also provides a large model-based intelligent television terminal cloud collaborative interaction system, which can be realized by executing the process steps of the large model-based intelligent television terminal cloud collaborative interaction method, that is, the large model-based intelligent television terminal cloud collaborative interaction method can be understood by those skilled in the art as the preferred embodiment of the large model-based intelligent television terminal cloud collaborative interaction system.

[0075] According to the application, a large model-based intelligent television terminal cloud collaborative interaction system is provided, which comprises a personalized service application module, a terminal cloud collaboration module, a multi-modal fusion analysis module and a hardware module. According to the application, a large model-based intelligent television terminal cloud collaborative interaction method is provided, which comprises that the personalized service application module calls the services provided by the system through the lightweight sdk api interface of the terminal cloud collaboration module; the service opens the sensor of the underlying hardware module to obtain multi-modal data; the multi-modal data is input into the multi-modal fusion analysis module for data synchronization and alignment, and feature extraction and fusion analysis, and then enters the terminal cloud collaboration module; the terminal cloud collaboration module determines the model to be called according to the task classification, processes the data according to the corresponding model, and then transmits the data results to the personalized service application module through the API interface of the service.

[0076] The model comprises a local model and a cloud large model. When the cloud large model is called, the data of the cloud is first desensitized by data security; when the local model is called for processing, the DeepSeek model supported by NPU computing power is called, and the local knowledge base is combined to optimize the personalized experience. The terminal cloud collaboration module determines the model to be called through double-dimensional dynamic adjustment to realize task allocation decision, and optimizes the task execution path in real time according to the hardware resource load and network state, comprising computing task classification score, and judging whether the current task classification score is less than the dynamic threshold value, if yes, the local model is called for execution; if not, the cloud large model is called for execution. The task classification score refers to the classification score of various tasks through real-time, computing complexity and information privacy, comprising the following modules:

[0077] Module M201: set the real-time score St, evaluate the score by measuring the sensitivity of the task to delay, the higher the delay tolerance, the higher the score;

[0078] Module M202: set the complexity score Sc, evaluate the score by evaluating the demand of the task for computing resources, the higher the resource demand, the higher the score;

[0079] Module M203: set the information privacy score Sp, evaluate the score by quantifying the data sensitivity, the higher the sensitivity, the lower the score;

[0080] Module M204: Calculate the individual task ranking score according to the real-time score, complexity score and information privacy score, the formula is as follows:

[0081] Score = a * St + b * Sc + g * Sp

[0082] a + b + g = 1

[0083] Wherein, a, b and g are the real-time, complexity and information privacy corresponding weights set in the end-cloud collaboration module. The dynamic threshold calculation formula is as follows:

[0084] Dynamic threshold = d * (NPU available computing power / total computing power) + q * (1 - network delay / reference delay)

[0085] Wherein, d represents the NPU computing power weight, and q represents the network quality weight.

[0086] The desensitization processing through data security includes camera / microphone data local desensitization processing, only uploading feature values to the cloud, and user privacy data encrypted storage in the end local knowledge base. The hardware module uses a television main control chip supporting INT8 instruction set, NPU computing power > 5TOPS, PCIe extensible interface and sensor. The sensor includes microphone array / image / environment sensor. The Deep Seek model is deployed locally by mixed precision quantization combined with pruning strategy; the quantization method includes weight quantization and activation value quantization; the pruning strategy includes hierarchical pruning and channel level pruning; the weight quantization adopts an asymmetric quantization strategy to reduce the influence of quantization error on model accuracy, and the module includes:

[0087] Module M101: Collect the minimum and maximum values of the weight data, determine the dynamic range of the weight, and directly use the actual range corresponding to the original weight for quantization, that is, the quantization range;

[0088] Module M102: Calculate the step size and quantization zero point according to the quantization bit number and quantization range;

[0089] Module M103: Map the original weight to the quantized value.

[0090] The activation value quantization adopts 4bit symmetric quantization; the hierarchical pruning first calculates the weight contribution degree of each layer, removes the redundant layers with weight contribution degree lower than 0.3 in the attention mechanism; the channel level pruning first performs importance evaluation on each channel in the fully connected layer, determines the pruning threshold, retains the channels with importance higher than the threshold, removes the channels with importance lower than the threshold, and the retention rate is greater than or equal to 60%, further compressing the model volume.

[0091] Those skilled in the art know that, in addition to implementing the system provided by the present application and each device, module and unit thereof in the form of pure computer readable program code, the system provided by the present application and each device, module and unit thereof can also be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. by logically programming the method steps to achieve the same functions. Therefore, the system provided by the present application and each device, module and unit thereof can be considered as a hardware component, and the devices, modules and units included therein for achieving various functions can also be considered as structures within the hardware component; the devices, modules and units for achieving various functions can also be considered as both software modules implementing methods and structures within hardware components.

[0092] The specific embodiments of the present application are described above. It needs to be understood that the present application is not limited to the specific embodiments described above, and various changes or modifications can be made by those skilled in the art within the scope of the claims, which does not affect the essential content of the present application. The embodiments of the present application and the features in the embodiments can be combined with each other in any manner without conflict.

Claims

1. A large model-based intelligent television terminal cloud interaction method, characterized in that, The application comprises the following steps: The personalized service application module calls the services provided by the light-weight sdk api interface of the end-cloud collaboration module; The service opens the sensors of the underlying hardware module to obtain multi-modal data; The multi-modal data is input into the multi-modal fusion analysis module for data synchronization and alignment, as well as feature extraction and fusion analysis, and then enters the end-cloud collaboration module; The end-cloud collaboration module determines the model to be called according to the task classification, processes the data according to the corresponding model, and then transmits the data results to the personalized service application module through the API interface of the service.

2. The large model-based intelligent television terminal cloud interaction method according to claim 1, characterized in that, The model includes a local model and a cloud large model. When the cloud large model is called, the data on the cloud is first desensitized by data security; when the local model is called for processing, the DeepSeek model supported by NPU computing power is called, and the local knowledge base is combined to optimize the personalized experience.

3. The large model-based intelligent television terminal cloud interaction method of claim 1, wherein, The end-cloud collaboration module determines the model to be called according to the task classification, and realizes task allocation decision through double-dimensional dynamic adjustment, including: According to the real-time optimization of hardware resource load and network state, the task execution path is optimized, including calculating the task classification score and determining whether the current task classification score is less than the dynamic threshold. If yes, the local model is executed; if no, the cloud large model is executed.

4. The large model-based intelligent television terminal cloud interaction method of claim 3, characterized in that, The task classification score refers to the classification score of various tasks in three dimensions of real-time, computational complexity and information privacy, including the following steps: Step S201: Set the real-time score St. The score is evaluated by measuring the sensitivity of the task to delay. The higher the delay tolerance, the higher the score. Step S202: Set the complexity score Sc. The score is evaluated by assessing the demand of the task for computing resources. The higher the resource demand, the higher the score. Step S203: Set the information privacy score Sp. The score is evaluated by quantifying the data sensitivity. The higher the sensitivity, the lower the score. Step S204: Calculate the classification score of each task according to the real-time score, complexity score and information privacy score. The formula is as follows: Score = α * St + β * Sc + γ * Sp α + β + γ = 1 Wherein, α, β and γ are the corresponding weights of real-time, complexity and information privacy set in the end-cloud collaboration module.

5. The large model-based intelligent television terminal cloud interaction method of claim 3, wherein, The dynamic threshold calculation formula is as follows: Dynamic threshold = δ × (NPU available computing power / total computing power) + θ × (1 - network delay / reference delay) Wherein, δ represents the NPU computing power weight, and θ represents the network quality weight.

6. The large model-based intelligent television terminal cloud interaction method of claim 1, wherein, The desensitization processing through data security includes local desensitization processing of camera / microphone data, uploading only feature values to the cloud, and encrypting user privacy data and storing it in the local knowledge base on the terminal side.

7. The large model-based intelligent television terminal cloud interaction method of claim 1, wherein, The hardware module uses a television main control chip supporting INT8 instruction set, NPU computing power > 5TOPS, PCIe extensible interface and sensors; The sensors include microphone array / image / environmental sensors.

8. The large model-based intelligent television terminal cloud interaction method of claim 1, wherein, The Deep Seek model is deployed locally through mixed precision quantization combined with pruning strategy; The quantization method includes weight quantization and activation value quantization; The pruning strategy includes hierarchical pruning and channel-level pruning; The weight quantization adopts an asymmetric quantization strategy, reduces the influence of quantization error on model accuracy, and includes the following steps: Step S101: Collect the minimum value and maximum value of the weight data, determine the dynamic range of the weight, and directly use the actual range corresponding to the original weight for quantization, i.e., the quantization range, for asymmetric quantization; Step S102: Calculate the step size and quantization zero point according to the quantization bit number and the quantization range; Step S103: Map the original weight to the quantized value.

9. The large model-based intelligent television terminal cloud interaction method of claim 1, wherein, The activation value quantization adopts 4-bit symmetric quantization; The hierarchical pruning first calculates the weight contribution degree of each layer, removes redundant layers in the attention mechanism whose weight contribution degree is lower than 0.3, and then performs channel-level pruning. The channel-level pruning first performs importance evaluation on each channel in the full connection layer, determines a pruning threshold, retains channels with importance higher than the threshold, removes channels with importance lower than the threshold, and compresses the model volume.

10. A large model-based intelligent television terminal cloud collaborative interaction system, characterized in that, It includes: A personalized service application module, an end-cloud collaboration module, a hardware module, and a multi-modal fusion analysis module; The personalized service application module calls the services provided by the system through the lightweight sdk api interface of the end-cloud collaboration module; the service opens the sensor of the underlying hardware module to obtain multi-modal data; the multi-modal data is input into the multi-modal fusion analysis module for data synchronization and alignment, as well as feature extraction and fusion analysis, and then enters the end-cloud collaboration module; The end-cloud collaboration module determines the model to be called according to the task classification, processes the data according to the corresponding model, and then transmits the data results to the personalized service application module through the API interface of the service.

Citation Information

Patent Citations

  • Television interaction method and system based on application scene large model intelligent decision

    CN119402690A