A device low-delay interaction method and system under an end-cloud coordination architecture

By using a low-latency interaction method under the edge-cloud collaborative architecture, the processing on the edge and cloud is synchronized, and the content is dynamically adapted and filled in and semantic correction is performed. This solves the problems of high latency and semantic fragmentation in edge-cloud collaborative voice interaction, and improves the real-time performance and naturalness of the interaction.

CN122135699APending Publication Date: 2026-06-02GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN Β· China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-04-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing edge-cloud collaborative voice interaction suffers from high latency, semantic fragmentation, and a harsh listening experience, making it unsuitable for interaction scenarios with high real-time requirements. Furthermore, edge-side filling solutions cannot dynamically adapt to cloud-based inference time and user real-time intent.

Method used

A low-latency interaction method under the edge-cloud collaborative architecture is adopted. By synchronously executing edge and cloud processing, an intent vector is generated and dynamically adapted to fill content. Combined with real-time semantic similarity correction and acoustic emotion smooth transition, low-latency, semantically coherent and acoustically smooth interaction is achieved.

Benefits of technology

Significantly reduces user perception latency, avoids semantic conflicts between fill content and cloud response, eliminates abruptness in voice splicing, and improves real-time performance, coherence, and naturalness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135699A_ABST
    Figure CN122135699A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of voice interaction technology, specifically relating to a low-latency interaction method and system for devices under an edge-cloud collaborative architecture. The method includes simultaneous execution steps of edge-side processing and cloud-side processing of streaming user voice. Edge-side processing generates a first intent vector, while cloud-side processing pushes voice features to a large language model for deep generation. This involves obtaining the inference prediction time and the prediction variance of the cloud-based deep generation, adjusting the initial response content based on alignment parameters, and then outputting the content immediately following the edge-side filling. This invention effectively shields cloud inference and network latency, avoids semantic conflicts and abrupt acoustic splicing, and improves the real-time performance and naturalness of edge-cloud collaborative voice interaction.
Need to check novelty before this filing date? Find Prior Art