A doll voice call method and device based on end-cloud cooperation and a medium

By binding the terminal to the puppet's voice call, configuring preference weights, performing end-side recognition and collaborative priority generation, and dynamically selecting edge cloud nodes and data modalities, the problem of insufficient resource allocation in existing technologies is solved, achieving a balance between real-time emergency communication and privacy security.

CN121334115BActive Publication Date: 2026-02-24SHANGHAI FANHEYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511882116.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-02-24
Estimated Expiration
2045-12-15

AI Technical Summary

Technical Problem

Existing doll voice calling technology lacks a dynamic resource allocation mechanism at the end-to-cloud collaboration level, making it difficult to achieve fine-grained adaptive allocation between emergency and normal scenarios. Furthermore, it lacks hierarchical control measures, failing to reduce invalid data transmission while ensuring the real-time nature and privacy of emergency communications.

Method used

By binding mobile terminals and doll terminals, configuring preference weights, generating an initial unified collaborative parameter set, performing end-side intent recognition and privacy sensitivity recognition, generating a collaborative priority index, selecting target edge cloud nodes, dynamically adjusting data modality and service quality, and triggering differential updates only when the importance of the session changes.

Benefits of technology

It achieves dynamic matching between terminal upload scale and network scheduling capacity during voice calls, ensuring the real-time nature of emergency communications and suppressing the leakage of voice content in privacy-sensitive scenarios, while optimizing the reconfiguration of network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334115B_ABST
    Figure CN121334115B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on end cloud cooperation's doll voice calling method, equipment and medium, it is related to cooperative communication technical field, including, mobile terminal is bound with doll terminal and is configured preference weight, generates initial unified cooperation parameter set;Doll terminal is collected voice after wake-up and completes end side intention, urgency degree and privacy sensitivity identification, generates cooperation priority index and conversation priority label;Through priority information selection target edge cloud node and generates cooperation calling strategy token;Doll terminal forms conversation data according to strategy token and uploads according to difference update condition;Target edge cloud node forms conversation operation record according to priority information and difference data, and aggregate update unified cooperation parameter set.The application generates bandwidth, time delay and update frequency budget of end cloud cooperation control mechanism based on conversation cooperation priority index, realizes the dynamic matching of terminal upload scale, processing complexity and network scheduling ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of collaborative communication technology, and in particular to a method, device and medium for doll voice calling based on end-to-cloud collaboration. Background Technology

[0002] With the continuous development of IoT, mobile communication, and cloud computing technologies, doll-based voice terminals for companionship, caregiving, and human-computer interaction scenarios are gradually evolving from simple local voice interaction to a terminal-network-cloud collaborative processing model. In existing technologies, doll voice calls typically rely on mobile communication networks to access cloud-based voice service platforms, enabling functions such as voice acquisition, voice recognition, semantic parsing, and call scheduling. Edge computing technology can be used to complete some voice processing and service forwarding closer to the terminal, thereby reducing communication latency and improving real-time interaction. Simultaneously, with the development of 5G, cloud-edge collaboration, and low-power terminal technologies, voice services are gradually evolving from fixed resource allocation to on-demand dynamic scheduling. The collaborative capabilities of terminal-side state awareness, network load awareness, and cloud-side computing power scheduling are continuously strengthening, providing a technological foundation for building a low-latency, highly reliable voice call system for complex caregiving scenarios.

[0003] However, existing doll voice calling technology still has significant shortcomings at the end-to-cloud collaboration level: First, existing technologies mostly adopt fixed bandwidth, fixed latency, or static quality of service strategies for voice transmission and call scheduling. There is a lack of dynamic matching mechanism based on the importance of the session between the scale of voice data uploaded by the terminal, the processing complexity, and the network-side scheduling capabilities, making it difficult to achieve fine-grained adaptive allocation of resources between emergency and normal scenarios. Second, existing technologies mainly transmit and store voice data in full, lacking hierarchical control methods for different session priorities, privacy sensitivities, and data modalities. This makes it difficult to ensure both the real-time nature of emergency communication and privacy security, and also fails to reduce invalid data transmission and network load during the call through differential updates. This restricts the security and overall system efficiency of doll voice calling in complex care scenarios. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a doll voice calling method based on end-to-cloud collaboration to solve the problems of existing technologies, such as the difficulty in dynamically matching terminal upload scale, processing complexity and network scheduling capabilities according to the importance of the session, the lack of hierarchical control of voice data under different session priorities and privacy-sensitive scenarios, and the difficulty in reducing invalid transmission load through differential methods.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a doll voice call method based on end-to-cloud collaboration, which includes binding a mobile terminal to a doll terminal, configuring preference weights, and generating an initial unified collaboration parameter set based on the preference weights.

[0008] After detecting a wake-up trigger, the doll terminal collects voice data, performs end-side intent recognition, urgency level recognition, and privacy sensitivity recognition, generates a collaborative priority index based on an initial unified collaborative parameter set, and determines the session priority label.

[0009] The collaboration priority index and session priority label are reported to the core network. The core network selects the target edge cloud node based on the unified collaboration parameter set and generates a collaboration call policy token through the target edge cloud node.

[0010] The doll terminal selects the data modality through the collaborative call strategy token, forms session data and sends it to the target edge cloud, updates the collaborative priority index, calculates the change in the collaborative priority index, and obtains differential data when the change in the collaborative priority index meets the differential update condition.

[0011] The differential data is uploaded to the target edge cloud node. The target edge cloud node forms a session execution record based on the collaboration priority index and the differential data. Multiple session execution records are aggregated and the unified collaboration parameter set is updated.

[0012] As a preferred embodiment of the doll voice call method based on end-to-cloud collaboration described in this invention, the steps of binding the mobile terminal and the doll terminal, configuring preference weights, and generating an initial unified collaboration parameter set based on the preference weights are as follows:

[0013] Bind the mobile terminal to the doll terminal, initialize the doll terminal, and configure privacy preference weight, response speed preference weight, and network resource preference weight.

[0014] The privacy preference weight, response speed preference weight, and network resource preference weight are normalized to generate privacy preference weight coefficient, response speed preference weight coefficient, and network resource preference weight coefficient, which serve as the initial unified collaborative parameter set.

[0015] As a preferred embodiment of the doll voice call method based on end-to-cloud collaboration described in this invention, the doll terminal collects voice data after detecting a wake-up trigger, and performs end-side intent recognition, urgency level recognition, and privacy sensitivity recognition. The specific steps are as follows.

[0016] When the doll terminal detects a wake-up trigger, it digitally samples the voice signal to generate the original voice sequence.

[0017] The original speech sequence is decoded, and template matching is performed using a pre-defined intent phrase template library to obtain the structured intent.

[0018] Using the original speech sequence, the urgency of speech energy, the urgency of speech rate, and the urgency of keywords are calculated, and the maximum value among these three factors is taken as the quantification value of urgency.

[0019] The original speech sequence is converted into a text command string. A pre-defined privacy keyword list is searched word by word in the text command string, and a privacy sensitivity quantification value is set.

[0020] As a preferred embodiment of the doll voice call method based on end-to-cloud collaboration described in this invention, the specific steps for generating a collaboration priority index based on an initial unified collaboration parameter set and determining a session priority label are as follows:

[0021] The collaboration priority index is calculated based on the initial unified collaboration parameter set, structured intent, urgency quantification value, and privacy sensitivity quantification value.

[0022] When the collaboration priority index is zero, the current session is identified as a privacy-suppressing session label.

[0023] When the collaboration priority index is greater than zero and less than or equal to the first priority threshold, the current session is identified as a low-priority session label.

[0024] When the collaboration priority index is greater than the first priority threshold and less than or equal to the second priority threshold, the current session is identified as a medium priority session.

[0025] When the collaboration priority index is greater than the second priority threshold, the current session is determined to be a high-priority session and the session priority label is obtained.

[0026] As a preferred embodiment of the doll voice call method based on edge-cloud collaboration described in this invention, the steps of reporting the collaboration priority index and session priority label to the core network, the core network selecting the target edge cloud node according to the unified collaboration parameter set, and generating a collaboration call policy token through the target edge cloud node are as follows.

[0027] The collaboration priority index and session priority label are reported to the core network. The collaboration adaptation score is calculated for each candidate edge cloud node, and the edge cloud node with the highest score is selected as the target edge cloud node.

[0028] Select the corresponding target service quality level from the preset service quality policy table by using the session priority label;

[0029] On the target edge cloud node, the resource budget parameters on the edge side are calculated using the collaborative priority index.

[0030] Combine the target edge cloud node, target service quality level, and edge resource budget parameters to obtain the collaborative call policy token.

[0031] As a preferred embodiment of the doll voice call method based on edge-cloud collaboration described in this invention, the doll terminal selects a data modality through a collaborative call strategy token, constructs session data, and sends it to the target edge cloud. The specific steps are as follows:

[0032] Using the session priority tag, select the allowed combination of data modalities for upload from the preset set of allowed data modalities;

[0033] The maximum duration of the original voice segments that can be carried is determined based on the collaborative call policy token. Within the maximum duration, voice segments that have been desensitized are extracted to limit the processing complexity of local voice desensitization and feature extraction, thereby forming session data and sending it to the target edge cloud.

[0034] As a preferred embodiment of the doll voice call method based on end-to-cloud collaboration described in this invention, the steps for updating the collaboration priority index, calculating the change in the collaboration priority index, and obtaining differential data when the change in the collaboration priority index meets the differential update condition are as follows:

[0035] After the session data is sent to the target edge cloud node, the collaboration priority index is recalculated, and the change in the collaboration priority index is calculated.

[0036] When the change in the collaboration priority index is not less than the differential trigger threshold, the differential update condition is determined to be met, and the structured information that has changed compared to the session data is encapsulated into session differential data.

[0037] When the differential update conditions are not met, the existing session data state is maintained and no new session differential data is sent.

[0038] As a preferred embodiment of the doll voice call method based on edge-cloud collaboration described in this invention, the following steps are taken: The differential data is uploaded to the target edge cloud node; the target edge cloud node forms a session execution record based on the collaboration priority index and the differential data; multiple session execution records are aggregated and a unified collaboration parameter set is updated.

[0039] Upload the differential data to the target edge cloud node, receive the collaborative priority index, differential data, network scheduling and resource usage in this session to generate a session running record, update the initial unified collaborative parameter set, and obtain the unified collaborative parameter set;

[0040] Session execution records are aggregated and managed, execution metrics are calculated, and the unified collaborative parameter set is adaptively updated based on the execution metrics.

[0041] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the doll voice call method based on end-to-cloud collaboration as described in the first aspect of the present invention.

[0042] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the doll voice call method based on end-to-cloud collaboration as described in the first aspect of the present invention.

[0043] The beneficial effects of this invention are as follows: A cloud-end collaborative control mechanism based on session collaboration priority index to generate bandwidth, latency, and update frequency budgets achieves dynamic matching of terminal upload scale, processing complexity, and network scheduling capabilities during voice calls; a privacy-controlled pruning mechanism based on hierarchical mapping of session priority tags and data modality allowable sets ensures real-time emergency communication while suppressing voice content leakage in low-priority and privacy-sensitive scenarios; and a closed-loop collaborative mechanism based on changes in session collaboration priority index to trigger differential reporting and adjust service quality in a coordinated manner ensures that network and computing resource reconfiguration is triggered only when the importance of the session changes effectively. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart of a doll voice call method based on edge-cloud collaboration.

[0046] Figure 2 A flowchart generated for coordination priorities.

[0047] Figure 3 A flowchart for generating collaborative call policy tokens and reporting session data.

[0048] Figure 4 This is a flowchart for closed-loop optimization of session differential updates and unified collaborative parameter sets. Detailed Implementation

[0049] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0050] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0051] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0052] Reference Figures 1-4 This is one embodiment of the present invention, which provides a doll voice calling method based on end-to-cloud collaboration, including the following steps:

[0053] S1. Bind the mobile terminal to the doll terminal, configure preference weights, and generate an initial unified collaborative parameter set based on the preference weights.

[0054] Guardians bind their mobile devices to the doll's terminal one-to-one, initialize the doll's configuration, and configure contact information, including contact identifier, contact method, and whether emergency calls are allowed. They also configure preference weights, including three weights: privacy preference weight, response speed preference weight, and network resource preference weight. For example, a higher privacy preference weight means less voice transmission and more structured information transmission; a higher response speed preference weight means that even if more data is transmitted, the connection will be established as quickly as possible; and a higher network resource preference weight means that voice transmission should be minimized and reporting reduced. Finally, they configure privacy settings, such as whether to allow the uploading of encrypted voice clips and de-identified text.

[0055] Three preference weights are extracted from the guardian configuration. The three preference weights are normalized to obtain three core weight coefficients: privacy preference weight coefficient, response speed preference weight coefficient, and network resource preference weight coefficient. The sum of the three core weight coefficients is made to one, and the three core weight coefficients are used as the initial unified collaborative parameter set.

[0056] S2. After detecting a wake-up trigger, the doll terminal collects voice data, performs end-side intent recognition, urgency recognition, and privacy sensitivity recognition, generates a collaborative priority index based on the initial unified collaborative parameter set, and determines the session priority label.

[0057] When the doll terminal detects that the energy of consecutive voice frames exceeds the preset acoustic threshold and the duration exceeds the preset wake-up time, it determines that the wake-up is triggered.

[0058] It should be noted that the preset acoustic threshold is obtained by collecting ambient audio data and statistically analyzing the long-term mean and standard deviation of the audio energy, and then using the sum of the mean and the multiple of the standard deviation as the acoustic threshold; the preset wake-up duration is obtained by statistically analyzing the minimum stable duration of a continuous effective speech segment in real speech trigger samples, and then using the minimum stable duration as the wake-up duration.

[0059] When a wake-up trigger is detected, the current speech signal is digitally sampled to obtain the original speech sequence. The original speech sequence is then decoded to obtain the structured intent category. Specifically, the original speech sequence is converted into a text command string, which is then segmented to obtain a word sequence. This sequence is then matched against a pre-defined intent phrase template library. If a template match is successful, the intent category bound to the template is used as the structured intent. If no template match is found directly, the appellations and action words in the word sequence are combined and matched according to a pre-defined keyword dictionary and relational word rules. The combined word sequence is then matched against the pre-defined intent phrase template library again to obtain the structured intent category. An intent quantization mapping table is set up, pre-configuring a corresponding intent weight value for each intent category. The intent weight values ​​for different intent categories are monotonically arranged according to the strength of the intent's influence on the importance and urgency of the conversation. The structured intent categories are mapped using the intent quantization mapping table to obtain the structured intent. Finally, the urgency of the speech energy and the speech rate are calculated from the original speech sequence. The urgency level is quantified by taking the maximum value among the urgency levels of speech energy, speech rate, and keyword urgency. Specifically, the ratio of the average frame energy of the current speech to the average background noise energy obtained during the standby phase is used as the speech energy urgency. The speech energy urgency is then segmented and mapped to 0 to 1 according to a preset energy range. The number of words or syllables per unit time is compared with a preset normal speech rate range. When the normal speech rate exceeds the upper limit, it is mapped to 0 to 1 according to the excess ratio to obtain the speech rate urgency. A preset urgent keyword table is searched in the text command string. When an urgent keyword is matched, the keyword urgency is set to 1 directly; otherwise, it is set to 0. A preset privacy keyword table is searched word by word in the text command string. When a keyword involving contact information or account and password information is detected, the privacy sensitivity quantification value is set to 1 directly. When only keywords involving identity information or location information are detected, the privacy sensitivity quantification value is set to 0.5. When no privacy keywords are detected, the privacy sensitivity quantification value is set to 0.

[0060] It should be noted that the preset intent phrase template library is constructed through historical children's voice command samples. The standard intent phrase templates are generated by manually deduplicating and merging synonyms in the speech-to-text results and classifying them according to business functions. The preset energy range is determined by statistically obtaining the overall distribution range of speech frame energy, using the concentrated range of energy distribution as the medium energy range, the energy range below the concentrated range as the low energy range, and the energy range above the concentrated range as the high energy range. The preset normal speech rate range is determined by statistically analyzing the number of words or syllables per unit time of speech samples to obtain the overall distribution of speech rate, and using the concentrated range of the overall distribution as the normal speech rate range.

[0061] The collaboration priority index is calculated using a unified set of collaboration parameters, structured intent, urgency quantification, and privacy sensitivity quantification. The expression is as follows:

[0062] ;

[0063] in, Indicates the priority index for collaboration. This represents a quantitative value indicating the urgency level. Represents the network state quantization value. This represents a privacy-sensitive metric. This represents the response speed preference weighting coefficient. This represents the weighting coefficient of network resource preference. This represents the privacy preference weighting coefficient.

[0064] Set priority label mapping rules, when When the value is zero, the current session is identified as a privacy-suppressed session label; when... When the value is greater than zero and less than or equal to the first priority threshold, the current session is determined to be a low-priority session label; when... When the priority threshold is greater than the first priority threshold but less than or equal to the second priority threshold, the current session is identified as a medium-priority session; when... If the priority value is greater than the second priority threshold, the current session is determined to be a high-priority session and the session priority label is obtained.

[0065] It should be noted that the first priority threshold and the second priority threshold are determined by pre-constructing a test voice sample set covering various daily dialogue scenarios, notification scenarios, and emergency help scenarios, and labeling each sample with a corresponding business response level. A collaborative priority index is calculated for each test sample, and the collaborative priority index is statistically correlated with the manually labeled business response levels to determine the boundary interval between adjacent business response levels. The boundary point used to distinguish between normal response and enhanced response is determined as the first priority threshold, and the boundary point used to distinguish between enhanced response and emergency response is determined as the second priority threshold.

[0066] S3. The collaborative priority index and session priority label are reported to the core network. The core network selects the target edge cloud based on the unified collaborative parameter set and generates a collaborative call policy token through the target edge cloud.

[0067] The coordination priority index and session priority tag are encapsulated into a session coordination summary message and reported to the core network. After receiving the session coordination summary message, the core network calculates the coordination adaptation score for each candidate edge cloud node from the set of candidate edge cloud nodes corresponding to the terminal's current access area, and selects the edge cloud node with the highest score as the target edge cloud node. The expression is as follows:

[0068] ;

[0069] in, Indicates the first Collaborative adaptation score of candidate edge cloud nodes Indicates the first The current load of each candidate edge cloud node. Indicates the terminal to the number The current link latency of each candidate edge cloud node. Indicates the load impact factor. This represents the time delay impact coefficient.

[0070] It should be noted that, It obtains the current processor usage, memory consumption, and the number of concurrent voice-related tasks being processed through calibration; It is obtained by the core network by sending latency detection messages to obtain the round-trip transmission latency from the terminal to each candidate edge cloud node, and then normalizing the round-trip transmission latency. The final load impact coefficient is selected by comparing the scheduling success rate and latency stability under different load impact coefficient values. The delay impact coefficient is selected by comparing the session establishment success rate, end-to-end delay stability, and packet loss corresponding to different delay impact coefficients, and selecting the coefficient value corresponding to the optimal overall communication quality.

[0071] Using the session priority tag, the corresponding target service quality level is selected from the preset service quality policy table. When the session priority tag is privacy suppression level, the corresponding target service quality level is control signaling level; when the session priority tag is low priority, the corresponding target service quality level is normal voice level; when the session priority tag is medium priority, the corresponding target service quality level is enhanced voice level; and when the session priority tag is high priority, the corresponding target service quality level is emergency voice level.

[0072] On the target edge cloud node, the resource budget parameters on the edge are calculated using the collaborative priority index, expressed as follows:

[0073] ;

[0074] ;

[0075] ;

[0076] in, Indicates the uplink bandwidth budget. The session priority label is The corresponding reference bandwidth parameters, This represents the bandwidth adjustment coefficient. Indicates the delay budget. The session priority label is The corresponding reference delay parameters, This represents the time delay adjustment coefficient. Indicates the budget for state update frequency. The session priority label is The corresponding reference state update frequency parameters, This represents the state update frequency adjustment coefficient.

[0077] It should be noted that, It is based on the corresponding session priority tag The minimum coding bit rate requirement for the next voice data and the minimum bandwidth requirement for control signaling are determined together. It is based on the corresponding session priority tag The maximum allowable end-to-end queuing delay and the upper limit of single voice processing delay on the edge side are combined to determine the delay. It is based on the corresponding session priority tag The minimum perception period requirement for session state changes is used to deduce the maximum allowed number of state updates per unit time. The bandwidth adjustment coefficient, latency adjustment coefficient, and state update frequency adjustment coefficient are determined by using the lowest session priority corresponding to the baseline parameter as the lower limit and the highest session priority corresponding to the baseline parameter as the upper limit on the cloud side. This ensures that the effect of the adjustment coefficients can cover the entire range of resource changes from the lowest to the highest level. As a result, it ensures that as the session coordination priority index changes from low to high, the uplink bandwidth budget can be continuously increased from small to large, the end-to-edge processing latency budget can be continuously tightened from large to small, and the state update frequency budget can be continuously increased from low to high.

[0078] The uplink bandwidth budget, latency budget, and status update frequency budget are combined to obtain the end-side resource budget parameter set; the target edge cloud node, target service quality level, and end-side resource budget parameter set are combined to obtain the collaborative call policy token.

[0079] S4. The doll terminal selects the data modality through the collaborative call strategy token, forms session data and sends it to the target edge cloud, updates the collaborative priority index, calculates the change in the collaborative priority index, and obtains differential data when the change in the collaborative priority index meets the differential update condition.

[0080] Using a session priority tag, users can select the allowed data modal combinations and their upload limits from a preset set of allowed data modalities. When the session priority tag is at the privacy suppression level, the allowed data modal set is limited to control status identifier data, session trigger flag data, and device operation status summary data; uploading any raw voice data, environmental audio data, or user-expressed content data is not allowed. When the session priority tag is at a low priority level, the allowed data modal set is limited to structured intent identifier data, contact logical identifier data, and device status summary data; uploading raw voice waveform data and continuous environmental audio data is not allowed. When the session priority tag is at a medium priority level, the allowed data modal set is limited to structured intent identifier data, contact logical identifier data, and environmental status data. Summary data and limited-length voice feature data are not allowed to be uploaded as complete raw voice waveform data. When the session priority tag is high priority, the data modality set is limited to structured intent identifier data, contact logical identifier data, environmental status summary data, limited-length voice feature data, and privacy-processed raw voice data to support real-time processing of emergency voices. During the construction process, the doll terminal determines the upper limit of the duration of the raw voice segments that can be carried in the initial report of a single session based on the uplink bandwidth budget in the collaborative call policy token. Within the duration limit, the voice segments that have been desensitized are extracted. The processing complexity of local voice desensitization and feature extraction is limited based on the end-to-edge processing latency budget, and structured intent identifier data and contact logical identifier data are retained first.

[0081] During a call, the doll terminal calculates the state update frequency budget based on the budget, converting it into the minimum allowed update interval for the current session. It then periodically initiates session state collection and coordination priority index updates according to this minimum update interval, recalculates the coordination priority index, and calculates the change in the coordination priority index. The expression for this is:

[0082] ;

[0083] in, Indicates the current The relative change of the coordination priority index at any given time compared to the initial coordination priority index. Indicates the current Time-of-use coordination priority index This indicates the coordination priority index corresponding to the time the call is established. It represents a very small positive constant to prevent the denominator from being zero.

[0084] When the relative change is not less than the differential trigger threshold, the differential update condition is determined to be met, and the structured information that has changed compared to the session data is encapsulated into session differential data; when the differential update condition is not met, the original session data state is maintained, and no new session differential data is sent.

[0085] It should be noted that the differential trigger threshold is to divide the effective range of the collaborative priority index corresponding to each session priority label in ascending order, and use the minimum interval span between any two adjacent priority intervals as the differential trigger threshold corresponding to the priority level, so that when the relative change of the collaborative priority index is sufficient to cross the boundary of the current priority interval, a differential update will be triggered.

[0086] S5. Upload the differential data to the target edge cloud. The edge cloud and the core network form a session operation record based on the collaboration priority index and the differential data. Aggregate multiple session operation records and update the unified collaboration parameter set.

[0087] After uploading the differential data to the target edge cloud, the original voice frame file and temporary cache are completely deleted, and only the summary information structure associated with this session is retained. The summary information structure includes at least the session identifier, the initial coordination priority index, the coordination priority index at the end of the call, the session priority tag, and the session start and end timestamps.

[0088] When a voice call is released, the target edge cloud and the core network generate a session operation record based on the coordination priority index, session differential data, network scheduling and resource usage received during the session. The session priority label, differential trigger threshold, baseline bandwidth parameters, baseline latency parameters, baseline state update frequency parameters, bandwidth adjustment coefficient, latency adjustment coefficient and state update frequency adjustment coefficient are put into the initial unified coordination parameter set to form a unified coordination parameter set.

[0089] The cloud side aggregates and manages session operation records according to the version number of the unified collaborative parameter set. Multiple records with the same version number are selected from the session operation record library, and these records are divided into multiple statistical batches. For each batch, several operation indicators related to the effectiveness of end-cloud-network collaborative control are calculated, such as the session service quality compliance rate, resource budget utilization deviation, differential update trigger rate, and the frequency of session upgrades and downgrades.

[0090] To achieve closed-loop optimization of the edge-cloud-network collaborative control parameters, adaptive updates are performed on the adjustable collaborative control parameters in the unified collaborative parameter set, expressed as follows:

[0091] ;

[0092] in, This indicates the updated values ​​of the coordinated control parameters. This indicates the original value of the corresponding collaborative control parameter in the current version's unified collaborative parameter set. This indicates the step size coefficient for parameter updates. This represents the expected target value of the target operating index corresponding to the coordinated control parameters. This represents the actual observed values ​​of the operational metrics obtained through statistics from multiple session runs under the current version's unified collaborative parameter set.

[0093] It should be noted that, Based on the service quality constraints required by the service quality level corresponding to different session priority labels, such as the minimum voice clarity, maximum acceptable latency, and minimum reliability guarantee level, the target operating range that such sessions should achieve under ideal operating conditions is determined. Combined with the upper limit of the carrying capacity of the access network, core network, and edge cloud, the upper limit constraint correction of the target operating range is performed, and a representative target value is selected as the expected target value within the corrected target operating range. The maximum allowable adjustment range of the collaborative control parameters under safety constraints and the version update cycle of the unified collaborative parameter set are jointly determined to limit the maximum change of the collaborative control parameters in a single version update.

[0094] This embodiment also provides a computer device applicable to the doll voice call method based on end-to-cloud collaboration, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the doll voice call method based on end-to-cloud collaboration as proposed in the above embodiment.

[0095] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0096] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the puppet voice call method based on end-to-cloud collaboration as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0097] In summary, this invention achieves dynamic matching of terminal upload scale, processing complexity, and network scheduling capabilities during voice calls through an end-to-cloud collaborative control mechanism that generates bandwidth, latency, and update frequency budget based on a session collaboration priority index; it also achieves privacy-controlled pruning through a privacy-controlled pruning mechanism based on hierarchical mapping of session priority tags and data modality allowable sets, ensuring the real-time performance of emergency communications while suppressing the leakage of voice content in low-priority and privacy-sensitive scenarios; and it achieves network and computing resource reconfiguration only when the importance of the session changes effectively, through a closed-loop collaborative mechanism that triggers differential reporting based on changes in the session collaboration priority index and adjusts service quality accordingly.

[0098] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for voice-calling a doll based on end-to-end cloud collaboration, characterized in that: include, S1. Bind the mobile terminal to the doll terminal, configure preference weights, and generate an initial unified collaborative parameter set based on the preference weights; S2. After detecting a wake-up trigger, the doll terminal collects voice data, performs end-side intent recognition, urgency recognition, and privacy sensitivity recognition, generates a collaborative priority index based on the initial unified collaborative parameter set, and determines the session priority label. S3. The collaborative priority index and session priority label are reported to the core network. The core network selects the target edge cloud node according to the unified collaborative parameter set and generates a collaborative call policy token through the target edge cloud node. S4. The doll terminal selects the data modality through the collaborative call strategy token, forms session data and sends it to the target edge cloud, updates the collaborative priority index, calculates the change in the collaborative priority index, and obtains differential data when the change in the collaborative priority index meets the differential update condition. S5. Upload the differential data to the target edge cloud node. The target edge cloud node forms a session execution record based on the collaboration priority index and the differential data. It then aggregates multiple session execution records and updates the unified collaboration parameter set. The specific steps of step S1 are as follows: bind the mobile terminal to the doll terminal, initialize the configuration of the doll terminal, and configure the privacy preference weight, response speed preference weight and network resource preference weight. The privacy preference weight, response speed preference weight, and network resource preference weight are normalized to generate privacy preference weight coefficient, response speed preference weight coefficient, and network resource preference weight coefficient, which serve as the initial unified collaborative parameter set. The specific steps of step S2 are as follows: when the doll terminal detects the wake-up trigger, it digitally samples the voice signal to generate the original voice sequence. The original speech sequence is decoded, and template matching is performed using a pre-defined intent phrase template library to obtain the structured intent. Using the original speech sequence, the urgency of speech energy, the urgency of speech rate, and the urgency of keywords are calculated, and the maximum value among these three factors is taken as the quantification value of urgency. The original speech sequence is converted into a text command string. A pre-set privacy keyword list is searched word by word in the text command string, and a privacy sensitivity quantification value is set. The collaboration priority index is calculated based on the initial unified collaboration parameter set, structured intent, urgency quantification, and privacy sensitivity quantification.

2. The doll voice calling method based on end-to-cloud collaboration as described in claim 1, characterized in that: The steps involve reporting the collaboration priority index and session priority label to the core network, the core network selecting a target edge cloud node based on a unified collaboration parameter set, and generating a collaboration call policy token through the target edge cloud node. The collaboration priority index and session priority label are reported to the core network. The collaboration adaptation score is calculated for each candidate edge cloud node, and the edge cloud node with the highest score is selected as the target edge cloud node. Select the corresponding target service quality level from the preset service quality policy table by using the session priority label; On the target edge cloud node, the resource budget parameters on the edge side are calculated using the collaborative priority index. Combine the target edge cloud node, target service quality level, and edge resource budget parameters to obtain the collaborative call policy token.

3. The doll voice calling method based on end-to-cloud collaboration as described in claim 2, characterized in that: The doll terminal selects a data modality through a collaborative call policy token, constructs session data, and sends it to the target edge cloud. The specific steps are as follows: Using the session priority tag, select the allowed combination of data modalities for upload from the preset set of allowed data modalities; The maximum duration of the original voice segments that can be carried is determined based on the collaborative call policy token. Within the maximum duration, voice segments that have been desensitized are extracted to limit the processing complexity of local voice desensitization and feature extraction, thereby forming session data and sending it to the target edge cloud.

4. The doll voice calling method based on end-to-cloud collaboration as described in claim 3, characterized in that: The updated collaboration priority index involves calculating the change in the collaboration priority index. When the change in the collaboration priority index meets the differential update condition, the differential data is obtained. The specific steps are as follows: After the session data is sent to the target edge cloud node, the collaboration priority index is recalculated, and the change in the collaboration priority index is calculated. When the change in the collaboration priority index is not less than the differential trigger threshold, the differential update condition is determined to be met, and the structured information that has changed compared to the session data is encapsulated into session differential data. When the differential update conditions are not met, the existing session data state is maintained and no new session differential data is sent.

5. The doll voice calling method based on end-to-cloud collaboration as described in claim 4, characterized in that: The process involves uploading the differential data to the target edge cloud node, where the target edge cloud node generates a session execution record based on the collaboration priority index and the differential data. Multiple session execution records are then aggregated and a unified collaboration parameter set is updated. The specific steps are as follows: Upload the differential data to the target edge cloud node, receive the collaborative priority index, differential data, network scheduling and resource usage in this session to generate a session running record, update the initial unified collaborative parameter set, and obtain the unified collaborative parameter set; Session execution records are aggregated and managed, execution metrics are calculated, and the unified collaborative parameter set is adaptively updated based on the execution metrics.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the doll voice call method based on end-to-cloud collaboration as described in any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the doll voice call method based on end-to-cloud collaboration as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice interaction method, device, system and equipment for protecting privacy and storage medium

    CN113472806A

  • First-aid method and system based on 5G communication and edge device

    CN120499732A