Intercom communication delay optimization method and system for mobile users
By using a walkie-talkie communication delay optimization method and system, noise reduction and echo cancellation are performed on the input voice, semantic units are dynamically divided, and transmission channel parameters and confidence indicators are configured differently according to the scenario mode. This solves the transmission delay problem of key voice information when network resources are limited, and improves the communication efficiency of walkie-talkies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN AUGOO COMM EQUIP CO LTD
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-31
AI Technical Summary
In existing walkie-talkie communications for mobile users, the passive adaptation to network fluctuations and the indiscriminate processing of all communication content by traditional vocoders result in the inability to prioritize the transmission of critical voice information when network resources are limited, thus affecting communication efficiency.
By performing noise reduction and echo cancellation on the input speech, dynamically dividing semantic units, and configuring transmission channel parameters and confidence indicators differently according to scene modes, low-latency and high-reliability transmission of key speech information can be achieved.
While ensuring low-latency and highly reliable delivery of critical voice information, it intelligently maintains the overall fluency of the conversation, thereby improving the efficiency of walkie-talkie communication.
Smart Images

Figure CN122496844A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication technology, specifically to a method and system for optimizing communication delay of walkie-talkies for mobile users. Background Technology
[0002] Current walkie-talkie communications for mobile users generally employ passive optimization strategies to adapt to network fluctuations. These strategies combat network uncertainty by increasing buffering to mitigate jitter and addressing packet loss during retransmissions. Essentially, this trades transmission reliability for increased latency; larger buffers mean longer playback wait times, retransmission mechanisms introduce at least one round-trip delay, and fixed redundant coding constantly consumes additional bandwidth. Furthermore, traditional encoders are commonly used in the voice encoding and decoding stages, indiscriminately encoding and transmitting all voice content. This limits voice quality and naturalness at low bit rates. When network congestion or signal fluctuations occur, all voice data packets are forced to endure the same level of buffer accumulation or retransmission wait times, significantly increasing end-to-end latency and causing severe fluctuations. Lacking the ability to perceive and differentiate the semantic importance of voice information, critical and non-critical information are treated equally in the competition for network resources. When network bandwidth is strained or the packet loss rate increases, the transmission latency of critical information is also indiscriminately increased, and it may even be lost due to retransmission timeouts or buffer overflows, directly impacting communication efficiency.
[0003] In summary, existing technologies suffer from the technical problem that, due to the passive adaptation to network fluctuations and the fact that traditional vocoders process all communication content indiscriminately, trading time for reliability, the transmission delay of critical voice information cannot be prioritized when network resources are limited, further affecting the communication efficiency of walkie-talkies. Summary of the Invention
[0004] The purpose of this application is to provide a method and system for optimizing communication delays in walkie-talkies for mobile users, in order to solve the technical problem in the prior art where the transmission delay of key voice information cannot be prioritized when network resources are limited due to the passive adaptation to network fluctuations and the fact that traditional vocoders process all communication content indiscriminately, trading time for reliability. This further affects the efficiency of walkie-talkie communication.
[0005] To achieve the above objectives, this application provides a method and system for optimizing walkie-talkie communication delay for mobile users.
[0006] Firstly, this application provides a method for optimizing walkie-talkie communication latency for mobile users. This method is implemented through a walkie-talkie communication latency optimization system for mobile users. The method includes: feeding input speech sequentially into a noise reduction unit and an echo cancellation unit to filter out background noise and echo, obtaining clean input speech with a high signal-to-noise ratio; continuously scoring the clean input speech, dynamically dividing the speech stream into semantic units based on continuous score changes and semantic triggering features, and identifying and configuring the current scene's intercom mode; setting entry and exit thresholds based on the current scene's intercom mode, and determining the transmission type for each semantic unit, where the transmission type includes a main transmission unit, a summary transmission unit, an event transmission unit, or a suppression unit; configuring transmission channel parameters based on the transmission type and the current scene's intercom mode, where the transmission channel parameters include heterogeneous coding parameters, fidelity enhancement parameters, hierarchical scheduling parameters, redundancy allocation parameters, and packet loss recovery parameters; and attaching a confidence level identifier to the transmitted information based on the transmission channel parameters, with the receiving end performing adaptive playback control based on the confidence level identifier.
[0007] Optionally, the transmitting end performs continuous frame-level semantic importance scoring on the clean input speech according to a fixed time window, generating a scoring time series; based on the time evolution trend of the scoring values in the scoring time series and predefined semantic triggering features, the speech stream is dynamically segmented to construct semantic units. When a semantic triggering feature is detected, a preset number of frames are traced back and a preset number of frames are extended backward from the trigger frame to construct a command protection zone. Speech frames within the protection zone have a higher retention priority than frames outside the protection zone in subsequent decisions; based on automatic reasoning of user explicit input or command word density, context marker words, and historical communication tolerance rate, the current scene intercom mode is identified and configured. The scene intercom mode includes at least a continuous naturalness priority mode and a key slot fidelity priority mode.
[0008] Optionally, the semantic triggering feature includes at least one of command words, numeric words, locative words, object words, or call sign words. When the score values of multiple consecutive frames in the scoring time series show a monotonically increasing trend and exceed the starting threshold, the delineation of the semantic unit's starting boundary is triggered; when the score value sequence shows a monotonically decreasing trend and remains below the ending threshold, the delineation of the ending boundary is triggered. If at least one semantic triggering feature exists between the starting boundary and the ending boundary, then, taking the frame where the triggering feature is located as the center, a first preset number of frames are traced back and a second preset number of frames are extended backward to form a command protection zone. All frames within the protection zone are then weighted and fused with the original scoring sequence to ensure that the protected zone is maintained. The effective score of a frame within the protected area is not less than a predetermined percentage of the maximum score of a frame outside the protected area. Specifically, when a local minimum occurs in the score value sequence and the scores of the preceding and following frames are significantly higher than the minimum, the frame containing the local minimum is considered a natural segmentation candidate point within the semantic unit. Based on the temporal distance between adjacent candidate points and a preset minimum unit duration constraint, it is determined whether to split a single semantic unit into multiple sub-units. If the score value sequence remains flat within a preset time window and has no semantic triggering features, consecutive frames within the preset time window are merged into a non-instruction semantic unit, whose transmission type is preset as a candidate for either a suppression unit or a digest transmission unit.
[0009] Optionally, the system obtains the explicit mode selection provided by the user through a physical interface or voice command, and determines the current scene intercom mode based on the user's explicit input. When there is no explicit mode selection, the command word density, context identifier word detection results, and historical communication tolerance rate are input into a preset decision mapping library, and a scene mode confidence vector is output. When the confidence of the key slot fidelity priority mode in the scene mode confidence vector exceeds a first threshold and is consistent in N consecutive judgments, the current scene mode is configured as the key slot fidelity priority mode. When the confidence of the continuous naturalness priority mode exceeds a second threshold and is consistent in N consecutive judgments, it is configured as the continuous naturalness priority mode, wherein the first threshold is higher than the second threshold to construct hysteresis characteristics.
[0010] Optionally, an importance metric for semantic units is calculated, which is determined based on the weighted average of frame scores within the unit and the presence of an instruction protection zone. When the importance metric of a semantic unit is not lower than the entry threshold, the corresponding unit is marked as a main transmission unit and enters the main transmission active state. When the importance metrics of multiple consecutive semantic units are all lower than the exit threshold, the main transmission active state is exited. For semantic units not marked as main transmission units, they are judged sequentially according to the unit content features. If they contain non-verbal human voice events, they are marked as event transmission units. Otherwise, if they contain at least one key semantic slot, they are marked as summary transmission units, and the rest are marked as suppression units. The key semantic slots are extracted using a slot filling model based on conditional random fields, and the slot filling model shares the underlying feature extraction network with the semantic scoring model.
[0011] Optionally, according to the transmission type, corresponding heterogeneous coding parameters are matched, wherein the main transmission unit adopts neural low-bit-rate speech coding, the summary transmission unit extracts key slots for packaging, the event transmission unit performs parameterized coding, the suppression unit only generates time placeholders, and the coding bit rate and frame length are adaptively adjusted according to the scene mode; according to the current scene intercom mode, the corresponding mode's fidelity enhancement parameters are configured, wherein, in the key slot fidelity priority mode, double redundant transmission and fast acknowledgment retransmission are enabled for the main transmission unit within the command protection zone; in the continuous naturalness priority mode, a disabling retransmission mechanism is enabled; according to the transmission type, the corresponding heterogeneous coding parameters are matched ... neural low-bit-rate speech coding is adopted, the summary transmission unit extracts key slots for packaging, the event transmission unit performs parameterized coding, the suppression unit only generates time placeholders, and the coding bit rate and frame length are adaptively adjusted according to the scene mode. The transmission characteristics of the transmission type are configured with hierarchical scheduling parameters and redundancy allocation parameters. Among them, the sending queue is dynamically scheduled according to semantic priority and historical backlog value, and high-priority packets can preempt low-priority packets; forward error correction redundancy is allocated according to semantic loss cost and predictability recoverability, and high redundancy is allocated to content with high loss cost and difficult recovery; corresponding packet loss recovery parameters are configured according to the current intercom mode. In the continuous naturalness priority mode, feature extrapolation, semantic template reconstruction or duration-maintaining path is used for packet loss prediction and recovery; in the critical slot fidelity priority mode, prediction is disabled and timed waiting and silent position are used.
[0012] Optionally, when the current intercom mode is the critical slot fidelity priority mode, the transmission type is the main transmission unit, and the semantic unit contains an instruction protection zone, the double redundancy transmission and fast acknowledgment retransmission mechanism is activated. The sending end generates two independently encoded redundant packets for the same voice content and transmits them through two physical links or time-division multiplexing of the same link. The fast acknowledgment timeout threshold is set to 0.5 times the current round-trip delay and the minimum value is not lower than the minimum threshold. If no acknowledgment signal is received within the timeout threshold, the data packet is immediately retransmitted, and the retransmitted data packet has the highest transmission priority.
[0013] Optionally, based on the configuration results of the transmission channel parameters, content confidence and timing confidence are calculated for each data packet, and confidence is identified. The content confidence is determined based on the transmission type and corresponding heterogeneous coding parameters, and the timing confidence is determined based on the priority and link state parameters in the hierarchical scheduling parameters. The receiving end adjusts the jitter buffer depth according to the timing confidence and adjusts the prediction replacement strategy according to the content confidence. When the timing confidence is high, the buffer is shallower to reduce playback latency, and when the timing confidence is low, the buffer is deepened to tolerate network jitter. When the content confidence is high, the prediction result is played directly, and when the content confidence is low, the prediction is suppressed and a mute is returned to placeholder. When the real data packet arrives after the prediction recovery result has been played, the subsequent unplayed decoded output is smoothed and corrected. If the delayed data packet corresponds to an instruction protection zone, the semantic boundary of the output command is prohibited from being changed.
[0014] Optionally, the receiving end periodically obtains one or more feedback indicators, including end-to-end delay, packet loss rate, number of critical slot recovery failures, command deadline default rate, prediction recovery confidence, and retransmission success rate, and transmits them back to the sending end via the reverse channel. The sending end switches between continuous naturalness priority target and critical slot fidelity priority target according to the current intercom mode, and dynamically adjusts one or more parameters based on the feedback indicators, including semantic unit segmentation threshold, command protection zone backtracking and extension length, transmission type decision boundary, encoded frame length and code rate, forward error correction redundancy allocation, priority mapping rules, and retransmission timeout threshold in fidelity mode.
[0015] Secondly, this application also provides a walkie-talkie communication delay optimization system for mobile users, used to execute the walkie-talkie communication delay optimization method for mobile users as described in the first aspect, wherein the walkie-talkie communication delay optimization system for mobile users includes: a noise reduction processing module, used to feed the input speech sequentially into a noise reduction unit and an echo cancellation unit to filter out background noise and echo, thereby obtaining clean input speech with a high signal-to-noise ratio; a semantic unit segmentation module, used to continuously score the clean input speech, dynamically segment the speech stream into semantic units based on continuous score changes and semantic triggering features, and identify and configure the current scene's walkie-talkie mode; and a transmission type determination module. The module is used to set entry and exit thresholds based on the current intercom mode and to determine the transmission type for each semantic unit. The transmission type includes a main transmission unit, a summary transmission unit, an event transmission unit, or a suppression unit. The transmission channel parameter configuration module is used to configure transmission channel parameters according to the transmission type and the current intercom mode. The transmission channel parameters include heterogeneous coding parameters, fidelity enhancement parameters, hierarchical scheduling parameters, redundancy allocation parameters, and packet loss recovery parameters. The confidence level identification module is used to attach a confidence level identification to the transmitted information based on the transmission channel parameters. The receiving end performs adaptive playback control based on the confidence level identification.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: by continuously semantically scoring the voice stream and dynamically dividing it into semantic units, setting thresholds according to scene modes to determine the transmission type, configuring transmission channel parameters differently according to the transmission type and scene mode, and adding confidence labels to the transmitted information to drive the receiver to perform adaptive playback control, while ensuring low latency and high reliability delivery of key semantic information, the overall fluency of the dialogue is intelligently maintained, thereby improving the communication efficiency of the walkie-talkie. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the walkie-talkie communication delay optimization method for mobile users proposed in this application.
[0018] Figure 2This is a schematic diagram of the walkie-talkie communication delay optimization system for mobile users, as described in this application.
[0019] Figure labeling: Noise reduction processing module 11, semantic unit segmentation module 12, transmission type determination module 13, transmission channel parameter configuration module 14, confidence level identification module 15. Detailed Implementation
[0020] This application provides a method and system for optimizing walkie-talkie communication latency for mobile users. It addresses the technical problem in existing technologies where passively adapting to network fluctuations and the indiscriminate processing of all communication content by traditional vocoders—a trade-off of time for reliability—lead to a failure to prioritize the transmission of critical voice information when network resources are limited, further impacting walkie-talkie communication efficiency. By continuously semantically scoring the voice stream and dynamically dividing it into semantic units, setting thresholds based on scene modes to determine the transmission type, configuring transmission channel parameters differently according to the transmission type and scene mode, and adding confidence markers to the transmitted information to drive adaptive playback control at the receiver, this approach ensures low-latency, high-reliability delivery of critical semantic information while intelligently maintaining overall dialogue fluency, thereby improving walkie-talkie communication efficiency.
[0021] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0022] Example 1, please refer to the appendix. Figure 1 This application provides a method for optimizing walkie-talkie communication latency for mobile users. The method is applied to a walkie-talkie communication latency optimization system for mobile users, and specifically includes the following steps: The input speech is sequentially fed into the noise reduction unit and the echo cancellation unit to filter out background noise and echo, thereby obtaining clean input speech with a high signal-to-noise ratio.
[0023] The system continuously scores the clean input speech, dynamically divides the speech stream into semantic units based on the changes in continuous scores and semantic trigger features, and identifies and configures the intercom mode for the current scene.
[0024] Furthermore, this application also includes the following steps: the transmitting end performs continuous frame-level semantic importance scoring on clean input speech according to a fixed time window, generating a scoring time series; based on the time evolution trend of the scoring values in the scoring time series and predefined semantic triggering features, the speech stream is dynamically segmented to construct semantic units, wherein, when a semantic triggering feature is detected, a preset number of frames are traced back and a preset number of frames are extended backward from the trigger frame to construct a command protection zone, and speech frames within the protection zone have a higher retention priority than frames outside the protection zone in subsequent decisions; based on automatic reasoning of user explicit input or command word density, context marker words, and historical communication tolerance rate, the current scene intercom mode is identified and configured, wherein the scene intercom mode includes at least a continuous naturalness priority mode and a key slot fidelity priority mode.
[0025] Furthermore, this application also includes the following steps: the semantic triggering feature includes at least one of command words, numeric words, locative words, object words, or call sign words; when the score values of multiple consecutive frames in the scoring time series show a monotonically increasing trend and exceed the starting threshold, the semantic unit starting boundary is triggered; when the score value sequence shows a monotonically decreasing trend and remains below the ending threshold, the termination boundary is triggered; if at least one semantic triggering feature exists between the starting boundary and the termination boundary, then, taking the frame where the triggering feature is located as the center, a first preset number of frames are traced back and a second preset number of frames are extended backward to form a command protection zone, and all frames within the protection zone are weighted with the original scoring sequence. The fusion process ensures that the effective score of frames within the protected area is not lower than a predetermined percentage of the maximum score of frames outside the protected area. Specifically, when a local minimum occurs in the score value sequence and the scores of the preceding and following frames are significantly higher than the minimum, the frame containing the local minimum is used as a natural segmentation candidate point within the semantic unit. Based on the time distance between adjacent candidate points and a preset minimum unit duration constraint, it is determined whether to split a single semantic unit into multiple sub-units. If the score value sequence remains flat within a preset time window and has no semantic triggering features, then consecutive frames within the preset time window are merged into a non-instruction semantic unit, whose transmission type is preset as a candidate for either a suppression unit or a digest transmission unit.
[0026] Furthermore, this application also includes the following steps: obtaining the explicit mode selection provided by the user through a physical interface or voice command, and determining the current scene intercom mode based on the user's explicit input; when there is no explicit mode selection, inputting the command word density, context identifier word detection results, and historical communication tolerance rate into a preset decision mapping library, and outputting a scene mode confidence vector; when the confidence of the key slot fidelity priority mode in the scene mode confidence vector exceeds a first threshold and is consistent N times consecutively, configuring the current scene mode as the key slot fidelity priority mode; when the confidence of the continuous naturalness priority mode exceeds a second threshold and is consistent N times consecutively, configuring it as the continuous naturalness priority mode, wherein the first threshold is higher than the second threshold to construct hysteresis characteristics.
[0027] Specifically, the transmitting end continuously samples the input speech according to a fixed time window length, with each time window corresponding to a speech frame, the frame length of which can be set to 20ms. The original signal contains a mixture of user speech and environmental noise, which is filtered out by a noise reduction unit. A Fast Fourier Transform is performed on each frame to convert the time-domain signal to the frequency domain, obtaining the amplitude spectrum and phase spectrum. To estimate the noise power spectrum of the current frame, a minimum tracking method is used, maintaining a sliding window of 1s in length, storing the amplitude values of each frequency point within the past 1s, and taking the minimum value of each frequency point within the window as the noise power estimate for that frequency point. Simultaneously, a speech activity detector is used to determine whether the current frame is a pure noise frame; if so, the noise estimate is updated more aggressively. Based on the noise power estimate, the posterior signal-to-noise ratio (SNR) and prior signal-to-noise ratio (SNR) of each frequency point are calculated. The prior SNR is obtained recursively from the estimated value of the previous frame and the posterior SNR of the current frame using a recursive averaging method. Then, based on the Wiener filtering principle, the gain coefficient for each frequency point is calculated, ranging between 0 and 1. The gain is close to 1 for frequencies with high signal-to-noise ratio (SNR) and close to 0 for frequencies with low SNR. The gain coefficient is multiplied by the original amplitude spectrum to obtain the denoised amplitude spectrum. The original phase spectrum is preserved, and an inverse Fourier transform is performed to recover the denoised speech frame in the time domain. The frames are then concatenated using an overlap-addition method to form a continuous denoised speech stream.
[0028] The denoised speech stream enters the echo cancellation unit. The echo cancellation unit requires two input signals: the denoised microphone signal, which may still contain echoes from distant speech; and the distant reference signal, i.e., the speech signal being played by the local speaker. Internally, the echo cancellation unit maintains an adaptive transverse filter, with the filter order covering the duration of a typical echo path, typically 32ms to 64ms. The filter coefficients are initialized to zero. For each sampling point or each frame of data, the filter calculates an estimated echo signal using a delayed version of the distant reference signal and subtracts this estimate from the denoised microphone signal to obtain the error signal, which is the echo-cancelled output signal. The filter coefficients are updated using a normalized least mean square algorithm, with a moderate step size factor to ensure a balance between convergence speed and steady-state error. The update formula uses the energy of the error signal and the reference signal to normalize the step size, ensuring stable convergence characteristics of the filter under different input amplitudes. To handle simultaneous two-way conversations (where the local and remote users speak at the same time), the energy ratio of the reference and error signals is detected. If the error energy is significantly larger than the reference energy, it indicates the presence of local speech, at which point filter coefficient updates are paused to prevent filter divergence. Furthermore, a nonlinear processor can be added to the error signal path to further suppress residual minor echoes, such as through center clipping or comfort noise injection, resulting in a cleaner output sound. The signal processed by the echo cancellation unit is the clean input speech. Background noise in the clean input speech has been significantly reduced, echo components have been effectively eliminated, and the signal-to-noise ratio meets the requirements for subsequent semantic scoring and coding.
[0029] For each speech frame, a speech recognition engine is used to obtain the confidence level and keyword information of the text content in that frame. The acoustic features of the frame, such as energy and zero-crossing rate, are analyzed to determine whether it is clear human voice, indistinct human voice, or noise. Semantic importance is scored based on the dialogue context, with the score mapped to a range of 0 to 1, where 0 represents completely unimportant and 1 represents important. The scores are categorized into at least four levels: key information level, ordinary statement level, non-verbal voice level, and environmental noise level. The scores of all frames are arranged chronologically to form a scoring time series, reflecting the trend of speech content importance changing over time.
[0030] The system monitors the numerical trend of the scoring time series in real time and simultaneously detects predefined semantic trigger features, including command words, numeric words, locative words, object words, or call sign words. Their presence indicates that the current speech has high instructiveness or criticality. When multiple consecutive frames show a monotonically increasing scoring trend, and these scores all exceed a pre-set starting threshold (e.g., 0.6), the first frame exceeding the threshold is identified as the starting boundary of the semantic unit. Monitoring of scoring changes continues. When the scoring sequence turns into a monotonically decreasing trend and remains below a pre-set ending threshold (e.g., 0.3), the frame before the last frame below the threshold is identified as the ending boundary of the semantic unit. The starting and ending thresholds are two preset scoring value limits. The starting threshold determines whether a semantic unit begins (i.e., the score continuously increases and exceeds this value), and the ending threshold determines whether a semantic unit ends (i.e., the score continuously decreases and falls below this value). This yields an original semantic unit from the starting boundary to the ending boundary.
[0031] Check if at least one semantic trigger feature exists within the boundary of the original semantic unit. If it exists, extract the frame position where the trigger feature is located. Centered on this frame, backtrack for a first preset number of frames and extend backward for a second preset number of frames, such as backtracking 500ms and extending 500ms, thus forming a command protection zone. For all frames within the protection zone, weighted and fused with a high-weight coefficient, the original score of each frame within the protection zone is multiplied by the weight factor and then averaged with the original score, ensuring that the effective score of each frame within the protection zone is not lower than 90% of the highest score among all frames outside the protection zone. This forces speech within the protection zone to have a higher preservation priority than speech outside the protection zone in subsequent decisions.
[0032] For a segmented semantic unit, if a frame's score is found to be lower than the scores of the two frames before and after it (forming a local minimum), and this minimum score is more than 30% lower than the scores of the frames before and after it, then that frame is marked as a natural segmentation candidate. The temporal distance between two adjacent candidate points is calculated and compared to a preset minimum unit duration constraint. If the temporal distance is greater than or equal to the preset minimum unit duration constraint, the original semantic unit is split into two independent sub-semantic units at that candidate point; if the temporal distance is less than the preset minimum unit duration constraint, the candidate point is ignored, and the unit remains intact. The minimum unit duration constraint is a preset minimum time length, such as 100ms. Only when the temporal distance between two natural segmentation candidate points is greater than or equal to this duration is it permissible to split the original semantic unit into multiple sub-units, avoiding the generation of excessively short and meaningless fragments.
[0033] A flat detection window is set up. A flat time window means that the score value remains relatively stable over a period of time, without a significant upward or downward trend, and there are no semantic trigger features within the window. If the change in the score value of all frames within this window does not exceed 0.1, and there are no semantic trigger features within the window, then these consecutive frames are merged into a non-instruction semantic unit, which does not contain key information, and its transmission type is preset as either a suppression unit (i.e., no transmission) or a summary transmission unit (i.e., only a small number of features are extracted as candidates).
[0034] Check whether the user has provided an explicit mode selection via physical interface or voice command. The physical interface can be a two-position toggle switch on the side of the device; toggling it upwards selects the continuous naturalness priority mode, and toggling it downwards selects the critical slot fidelity priority mode. Alternatively, it can be two virtual buttons on the touchscreen. The voice command module has a built-in voice keyword recognition module; when the user says "switch to natural mode" or "fidelity mode," it is considered an explicit selection. Once explicit input is detected, the current intercom mode is immediately configured to the user-specified mode, and the subsequent automatic reasoning process is skipped.
[0035] If the user does not make any explicit selection, the automatic inference process is initiated. Automatic inference consists of three steps: calculating command word density, detecting context marker words, and obtaining historical communication tolerance rate. A sliding time window is maintained, with a window duration of 3 seconds. For each received voice frame, it is determined whether the frame contains semantic trigger features, such as command words, numeric words, locative words, object words, or call sign words. The total number of occurrences of trigger features within the window is counted and divided by the window duration to obtain the command word density. The system checks whether any member of a predefined set of context marker words appears within the same sliding window, including words such as attention, listen, repeat, confirm, immediately, be sure, and correct. If at least one such word appears in the window, the context marker word detection result is true; otherwise, it is false. The historical communication tolerance rate is obtained from the closed-loop feedback module, recording the number of critical slot recovery failures reported by the receiver over a past period (i.e., the number of critical information items the receiver could not correctly recover), and the total number of critical slot failures. The tolerance rate equals the number of recovery failures divided by the total number of failures, and its value is between 0 and 1. The command word density, context marker detection results, and historical communication tolerance rate are input into a pre-defined decision mapping library. This results in a lightweight random forest classifier, pre-trained using a large amount of labeled communication data. The inputs are command word density, context marker detection results, and historical communication tolerance rate. The output is a two-dimensional confidence vector, corresponding to the confidence scores of the key slot fidelity-first mode and the continuous naturalness-first mode, respectively. For example, the output could be fidelity confidence and naturalness confidence. The output scene mode confidence vector contains two elements: fidelity confidence (representing the confidence score of the key slot fidelity-first mode) and naturalness confidence (representing the confidence score of the continuous naturalness-first mode). The sum of the two confidence scores is typically 1.
[0036] Read the fidelity confidence score and compare it with a pre-set first threshold. The first threshold is set to a high value, such as 0.8. Simultaneously, read the natural confidence score and compare it with a second threshold, which is set to a low value, such as 0.6. Maintain two counters to record the number of consecutive judgments that result in either fidelity mode or natural mode. If the fidelity confidence score exceeds the first threshold and the number of consecutive fidelity mode judgments reaches N, configure the current scene mode as the critical slot fidelity priority mode and reset the natural mode consecutive counter to zero. If the natural confidence score exceeds the second threshold and the N consecutive judgments are consistent, configure it as the continuous naturalness priority mode and reset the fidelity mode consecutive counter to zero. N is a positive integer, such as 3.
[0037] Because the first threshold is higher than the second threshold, a hysteresis characteristic is created. This means that when the pattern is in the natural priority mode, a high fidelity confidence level is required to switch to the fidelity mode; conversely, when the pattern is in the fidelity mode, a low natural confidence level is required to switch back to the natural mode. This asymmetric threshold design avoids frequent mode switching near the threshold boundaries. The above reasoning process is repeated periodically, and the current pattern is dynamically adjusted based on the latest judgment result.
[0038] By employing frame-level scoring and dynamic segmentation, the system proactively identifies the value differences among different parts of the speech stream, providing a safety boundary for critical instructions and preventing them from being compromised by abrupt segmentation or network fluctuations. Through intelligent scene recognition based on multi-source information, it understands whether the intent of the current communication is for routine, fluent conversation or critical, reliable command, thereby setting the correct global optimization goals for subsequent processing.
[0039] Based on the current intercom mode settings, entry and exit thresholds are set, and the transmission type of each semantic unit is determined. The transmission type includes main transmission unit, summary transmission unit, event transmission unit, or suppression unit.
[0040] Furthermore, this application also includes the following steps: calculating the importance metric of semantic units, wherein the importance metric is determined based on the weighted average of frame scores within the unit and the presence of an instruction protection zone; when the importance metric of a semantic unit is not lower than the entry threshold, the corresponding unit is marked as a main transmission unit and enters the main transmission active state; when the importance metrics of multiple consecutive semantic units are all lower than the exit threshold, the main transmission active state is exited; for semantic units not marked as main transmission units, a decision is made sequentially based on the unit content features, wherein if a non-verbal voice event is included, it is marked as an event transmission unit; otherwise, if at least one key semantic slot is included, it is marked as a summary transmission unit, and the remaining slots are marked as suppression units; the key semantic slots are extracted through a slot filling model based on a conditional random field, wherein the slot filling model shares the underlying feature extraction network with the semantic scoring model.
[0041] Specifically, entry and exit thresholds are set based on the current intercom mode. The entry threshold determines whether a semantic unit is important enough to be marked as a primary transmission unit; the exit threshold determines whether the appearance of multiple consecutive unimportant units means that the active state of primary transmission should end. In the continuous naturalness priority mode, the entry and exit thresholds are set lower, allowing more semantic units to be transmitted to ensure smooth and natural dialogue. In the critical slot fidelity priority mode, the entry and exit thresholds are set higher, ensuring that only truly critical semantic units are marked as primary transmission units, thereby reducing the transmission of unnecessary data and reserving bandwidth for critical information.
[0042] For each semantic unit, the semantic scores of all speech frames within that unit are obtained from the scoring time series. A weighted average of these scores is calculated. If a frame belongs to the instruction protection zone, its weight is higher; frames outside the protection zone have lower weights. The importance metric is calculated as (the sum of the scores of frames within the protection zone multiplied by their weight coefficients plus the sum of the scores of frames outside the protection zone multiplied by their weight coefficients) divided by the total number of frames. If the semantic unit contains at least one semantic triggering feature, the overall importance metric is multiplied by a gain coefficient, such as 1.2, to further highlight units containing key information, ultimately resulting in an importance metric value between zero and one.
[0043] A Boolean state variable, called the main transmission active state, is set, initially set to false. A counter is also set to record the number of semantic units whose importance metric is consecutively below the exit threshold, initially set to 0. For the current semantic unit, if its importance metric is greater than or equal to the entry threshold in the current mode, the semantic unit is marked as a main transmission unit. The main transmission active state is set to true, and the counter for consecutive values below the exit threshold is reset to zero. Furthermore, if the unit was previously inactive, it immediately enters an active state. If the importance metric is less than the entry threshold, it is not immediately marked as a main transmission unit; instead, the current main transmission active state is checked. If the active state is true, it is further determined whether the unit's importance metric is lower than the exit threshold. If it is lower, the counter for consecutive values below the exit threshold is incremented by 1. When the counter reaches a preset number of consecutive values, such as three consecutive semantic units below the exit threshold, the main transmission active state is set to false, and the unit exits the active state. If the importance metric is not lower than the exit threshold (i.e., between the exit and entry thresholds), the current active state remains unchanged, and the counter is neither reset nor incremented. If the main transmission active state is false and the importance metric is lower than the entry threshold, subsequent other type decisions are made directly, and the unit is not marked as a main transmission unit. This ensures that when a key voice segment begins, it quickly enters and remains in an active state until several unimportant units appear in succession, thus avoiding frequent entry and exit from the active state due to brief low-score intervals between main transmission units.
[0044] For semantic units not marked as primary transmission units, it is determined whether the semantic unit contains non-verbal voice events. A binary classifier is trained based on acoustic features to distinguish between normal speech, coughs, laughter, sighs, etc. If the detection result indicates the presence of at least one non-verbal voice event, such as a cough or laughter, the semantic unit is marked as an event transmission unit. For event transmission units, the original speech waveform is not transmitted; instead, a parameterized representation of the event is extracted, including event type, event duration, average intensity, pitch variation features, etc., and packaged into a small data packet and sent to the receiving end. The receiving end uses this data packet to synthesize the corresponding voice event or display emoticons. If the semantic unit does not contain any non-verbal voice events, it is further determined whether the unit contains at least one key semantic slot. The extraction of key semantic slots is accomplished through a slot-filling model based on conditional random fields. The slot-filling model shares the underlying feature extraction network with the semantic scoring model. The underlying network outputs a high-dimensional feature vector for each frame, which is simultaneously fed into the scoring head and the slot-filling head. The slot-filling model uses a large amount of labeled speech data during training. If the slot-filling model detects at least one key semantic slot in a semantic unit—that is, one or more frames are marked as command word slots, number slots, location slots, or object slots—then the unit is marked as a summary transmission unit. For a summary transmission unit, the complete speech is not transmitted; instead, the extracted key slots and their values are packaged in a structured format. For example, for the speech license plate number A8639, the summary packet would contain: slot type is license plate number, and slot value is A8639. The size of the summary packet is typically tens to over a hundred bytes, much smaller than the hundreds to thousands of bytes of the original encoded speech.
[0045] If a semantic unit is neither a main transmission unit nor contains non-verbal voice events, nor any key semantic slots, it is marked as a suppressed unit. For suppressed units, no actual content is transmitted; only a very short time placeholder, such as an empty data packet a few bytes long, is generated, containing the start and end timestamps of the unit. This allows the receiver to know that there was speech during this period but it was actively suppressed, so that it can fill the timeline with comfortable noise or silence during playback, avoiding timeline misalignment caused by no packets at all.
[0046] The semantic scoring model and the slot-filling model share the first few layers of a deep neural network, typically consisting of two convolutional layers and a gated recurrent unit layer. The input is the frequency domain features of the speech frame, and the output is a 256-dimensional feature vector for each frame. The transmitting end performs a short-time Fourier transform on every 20ms speech frame, extracting 80-dimensional log-Mel spectrum features as the raw input for each frame. The first convolutional layer uses 32 3×3 kernels with a stride of 1, outputting 32 feature maps. The second convolutional layer uses 64 3×3 kernels with a stride of 1, outputting 64 feature maps. Then, a global average pooling layer compresses the temporal dimension of each feature map before feeding it into a gated recurrent unit layer with 128 hidden units. This layer models the time series and outputs a 128-dimensional feature vector for each frame. This 128-dimensional vector is the shared low-level feature, containing the acoustic information and some contextual semantic information of the frame. The feature vectors are simultaneously input into two branches: the scoring branch is a network containing two fully connected layers, which ultimately outputs a scalar score; the slot filling branch is a conditional random field layer that outputs the probability distribution of each slot label for each frame. The shared network design reduces the number of model parameters by approximately 40% and computation latency by 30%, making it suitable for real-time operation on mobile devices.
[0047] When determining whether a semantic unit is a summary transmission unit, the slot-filling model is invoked. The 80-dimensional Mel-spectral features of all speech frames within the semantic unit are sequentially input into a shared underlying network, resulting in 128-dimensional shared features for each frame. These shared features are simultaneously input into the scoring branch to obtain a score for each frame, and also into the slot-filling branch. In the slot-filling branch, a linear transformation layer maps the 128-dimensional features to 20-dimensional emission scores. Then, the CRF layer uses the Viterbi algorithm based on the pre-learned label transition probability matrix to find a label sequence that maximizes the total score. All non-O label segments are extracted from the label sequence. For example, if a label sequence contains B-CMD followed by several I-CMDs, a command word slot is extracted, with its value being the text in the corresponding frame. Due to real-time requirements, a small end-to-end speech recognition model is pre-trained, recognizing only a limited vocabulary such as command words, numbers, and directional words. This model works in conjunction with the CRF model to convert frame sequences into slot values. All extracted slots are used as the summary information for the semantic unit.
[0048] Based on the transmission type and the current intercom mode, configure the transmission channel parameters, which include heterogeneous coding parameters, fidelity enhancement parameters, hierarchical scheduling parameters, redundancy allocation parameters, and packet loss recovery parameters.
[0049] Furthermore, this application also includes the following steps: matching corresponding heterogeneous coding parameters according to the transmission type, wherein the main transmission unit adopts neural low bit rate speech coding, the summary transmission unit extracts key slots for packaging, the event transmission unit performs parameterized coding, the suppression unit only generates time placeholders, and the coding bit rate and frame length are adaptively adjusted according to the scene mode; configuring the fidelity enhancement parameters of the corresponding mode according to the current scene intercom mode, wherein, in the key slot fidelity priority mode, double redundant transmission and fast acknowledgment retransmission are enabled for the main transmission unit within the command protection zone; and in the continuous naturalness priority mode, a retransmission disabling mechanism is enabled; Based on the transmission characteristics of the transmission type, hierarchical scheduling parameters and redundancy allocation parameters are configured. Specifically, the transmission queue is dynamically scheduled according to semantic priority and historical under-transmission debt value, and high-priority packets can preempt low-priority packets. Forward error correction redundancy is allocated according to semantic loss cost and predictability recoverability, with high redundancy allocated to content with high loss cost and difficult recovery. The corresponding packet loss recovery parameters are configured according to the current intercom mode. In the continuous naturalness priority mode, feature extrapolation, semantic template reconstruction, or duration-maintaining path are used for packet loss prediction and recovery. In the critical slot fidelity priority mode, prediction is disabled, and timed waiting and silent position holding are used.
[0050] Furthermore, this application also includes the following steps: when the current intercom mode is the critical slot fidelity priority mode, the transmission type is the main transmission unit, and the semantic unit contains an instruction protection zone, the double redundancy transmission and fast acknowledgment retransmission mechanism is activated. The transmitting end generates two independently encoded redundant packets for the same voice content, and transmits them through two physical links or the same link in a time-division manner. The fast acknowledgment timeout threshold is set to 0.5 times the current round-trip delay and the minimum value is not lower than the minimum threshold. When no acknowledgment signal is received within the timeout threshold, the data packet is immediately retransmitted, wherein the retransmitted data packet has the highest transmission priority.
[0051] Specifically, after determining the transmission type and current intercom mode for each semantic unit, specific transmission parameters are configured for each unit, including encoding parameters, fidelity enhancement parameters, scheduling and redundancy parameters, and packet loss recovery parameters. The main transmission unit uses a neural low-bit-rate speech encoder, which is an end-to-end deep neural network. The input is the original speech waveform, and the output is a code stream with an extremely low bit rate. In the continuous naturalness priority mode, a lower bit rate is selected to further reduce latency; in the critical slot fidelity priority mode, a higher bit rate is selected to ensure fidelity. The frame length is also adjusted accordingly: 20ms for natural mode and 10ms for fidelity mode. The summary transmission unit does not encode the speech waveform but packages the key semantic slots (slot type and value) extracted from the slot filling model in the previous step according to a fixed format. The event transmission unit performs parameterized encoding of non-verbal voice events, pre-defining an event type table, such as cough=1, laughter=2, sigh=3, throat clearing=4, whistling=5. Parameters such as the duration, average sound pressure level, and fundamental frequency change trajectory of each event are extracted and packaged. The suppression unit generates only one time placeholder packet containing the start and end timestamps of the unit.
[0052] Fidelity enhancement parameters are only effective in the critical slot fidelity priority mode and are applied only to the main transmission unit within the command protection zone. For the same semantic unit of speech content, the transmitter uses two independent neural encoder instances to generate two different bitstream packets, which are transmitted through two different physical channels to combat packet loss. The receiver can recover the speech upon receiving either packet. The transmitter maintains the state of each data packet. For packets with double redundancy enabled, the transmitter also sets a fast acknowledgment timer with a timeout threshold of 0.5 times the current round-trip time (RTT), and a minimum value of 20ms. If no acknowledgment for the packet is received from the receiver within the timeout period, the packet is immediately retransmitted, selecting either the unacknowledged redundant packet or regenerating it. The retransmitted packet has the highest transmission priority and is inserted at the head of the transmission queue.
[0053] In continuous naturalness-first mode, the retransmission mechanism is completely disabled: the sender does not request retransmission, and the receiver does not send negative acknowledgments; all packet loss is directly handled by the packet loss recovery module. In continuous naturalness-first mode with the retransmission mechanism disabled, the receiver does not send a retransmission request after detecting lost voice data packets, and the NACK or ACK times out. This is because retransmission requires waiting for a round-trip time (RTT), such as 50-200ms, which would affect the real-time performance of the conversation. The receiver utilizes the features of previously received historical voice data to generate lost voice segments in real time using a lightweight neural network, directly filling in the gaps. Although the prediction may not be as accurate as the original audio, it maintains the fluency and low latency of conversations in everyday chat.
[0054] A comprehensive scheduling score is calculated for each data packet to be sent. The score = semantic priority cardinality + under-transmission debt value × debt weight. The semantic priority cardinality is defined as follows: main transmission unit = 100, digest transmission unit = 60, event transmission unit = 40, and suppression unit = 10. The initial value of the under-transmission debt value is 0. Each time a semantic unit is skipped by the scheduler due to queue congestion (i.e., it fails to be sent in time), the debt value increases. The increase is equal to the number of key semantic slots in that unit × the scene importance coefficient. The number of key semantic slots is output by the slot filling model, and the scene importance coefficient is 1.0 in fidelity mode and 0.5 in natural mode. When the network is idle and the transmission queue length is less than 3 packets, the scheduler selects the unit with the highest debt value for compensation transmission according to the debt value from high to low. Only the digest information of that unit is sent. If it was originally a main transmission unit, it is downgraded to digest transmission, or the key features of boundary frames, start frames, and end frames are sent instead of being completely retransmitted. After transmission, the debt value of the unit is multiplied by 0.7 for decay until it is less than 0.01, at which point it is cleared to zero. For each data packet, the semantic loss cost and predicted recoverability are calculated. The redundancy allocation parameter is the proportion of forward error correction (FEC) redundancy allocated to each data packet. The allocation is based on the semantic loss cost and the predicted recoverability. Content with a high loss cost is allocated high redundancy; content with a high predicted recoverability is allocated low redundancy.
[0055] In the continuous naturalness-first mode, a predictive recovery algorithm is enabled at the receiver. When packet loss is detected, features from previously received speech packets are used to generate alternative speech through feature extrapolation, semantic template reconstruction, or duration-preserving path reconstruction via a lightweight neural network to maintain fluency. Feature extrapolation uses the speech features of the preceding frames to extrapolate backwards using a linear or non-linear prediction model, generating approximate features for the lost frames. Semantic template reconstruction selects the best-matching speech segment from a pre-trained semantic template library to fill in the missing segments based on the received semantic slots and context. Duration-preserving path reconstruction does not attempt to recover specific content; it only ensures the continuity of the playback timeline, filling the lost segments with comfortable noise or silence to avoid time misalignment.
[0056] In critical slot fidelity priority mode, all predictive recovery is completely disabled. Upon detecting packet loss, the receiver does not attempt to generate replacement content but instead initiates a timed wait. A wait timer is set for 1.5 times the current RTT. If redundant or retransmitted packets are received within this time, normal decoding and playback occur. If no packets are received within the timeout period, a silent placeholder is inserted at that position, and a report of critical slot recovery failure is sent to the upper layer to update the historical communication tolerance rate.
[0057] After determining the transmission type and recognizing the scene pattern, for each semantic unit to be transmitted, the following conditions are checked simultaneously: the current scene intercom mode is the critical slot fidelity priority mode, the transmission type is the main transmission unit, and the semantic unit contains a command protection zone. If all three conditions are met, the double redundancy transmission and fast acknowledgment retransmission mechanism is activated; otherwise, it is processed according to the normal transmission method. The transmitting end generates two independently encoded redundant packets for the same voice content, inputs the original voice waveform into the neural low bit rate voice encoder, and generates bitstream packet P1 using random seed A in the first step; and bitstream packet P2 using random seed B in the second step. The two seeds are different, resulting in different bitstreams, but either one can be independently decoded into understandable voice. The transmitting end adds header information to each packet, including semantic unit ID, frame sequence number, whether it is a redundant packet identifier, and encoding seed identifier. The transmission method is selected based on available network resources. If the device has both cellular network and wireless LAN, it transmits in parallel through two physical links: P1 uses the cellular network, and P2 uses Wi-Fi. If only one link is available, transmission is time-division multiplexed on the same link: P1 is sent first, followed by P2 after a 5ms interval. Upon receiving the first redundant packet, the receiver immediately decodes and plays it, discarding the second redundant packet. The sender maintains an acknowledgment timer for each redundant packet, measuring the current round-trip time (RTT) in real time and calculating the fast acknowledgment timeout threshold. After sending the first redundant packet, the sender starts the acknowledgment timer. If the receiver successfully receives any redundant packet, it immediately sends back an acknowledgment signal containing the semantic unit's ID. If the sender receives the acknowledgment signal within the timeout threshold, the timer is canceled, the semantic unit transmission is complete, and no further retransmission is needed. If no acknowledgment signal is received within the timeout threshold, all redundant packets for that semantic unit are considered lost or corrupted. The sender immediately triggers a retransmission, generating a new retransmission packet with the highest priority, inserting it at the head of the transmission queue, preempting any low-priority packets currently being sent or waiting to be sent, restarting the acknowledgment timer, and waiting for an acknowledgment signal again. If another timeout occurs, the retransmission is repeated until an acknowledgment signal is received or the number of retransmissions reaches a preset limit, such as 3 times.
[0058] Through heterogeneous coding and semantic debt mechanisms, the bandwidth allocation model has been completely transformed. High-value information receives high-fidelity coding, low-value information consumes almost no bandwidth, and medium-value information is transmitted with an extremely high signal-to-noise ratio. The debt repayment mechanism fills gaps in idle bandwidth, achieving peak shaving and valley filling of bandwidth.
[0059] Based on the transmission channel parameters, a confidence level identifier is added to the transmitted information, and the receiving end performs adaptive playback control based on the confidence level identifier.
[0060] Furthermore, this application also includes the following steps: based on the configuration results of the transmission channel parameters, calculate the content confidence and timing confidence for each data packet, and perform confidence identification, wherein the content confidence is determined based on the transmission type and the corresponding heterogeneous coding parameters, and the timing confidence is determined according to the priority and link state parameters in the hierarchical scheduling parameters; the receiving end adjusts the jitter buffer depth according to the timing confidence and adjusts the prediction replacement strategy according to the content confidence, wherein when the timing confidence is high, the buffer is shallower to reduce playback latency, and when the timing confidence is low, the buffer is deepened to tolerate network jitter; when the content confidence is high, the prediction result is played directly, and when the content confidence is low, the prediction is suppressed and a mute is returned to placeholder; when the real data packet arrives after the prediction recovery result has been played, the subsequent unplayed decoded output is smoothed and corrected, and if the delayed data packet corresponds to the command protection zone, the semantic boundary of the output command is prohibited from being changed.
[0061] Specifically, after the transmitting end completes the transmission channel parameter configuration, for each data packet to be sent, the content confidence and timing confidence are calculated. The base confidence is determined based on the transmission type of the semantic unit to which the data packet belongs; for example, the base confidence for the main transmission unit is 0.95, for the digest transmission unit it is 0.7, for the event transmission unit it is 0.5, and for the suppression unit it is 0.1. If the data packet is a retransmission packet (i.e., a packet retransmitted due to timeout without receiving an acknowledgment signal), the base value is multiplied by an attenuation factor of 0.9. If the data packet is generated through FEC recovery, the receiving end recovers the lost packet using redundant packets, which the transmitting end is unaware of. This step is not applicable at the transmitting end, but the transmitting end knows whether FEC redundancy is enabled, or if the packet is a digest packet and contains fewer slots than expected, allowing for further adjustments. However, the transmitting end typically only sets the base confidence based on the transmission type and the number of retransmissions. The final content confidence is equal to the base value, truncated to between 0 and 1, and quantized as an integer stored between 0 and 255. Content confidence is a value between 0 and 1, representing the reliability of the speech content carried by the current data packet. The higher the content confidence, the more reliable the content of the packet; the lower the content confidence, the more likely the packet has undergone predictive recovery or summary reconstruction, posing a risk of distortion.
[0062] Obtain the semantic priority cardinality of the data packet, such as main transport = 100, digest = 60, event = 40, suppression = 10. Obtain the current link state parameters, including the variance of the RTT samples and the current packet loss rate, and calculate the timing confidence score. The final timing confidence score is truncated to between 0.2 and 0.98 to avoid extreme values, and quantized as an integer between 0 and 255. Fill the calculated content confidence score and timing confidence score into the RTP extension header or a custom protocol header, and send it along with the encoded voice data. The timing confidence score is a value between 0 and 1, representing the predictability of the data packet arrival time. The higher the timing confidence score, the more punctual the packet arrival time is relative to the expected playback time, with less jitter; the lower the timing confidence score, the greater the network jitter, and the packet may arrive early or late. Timing confidence is determined based on the priority in the hierarchical scheduling parameters and the link state parameters: high-priority packets are usually sent first by the scheduler and have a higher timing confidence, such as 0.90; low-priority packets may be queued and delayed, and have a lower timing confidence, such as 0.30. Link state parameters include the variance of the current RTT, packet loss rate, etc. The larger the variance, the lower the timing confidence.
[0063] Two fields are added to the header of each data packet: content confidence and timing confidence. The receiver maintains a dynamic jitter buffer for each voice stream. The initial buffer depth is set to 60ms. Whenever a data packet is received, the receiver parses its timing confidence. If the timing confidence is high, it indicates low network jitter and timely packet arrival; the target buffer depth is reduced to 0.95 times the current depth, thus reducing playback latency and making the dialogue more real-time. If the timing confidence is low, it indicates high network jitter and packets may arrive early or late; the target buffer depth is increased to 1.1 times the current depth to allow more waiting time for late packets and reduce playback interruptions caused by buffer underload. If the timing confidence is between 0.50 and 0.85, the current buffer depth remains unchanged. The receiver adjusts the buffer depth every 100ms or every 10 received packets, using smoothing filtering to avoid drastic changes.
[0064] When the receiver detects a lost or delayed data packet, it needs to decide whether to use predictive recovery. Whether predictive recovery is enabled depends on the content confidence of the semantic unit to which the lost packet belongs, typically obtained from other packets of the same semantic unit previously received. Specifically, if the content confidence is high, it indicates the lost content is critical. In this case, the receiver disables predictive recovery and uses a silent placeholder: comfortable noise or complete silence is inserted during the playback time of the lost segment, and the critical slot loss is reported to the upper layer to update the historical communication tolerance rate. A waiting timer is also started; if subsequent retransmitted or redundant packets arrive, smoothing correction is performed. If the content confidence is low, it indicates the lost content is less important. The receiver enables predictive recovery, selecting based on the current scenario mode. In the continuous naturalness priority mode, the receiver calls a lightweight predictive model to generate an alternative speech segment and plays it directly. Because the content confidence is low, the user has a higher tolerance for distortion, and predictive recovery can maintain smooth dialogue. In the critical slot fidelity priority mode, even if the content confidence is low, prediction is disabled, and a silent placeholder is used uniformly.
[0065] After the receiver plays a replacement audio segment due to prediction recovery, the real data packet arrives. The receiver needs to decide whether to replace the already played prediction data with the real data. If the audio segment corresponding to the real data packet has not yet been played (i.e., it is still waiting in the jitter buffer), the prediction data is discarded, the real packet is decoded, and played in normal order. If the audio segment corresponding to the real data packet has already been played (i.e., the prediction result has been output to the user), the receiver needs to perform smoothing correction, transitioning the decoded output that has not yet been played, rather than modifying the already played history. For example, suppose the time range of the lost segment is [T1, T2], and the prediction segment is played out from T1 to T2. The real packet arrives at time T3 (T3>T2). The receiver does not modify the already played content from T1 to T2. For the audio that has not yet been played after T2, the receiver crossfades the subsequent frames obtained after decoding the real packet with the signal at the end of the previous prediction segment, takes the last 5ms of the prediction segment and the first 5ms of the real packet, and weights them together to make the waveform transition continuously and avoid abrupt changes.
[0066] If the actual data packet corresponds to speech within the command protection zone, altering the semantic boundaries of the output command is prohibited. In other words, even if the actual packet arrives, it cannot be used to correct already played predicted segments, as the user has already heard the predicted command words, and further correction would confuse the user. In this case, only the error is recorded, but no replacement is performed during smoothing correction. If the actual packet arrives after T2 and subsequent content has not yet been played, the actual packet can be used normally for the portion after T2, but the already played protected area content remains unchanged. By dynamically adjusting the jitter buffer depth using timing confidence, the buffer is boldly reduced to lower end-to-end latency when channel conditions are good and packet priority is high; when the channel is unstable, the buffer depth is automatically increased to absorb jitter and avoid intermittent playback, ensuring that the playback latency always approaches the optimal level under the current network conditions.
[0067] Furthermore, this application also includes the following steps: the receiving end periodically obtains one or more feedback indicators among end-to-end delay, packet loss rate, number of critical slot recovery failures, command deadline default rate, prediction recovery confidence, and retransmission success rate, and transmits them back to the sending end through the reverse channel; the sending end switches between continuous naturalness priority target and critical slot fidelity priority target according to the current intercom mode, and dynamically adjusts one or more parameters among semantic unit segmentation threshold, command protection zone backtracking and extension length, transmission type decision boundary, encoded frame length and code rate, forward error correction redundancy allocation, priority mapping rules, and retransmission timeout threshold in fidelity mode based on the feedback indicators.
[0068] Specifically, the receiving end periodically and continuously calculates end-to-end latency, packet loss rate, number of critical slot recovery failures, command deadline default rate, predicted recovery confidence, and retransmission success rate. These indicators are packaged into a feedback message and sent to the sending end every 2 seconds via the reverse channel. For each received data packet, the sending timestamp and receiving time are recorded, and the difference is calculated to obtain the end-to-end delay. Within each period, the receiver maintains the expected sequence number range of received packets, compares it with the actual received sequence numbers, and calculates the proportion of lost packets. Whenever the receiver determines that a key semantic slot has not been correctly recovered, a counter is incremented by 1, and the number is reported. For each identified command word, the end time of its voice acquisition and the time it takes for the receiver to fully play the command are recorded. If the playback time minus the acquisition time exceeds a preset deadline, it is considered a breach of contract. The breach rate is obtained by dividing the number of breached commands by the total number of commands within the period. Each time the receiver performs predictive recovery, a confidence score is output. The average of all predictive recovery confidence scores is used as the feedback value. If the sender has enabled a retransmission mechanism, it records whether each retransmitted packet is successfully received and used. The number of successfully received retransmitted packets within the period is divided by the total number of retransmissions triggered by the sender. The receiver sends the feedback report back to the sender through an available reverse channel.
[0069] Each time the transmitting end receives a feedback message, it parses out various metrics and dynamically adjusts multiple parameters based on the current intercom mode and the values of these metrics. Based on the latest metrics, it determines whether the current mode is still optimal. For example, in continuous naturalness priority mode, if the command deadline default rate suddenly spikes, it indicates that critical instructions are delayed, and it assesses whether to switch to critical slot fidelity priority mode. Conversely, in fidelity mode, if the end-to-end latency is too high and the number of critical slot recovery failures is extremely low, it assesses whether to relax safeguards and switch back to smooth mode to reduce latency. Mode switching follows hysteresis logic to avoid oscillations. Based on feedback metrics, it dynamically adjusts one or more of the following parameters: semantic unit segmentation threshold, instruction protection zone backtracking and extension length, transmission type decision boundary, encoded frame length and bitrate, forward error correction redundancy allocation, priority mapping rules, and retransmission timeout threshold in fidelity mode. The adjusted parameters take effect immediately, affecting subsequent voice processing and transmission. The next feedback cycle will evaluate the effect of these adjustments, thus forming a continuous closed-loop optimization.
[0070] In summary, the walkie-talkie communication latency optimization method for mobile users provided in this application has the following technical effects: by continuously semantically scoring the voice stream and dynamically dividing it into semantic units, setting thresholds to determine the transmission type according to the scene mode, configuring transmission channel parameters differently based on the transmission type and scene mode, and adding confidence labels to the transmitted information to drive the receiver to perform adaptive playback control, the method ensures low-latency and high-reliability delivery of key semantic information while intelligently maintaining the overall fluency of the dialogue, thereby improving the efficiency of walkie-talkie communication.
[0071] Example 2: Based on the same inventive concept as the walkie-talkie communication delay optimization method for mobile users in Example 1, this application also provides a walkie-talkie communication delay optimization system for mobile users. Please refer to the appendix. Figure 2The walkie-talkie communication delay optimization system for mobile users includes: a noise reduction processing module 11, used to feed the input speech sequentially into a noise reduction unit and an echo cancellation unit to filter out background noise and echo, obtaining clean input speech with a high signal-to-noise ratio; a semantic unit segmentation module 12, used to continuously score the clean input speech, dynamically segment the speech stream into semantic units based on continuous score changes and semantic triggering features, and identify and configure the current scene intercom mode; a transmission type determination module 13, used to set entry and exit thresholds based on the current scene intercom mode, and determine the transmission type for each semantic unit, wherein the transmission type includes a main transmission unit, a summary transmission unit, an event transmission unit, or a suppression unit; a transmission channel parameter configuration module 14, used to configure transmission channel parameters according to the transmission type and the current scene intercom mode, wherein the transmission channel parameters include heterogeneous coding parameters, fidelity enhancement parameters, hierarchical scheduling parameters, redundancy allocation parameters, and packet loss recovery parameters; and a confidence identification module 15, used to attach a confidence identification to the transmitted information based on the transmission channel parameters, and the receiving end performs adaptive playback control based on the confidence identification.
[0072] Furthermore, the semantic unit segmentation module 12 in the walkie-talkie communication delay optimization system for mobile users is also used for: the transmitting end performing continuous frame-level semantic importance scoring on clean input speech according to a fixed time window, generating a scoring time series; dynamically segmenting the speech stream based on the time evolution trend of the scoring values in the scoring time series and predefined semantic triggering features, constructing semantic units, wherein when a semantic triggering feature is detected, a preset number of frames are traced back and a preset number of frames are extended backward with the trigger frame as the center to construct a command protection zone, and speech frames within the protection zone have a higher retention priority than frames outside the protection zone in subsequent decisions; and automatically inferring the current scene intercom mode based on the density of explicit user input or command words, context marker words, and historical communication tolerance rate, wherein the scene intercom mode includes at least a continuous naturalness priority mode and a key slot fidelity priority mode.
[0073] Furthermore, the semantic unit partitioning module 12 in the walkie-talkie communication delay optimization system for mobile users is also used for: the semantic triggering feature includes at least one of command words, numeric words, directional words, object words, or call sign words; when the score values of multiple consecutive frames in the scoring time series show a monotonically increasing trend and exceed the starting threshold, the semantic unit starting boundary is triggered; when the score value sequence shows a monotonically decreasing trend and remains below the ending threshold, the ending boundary is triggered; if at least one semantic triggering feature exists between the starting boundary and the ending boundary, then, taking the frame where the triggering feature is located as the center, a first preset number of frames are traced back and a second preset number of frames are extended backward to form a command protection zone, and the units within the protection zone are... All frames are weighted and fused with the original scoring sequence to ensure that the effective score of frames within the protected area is not lower than a predetermined percentage of the maximum score of frames outside the protected area. Specifically, when a local minimum occurs in the scoring sequence and the scores of the preceding and following frames are significantly higher than the minimum, the frame containing the local minimum is considered a natural segmentation candidate point within the semantic unit. Based on the temporal distance between adjacent candidate points and a preset minimum unit duration constraint, it is determined whether to split a single semantic unit into multiple sub-units. If the scoring sequence remains flat within a preset time window and has no semantic triggering features, consecutive frames within the preset time window are merged into a non-instruction semantic unit, whose transmission type is preset as a candidate for either a suppression unit or a digest transmission unit.
[0074] Furthermore, the semantic unit partitioning module 12 in the walkie-talkie communication delay optimization system for mobile users is also used to: obtain the explicit mode selection provided by the user through a physical interface or voice command, and determine the current scene walkie-talkie mode based on the user's explicit input; when there is no explicit mode selection, input the command word density, context identifier word detection results, and historical communication tolerance rate into a preset decision mapping library, and output a scene mode confidence vector; when the confidence of the key slot fidelity priority mode in the scene mode confidence vector exceeds a first threshold and is consistent N times consecutively, configure the current scene mode as the key slot fidelity priority mode; when the confidence of the continuous naturalness priority mode exceeds a second threshold and is consistent N times consecutively, configure it as the continuous naturalness priority mode, wherein the first threshold is higher than the second threshold to construct hysteresis characteristics.
[0075] Furthermore, the transmission type determination module 13 in the walkie-talkie communication delay optimization system for mobile users is also used for: calculating the importance metric of semantic units, wherein the importance metric is determined based on the weighted average of frame scores within the unit and the presence of an instruction protection zone; when the importance metric of a semantic unit is not lower than the entry threshold, the corresponding unit is marked as a main transmission unit and enters the main transmission active state; when the importance metrics of multiple consecutive semantic units are all lower than the exit threshold, the main transmission active state is exited; for semantic units not marked as main transmission units, the unit content features are judged sequentially, wherein if non-verbal voice events are included, the unit is marked as an event transmission unit; otherwise, if at least one key semantic slot is included, the unit is marked as a summary transmission unit, and the remaining slots are marked as suppression units; the key semantic slots are extracted through a slot filling model based on a conditional random field, wherein the slot filling model shares the underlying feature extraction network with the semantic scoring model.
[0076] Furthermore, the transmission channel parameter configuration module 14 in the walkie-talkie communication delay optimization system for mobile users is also used for: matching corresponding heterogeneous coding parameters according to the transmission type, wherein the main transmission unit adopts neural low-bit-rate speech coding, the summary transmission unit extracts key slots for packaging, the event transmission unit performs parameterized coding, the suppression unit only generates time placeholders, and the coding bit rate and frame length are adaptively adjusted according to the scene mode; configuring the fidelity enhancement parameters for the corresponding mode according to the current scene intercom mode, wherein, in the key slot fidelity priority mode, double redundant transmission and fast acknowledgment retransmission are enabled for the main transmission unit within the command protection zone; in the continuous natural In the degree-priority mode, the retransmission mechanism is disabled; hierarchical scheduling parameters and redundancy allocation parameters are configured according to the transmission characteristics of the transmission type. Among them, the sending queue is dynamically scheduled according to semantic priority and historical backlog value, and high-priority packets can preempt low-priority packets; forward error correction redundancy is allocated according to semantic loss cost and predictability recoverability, and high redundancy is allocated to content with high loss cost and difficult recovery; corresponding packet loss recovery parameters are configured according to the current intercom mode. In the continuous natural degree-priority mode, feature extrapolation, semantic template reconstruction or duration-maintaining path is used for packet loss prediction and recovery; in the critical slot fidelity-priority mode, prediction is disabled, and timed waiting and silent position are used.
[0077] Furthermore, the transmission channel parameter configuration module 14 in the walkie-talkie communication delay optimization system for mobile users is also used to: activate a double redundancy transmission and fast acknowledgment retransmission mechanism when the current scenario intercom mode is the key slot fidelity priority mode, the transmission type is the main transmission unit, and the semantic unit contains an instruction protection zone. The transmitting end generates two independently encoded redundant packets for the same voice content, transmits them through two physical links or the same link in a time-division manner, and sets the fast acknowledgment timeout threshold to 0.5 times the current round-trip delay and the minimum value is not lower than the minimum threshold. When no acknowledgment signal is received within the timeout threshold, the data packet is immediately retransmitted, wherein the retransmitted data packet has the highest transmission priority.
[0078] Furthermore, the confidence identification module 15 in the walkie-talkie communication delay optimization system for mobile users is also used for: calculating content confidence and timing confidence for each data packet based on the configuration results of the transmission channel parameters, and identifying the confidence. The content confidence is determined based on the transmission type and corresponding heterogeneous coding parameters, and the timing confidence is determined based on the priority and link status parameters in the hierarchical scheduling parameters. The receiving end adjusts the jitter buffer depth according to the timing confidence and adjusts the prediction replacement strategy according to the content confidence. When the timing confidence is high, the buffer becomes shallower to reduce playback delay, and when the timing confidence is low, the buffer becomes deeper to tolerate network jitter. When the content confidence is high, the prediction result is played directly, and when the content confidence is low, the prediction is suppressed and a mute is returned to placeholder. When the real data packet arrives after the prediction recovery result has been played, the subsequent unplayed decoded output is smoothed and corrected. If the delayed data packet corresponds to the command protection zone, the semantic boundary of the output command is prohibited from being changed.
[0079] Furthermore, the confidence identification module 15 in the walkie-talkie communication delay optimization system for mobile users is also used for: the receiving end periodically obtaining one or more feedback indicators among end-to-end delay, packet loss rate, number of critical slot recovery failures, command deadline default rate, prediction recovery confidence, and retransmission success rate, and transmitting them back to the sending end through the reverse channel; the sending end switches between continuous naturalness priority target and critical slot fidelity priority target according to the current scenario walkie-talkie mode, and dynamically adjusts one or more parameters among semantic unit segmentation threshold, command protection zone backtracking and extension length, transmission type decision boundary, encoded frame length and code rate, forward error correction redundancy allocation, priority mapping rules, and retransmission timeout threshold in fidelity mode based on the feedback indicators.
[0080] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The walkie-talkie communication delay optimization method and specific examples for mobile users in the aforementioned embodiment one are also applicable to the walkie-talkie communication delay optimization system for mobile users in this embodiment. Through the foregoing detailed description of the walkie-talkie communication delay optimization method for mobile users, those skilled in the art can clearly understand the walkie-talkie communication delay optimization system for mobile users in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.
[0081] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0082] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for optimizing a delay of a walkie-talkie communication for mobile users, characterized in that, include: The input speech is sequentially fed into the noise reduction unit and the echo cancellation unit to filter out background noise and echo, thereby obtaining clean input speech with a high signal-to-noise ratio. Clean input speech is continuously scored, and the speech stream is dynamically divided into semantic units based on the changes in continuous scores and semantic triggering features. The current scene intercom mode is also identified and configured. Based on the current intercom mode, set entry and exit thresholds, and determine the transmission type for each semantic unit. The transmission type includes main transmission unit, summary transmission unit, event transmission unit, or suppression unit. Based on the transmission type and the current intercom mode, configure the transmission channel parameters, which include heterogeneous coding parameters, fidelity enhancement parameters, hierarchical scheduling parameters, redundancy allocation parameters, and packet loss recovery parameters. Based on the transmission channel parameters, a confidence level identifier is added to the transmitted information, and the receiving end performs adaptive playback control based on the confidence level identifier.
2. The walkie-talkie communication delay optimization method for mobile users according to claim 1, characterized in that, The system continuously scores the clean input speech, dynamically divides the speech stream into semantic units based on changes in the continuous scores and semantic triggering features, and identifies and configures the current scene's intercom mode, including: The transmitting end performs continuous frame-level semantic importance scoring on clean input speech according to a fixed time window, generating a scoring time series; Based on the time evolution trend of the scoring value in the scoring time series and the predefined semantic triggering features, the speech stream is dynamically segmented to construct semantic units. When a semantic triggering feature is detected, a preset number of frames are traced back and a preset number of frames are extended backward from the trigger frame to construct an instruction protection zone. Speech frames within the protection zone have a higher retention priority than frames outside the protection zone in subsequent decisions. Based on the user's explicit input or command word density, context identifier words, and historical communication tolerance rate, the current scene intercom mode is identified and configured automatically. The scene intercom mode includes at least a continuous naturalness priority mode and a key slot fidelity priority mode.
3. The walkie-talkie communication delay optimization method for mobile users according to claim 2, characterized in that, Based on the time evolution trend of score values in the score time series and predefined semantic triggering features, the speech stream is dynamically segmented to construct semantic units, including: The semantic triggering features include at least one of command words, numeric words, locative words, object words, or call sign words. When the score values of multiple consecutive frames in the scoring time series show a monotonically increasing trend and exceed the starting threshold, the delineation of the starting boundary of the semantic unit is triggered. When the score value sequence shows a monotonically decreasing trend and remains below the termination threshold, the termination boundary is triggered. If there is at least one semantic trigger feature between the start boundary and the end boundary, then take the frame where the trigger feature is located as the center, backtrack for a first preset number of frames and extend backward for a second preset number of frames to form an instruction protection zone, and then perform weighted fusion of all frames in the protection zone with the original scoring sequence so that the effective score of the frame in the protection zone is not less than a predetermined percentage of the maximum score of the frame outside the protection zone. Specifically, when a local minimum value appears in the score value sequence and the scores of the preceding and following frames are significantly higher than the minimum value, the frame where the local minimum value is located is taken as a natural segmentation candidate point within the semantic unit. Based on the time distance between adjacent candidate points and the preset minimum unit duration constraint, it is determined whether to split a single semantic unit into multiple sub-units. If the score value sequence remains flat and has no semantic triggering features within a preset time window, then the consecutive frames within the preset time window are merged into a non-instruction semantic unit, and its transmission type is preset as a candidate for suppression unit or summary transmission unit.
4. The walkie-talkie communication delay optimization method for mobile users according to claim 2, characterized in that, Based on user explicit input or command word density, contextual identifiers, and historical communication tolerance rates, the system automatically identifies and configures the current scene's intercom mode, including: Obtain the explicit mode selection provided by the user through a physical interface or voice command, and determine the intercom mode for the current scene based on the user's explicit input; When there is no explicit mode selection, the command word density, context identifier word detection results and historical communication tolerance rate are input into the preset decision mapping library, and the scene mode confidence vector is output. When the confidence of the key slot fidelity priority mode in the scene mode confidence vector exceeds the first threshold and the judgment is consistent for N consecutive times, the current scene mode is configured as the key slot fidelity priority mode. When the confidence level of the continuous naturalness priority mode exceeds the second threshold and the judgments are consistent for N consecutive times, the continuous naturalness priority mode is configured, wherein the first threshold is higher than the second threshold to construct the hysteresis characteristic.
5. The walkie-talkie communication delay optimization method for mobile users according to claim 1, characterized in that, Based on the current intercom mode settings, entry and exit thresholds are set, and the transmission type is determined for each semantic unit, including: The importance metric of a semantic unit is calculated, which is determined based on the weighted average of frame scores within the unit and the presence of an instruction protection zone. When the importance metric of a semantic unit is not lower than the entry threshold, the corresponding unit is marked as a main transmission unit and enters the main transmission active state. When the importance metric of multiple consecutive semantic units is lower than the exit threshold, the main transmission active state is exited. For semantic units that are not marked as primary transmission units, they are judged sequentially based on the characteristics of the unit content. If they contain non-verbal human voice events, they are marked as event transmission units. Otherwise, if it contains at least one key semantic slot, it is marked as a digest transmission unit, and the rest are marked as suppression units; The key semantic slots are extracted using a slot-filling model based on conditional random fields, and the slot-filling model shares the underlying feature extraction network with the semantic scoring model.
6. The walkie-talkie communication delay optimization method for mobile users according to claim 2, characterized in that, Configure transmission channel parameters according to the transmission type and the current intercom mode, including: According to the transmission type, the corresponding heterogeneous coding parameters are matched. The main transmission unit adopts neural low bit rate speech coding, the summary transmission unit extracts key slots and packages them, the event transmission unit performs parameterized coding, the suppression unit only generates time placeholders, and the coding bit rate and frame length are adaptively adjusted according to the scene mode. Configure the corresponding fidelity enhancement parameters according to the current intercom mode. In the critical slot fidelity priority mode, enable double redundant transmission and fast confirmation retransmission for the main transmission unit in the command protection zone. Enable the disabling retransmission mechanism in continuous naturalness priority mode; Configure hierarchical scheduling parameters and redundancy allocation parameters according to the transmission characteristics of the transmission type. The sending queue is dynamically scheduled according to semantic priority and historical under-transmission debt value. High-priority packets can preempt low-priority packets. Forward error correction redundancy is allocated based on semantic loss cost and predicted recoverability, with high redundancy allocated to content that has high loss cost and is difficult to recover. Configure the corresponding packet loss recovery parameters according to the current intercom mode. In the continuous naturalness priority mode, use feature extrapolation, semantic template reconstruction or duration-preserving path to perform packet loss prediction and recovery. In the critical slot fidelity priority mode, prediction is disabled, and timed waiting and silent preemption are used.
7. The walkie-talkie communication delay optimization method for mobile users according to claim 6, characterized in that, In the critical slot fidelity priority mode, double redundancy transmission and fast acknowledgment retransmission are enabled for the main transmission unit within the command protection zone, including: When the current intercom mode is the critical slot fidelity priority mode, the transmission type is the main transmission unit, and the semantic unit contains an instruction protection zone, the double redundancy transmission and fast acknowledgment retransmission mechanism is activated. The sending end generates two independently encoded redundant packets for the same voice content and transmits them through two physical links or time-division multiplexing of the same link. The fast acknowledgment timeout threshold is set to 0.5 times the current round-trip delay and the minimum value is not lower than the minimum threshold. If no acknowledgment signal is received within the timeout threshold, the data packet is immediately retransmitted, and the retransmitted data packet has the highest transmission priority.
8. The walkie-talkie communication delay optimization method for mobile users according to claim 1, characterized in that, Based on the transmission channel parameters, a confidence indicator is added to the transmitted information. The receiving end performs adaptive playback control based on the confidence indicator, including: Based on the configuration results of the transmission channel parameters, the content confidence score and the timing confidence score are calculated for each data packet, and the confidence score is identified. The content confidence score is determined based on the transmission type and the corresponding heterogeneous coding parameters, and the timing confidence score is determined based on the priority in the hierarchical scheduling parameters and the link state parameters. The receiver adjusts the jitter buffer depth based on timing confidence and the prediction replacement strategy based on content confidence. Specifically, when timing confidence is high, the buffer becomes shallower to reduce playback latency, while when timing confidence is low, the buffer becomes deeper to tolerate network jitter. When the content confidence is high, the prediction result is played directly; when the content confidence is low, the prediction is suppressed and the program is muted. When the actual data packet arrives late after the predicted recovery result has been played, smooth correction is performed on the subsequent decoded output that has not yet been played. If the delayed data packet corresponds to the instruction protection zone, the semantic boundary of the output command is prohibited from being changed.
9. The walkie-talkie communication delay optimization method for mobile users according to claim 1, characterized in that, Also includes: The receiving end periodically obtains one or more feedback indicators, including end-to-end delay, packet loss rate, number of critical slot recovery failures, command deadline default rate, predicted recovery confidence, and retransmission success rate, and transmits them back to the sending end through the reverse channel. The transmitting end switches between continuous naturalness priority target and critical slot fidelity priority target according to the current intercom mode, and dynamically adjusts one or more of the following parameters based on feedback indicators: semantic unit segmentation threshold, instruction protection zone backtracking and extension length, transmission type decision boundary, encoded frame length and bit rate, forward error correction redundancy allocation, priority mapping rules, and retransmission timeout threshold in fidelity mode.
10. A walkie-talkie communication delay optimization system for mobile users, characterized in that, The step of implementing the walkie-talkie communication delay optimization method for mobile users according to any one of claims 1 to 9, wherein the walkie-talkie communication delay optimization system for mobile users comprises: The noise reduction processing module is used to feed the input speech into the noise reduction unit and the echo cancellation unit in sequence to filter out background noise and echo, so as to obtain clean input speech with high signal-to-noise ratio. The semantic unit segmentation module is used to continuously score the clean input speech, dynamically segment the speech stream into semantic units based on the continuous score changes and semantic triggering features, and identify and configure the intercom mode for the current scene. The transmission type determination module is used to set entry and exit thresholds based on the current intercom mode and to determine the transmission type of each semantic unit. The transmission type includes main transmission unit, summary transmission unit, event transmission unit or suppression unit. The transmission channel parameter configuration module is used to configure transmission channel parameters according to the transmission type and the current intercom mode. The transmission channel parameters include heterogeneous coding parameters, fidelity enhancement parameters, hierarchical scheduling parameters, redundancy allocation parameters, and packet loss recovery parameters. The confidence level identification module is used to attach a confidence level identifier to the transmitted information based on the transmission channel parameters, and the receiving end performs adaptive playback control based on the confidence level identifier.