Power dispatching call real-time intelligent voice transcription system

CN122802491APending Publication Date: 2026-09-22国网陕西省电力有限公司 +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611243205.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0002]当前电力调度通信网络主要采用基于程控交换或软交换技术的音频混合桥接模式,将多路输入音频信号在数字层叠加后输出单一混合音频流,满足多方通话实时语音交互需求,音频混合过程中,各通信终端独立声纹特征与背景噪声频谱在物理层发生混叠,后端语音识别系统处理此类混合信号时,难以通过盲源分离算法还原重叠时段独立信源,导致多人争抢发言或高噪背景叠加工况下转写准确率降低

Benefits of technology

1、在实时智能语音转写中,媒体交换层面构建非混音逻辑通道分离机制,配合信令解析逻辑将业务属性编码写入RTP头部扩展字段,实现通信信源物理特征与业务身份绑定,改变传统电话会议桥接信号混合模式,将混合流物理拆解为多路携带确定性身份标签独立逻辑通道,从传输协议底层保留各信源频谱独立性与时序完整性,无需依赖盲源分离算法,消除多方通话信号混叠干扰,解决VoIP架构信令与媒体流路径分离导致身份匹配错位问题,确保下游处理系统接收媒体数据流具备源端可追溯性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802491A_ABST
    Figure CN122802491A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of telephone communication, and discloses a power dispatching call real-time intelligent voice transcription system, which comprises a service state mapping unit, a media stream orthogonal split gateway and a protocol frame level packaging unit.The service state mapping unit is used for analyzing SIP signaling in a call establishment stage and extracting a calling terminal physical address, and is used for mapping the calling terminal physical address into power grid service attribute coding to establish a session level source end identity mapping table; the media stream orthogonal split gateway is used for intercepting original RTP voice data packets, performing non-mixing channel separation to construct independent parallel logical transmission channels; and the protocol frame level packaging unit is used for retrieving the mapping table and writing corresponding service attribute coding into an RTP header extension field, and is used for outputting independent media streams carrying deterministic identity tags.The application constructs a non-mixing split mechanism in a media transmission layer and a signaling data frame level injection mechanism, eliminates aliasing interference of multi-party calls, and realizes the binding of dispatching instructions and dispatching subject identities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a real-time intelligent voice transcription system for power dispatching calls, belonging to the field of telephone communication technology. Background Technology

[0002] The current power dispatch communication network mainly adopts an audio hybrid bridging mode based on program-controlled switching or soft switching technology. It outputs a single hybrid audio stream after superimposing multiple input audio signals at the digital layer to meet the real-time voice interaction needs of multi-party calls. During the audio mixing process, the independent voiceprint features of each communication terminal and the background noise spectrum are superimposed at the physical layer. When the back-end speech recognition system processes such hybrid signals, it is difficult to restore the independent signal sources in the overlapping period through blind source separation algorithms, resulting in a decrease in transcription accuracy under conditions of multiple people competing to speak or high noise background superposition.

[0003] Existing terminal designs often focus on physical signal access while neglecting logical source identification. For example, the utility model patent with authorization announcement number CN201130995Y discloses a multimedia dispatch communication seat. Although it integrates a video capture card, voice card, and routing equipment to achieve digital access and distribution of audio and video, its core technology is still limited to traditional channel forwarding. The system only encodes analog signals into digital streams for transparent transmission and does not construct a digital fingerprint of the issuing entity at the RTP protocol layer. The media stream loses its service attributes as soon as it leaves the seat and enters the switching network. The signaling stream and media stream are transmitted along independent logical paths. The transcription system lacks real-time correlation between the real-time transmission protocol media stream and the session initialization protocol signaling data, resulting in the generated transcribed text being unable to reflect the signaling data. When data is sent to specific substations or dispatch consoles, there is a lack of identity anchoring in high-concurrency scenarios of emergency consultations across the entire network, and there is a risk of incorrect attribution determination in instruction records. The power dispatch network involves heterogeneous access environments such as fiber optics, microwave, and satellite. The physical delay and jitter of data packets sent by different terminals to the switching nodes vary. The existing real-time transmission protocol relies on the local clock of the sending end to generate timestamps. Without high-precision clock synchronization across the entire network, the time sequence restored by the receiving end based on the original timestamps does not match the physical auditory facts, resulting in a reversal of causal logic. The existing switching mechanism uses a uniform encoding strategy for all channels, which cannot distinguish the transmission priority of core instructions and background noise based on business roles, and occupies communication resources in bandwidth-limited private network environments.

[0004] Therefore, the technical problem to be solved by this invention is how to solve the problems of signal source feature loss caused by multi-party call signal aliasing, identity matching mismatch caused by signaling media separation, timing logic disorder caused by heterogeneous network transmission differences, and mismatch between communication resource allocation and service value without modifying existing terminal hardware. Summary of the Invention

[0005] To address the problems mentioned in the background art, the technical solution of the present invention is as follows: A real-time intelligent voice transcription system for power dispatching calls, comprising: The service status mapping unit is used to access the signaling control plane of the dispatching and switching network, parse the session initialization protocol SIP signaling message during the call setup phase and extract the physical address identifier of the calling terminal, map the physical address identifier to the power grid service attribute code according to the preset power grid topology database, and establish a session-level source identity mapping table in the signaling control plane that contains the correspondence between the physical address identifier and the power grid service attribute code. The media stream orthogonal splitting gateway is connected in series in the media transmission plane of the scheduling and switching network and communicates with the service state mapping unit. It is used to obtain the source identity mapping table and intercept the original Real-Time Transport Protocol (RTP) voice data packets sent by multiple communication terminals belonging to the same session. The gateway executes unmixed channel separation logic to build parallel logical transmission channels equal to the number of communication terminals, and uses a zero-copy memory mechanism to copy and split the original RTP voice data packets to the corresponding logical transmission channels, thereby building a physically isolated and time-synchronized independent transmission environment for each voice data packet. The protocol frame-level encapsulation unit is located at the output port of the logical transmission channel. Under the closed transmission constraints of the logical transmission channel, it retrieves the power grid service attribute code corresponding to the communication terminal to which the current channel belongs from the source identity mapping table, and writes the power grid service attribute code into the protocol header extension field of the original RTP voice data packet using the header extension format defined by the RFC5285 standard, forming an independent media stream output carrying a deterministic service identity label.

[0006] Preferably, when the media stream orthogonal splitter gateway intercepts voice streams sent simultaneously by multiple communication terminals, it maintains the independent encoding format and timing characteristics of each voice stream, does not perform digital mixing and overlay processing of multiple audio waveforms, and independently outputs a single audio stream in each logical transmission channel.

[0007] Preferably, the protocol frame-level encapsulation unit encapsulates the power grid service attributes into a fixed-length binary tag using network byte order. This binary tag is sent frame by frame with each RTP voice data packet, ensuring that the downstream receiver can parse the unique owner of the current voice frame under any network latency.

[0008] Preferably, the media stream orthogonal splitting gateway also includes unified time base reshaping logic, which maintains a monotonically increasing virtual reference clock for the current session in memory; when the gateway receives the original RTP voice data packet sent by any communication terminal in the kernel mode of the network interface card, it records the physical arrival time and calculates the relative offset between the physical arrival time and the virtual reference clock; before splitting the original RTP voice data packet to the logical transmission channel, the gateway rewrites the timestamp field of the RTP protocol header using the relative offset, and the rewritten timestamp field represents the logical alignment time of each communication terminal at the gateway.

[0009] Preferably, the media stream orthogonal splitter gateway stores a differentiated encoding strategy table based on service roles. This strategy table defines the pass-through strategy corresponding to core roles and the transcoding strategy corresponding to non-core roles. The gateway matches the encoding strategy for each logical transmission channel based on the role field in the power grid service attribute code extracted by the service status mapping unit. For logical transmission channels that match the pass-through strategy, the gateway keeps the payload encoding format of the original RTP voice data packets unchanged. For logical transmission channels that match the transcoding strategy, the gateway calls the digital signal processing core to transcode the payload of the original RTP voice data packets into a high compression ratio format and modifies the payload type field in the RTP protocol header.

[0010] Preferably, when the media stream orthogonal splitter gateway detects that a communication terminal corresponding to a non-core role has acquired the session speaking right signaling, it triggers the logical transmission channel corresponding to that terminal to switch from the transcoding strategy to the transparent transmission strategy.

[0011] Preferably, the media stream orthogonal splitter gateway also includes an adaptive load suppression module, which monitors the voice signal energy value in each logical transmission channel in real time. When the signal energy value of a certain logical transmission channel is detected to be continuously lower than the preset background noise threshold, the module controls the gateway to stop sending voice payload data packets for that channel and sends silent descriptor frames containing only protocol header extension fields at preset intervals to maintain the heartbeat connection of the service identity tag.

[0012] Preferably, the data structure for power grid business attribute encoding includes a unique substation identification code, a dispatch console seat number, and an equipment bay number. The business status mapping unit obtains business attribute data by parsing the INVITE or INFO message body of the SIP protocol.

[0013] Preferably, the unified time base refactoring logic uses the following formula to calculate the timestamp value used to rewrite the RTP header. : ,in, The physical time when the original RTP voice data packet arrives at the gateway. The reference time for the virtual reference clock. The sampling rate of the current audio encoding. This is the preset jitter buffer bias.

[0014] Preferably, the system also includes a transcription engine interface located downstream of the logical transmission channel. This interface is used to parse the RTP header extension fields of the independent media stream to extract the power grid service attribute code, and to call the corresponding acoustic model to perform speech recognition on the independent media stream based on the code.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. In real-time intelligent speech transcription, a non-mixed logical channel separation mechanism is constructed at the media exchange layer. In conjunction with the signaling parsing logic, the service attribute encoding is written into the extended field of the RTP header, realizing the binding of the physical characteristics of the communication source with the service identity. This changes the traditional telephone conference bridging signal mixing mode, physically decomposes the mixed stream into multiple independent logical channels carrying deterministic identity labels, and preserves the spectrum independence and timing integrity of each source from the transmission protocol layer. It does not rely on blind source separation algorithms, eliminates the interference of multi-party call signal aliasing, solves the problem of identity matching mismatch caused by the separation of signaling and media stream paths in VoIP architecture, and ensures that the downstream processing system has source traceability of the received media data stream.

[0016] 2. Configure unified time base resetting logic at the gateway to maintain the global virtual reference clock for the current session. Rewrite the timestamp field in the RTP header based on the physical time when the data packet arrives at the gateway. Establish a unified time measurement system based on the convergence time of the switching nodes. This will shield the local clock asynchrony and transmission jitter caused by the heterogeneous access environment of different substation terminals. There is no need to modify the hardware of existing terminals or deploy a high-precision synchronization protocol across the entire network. This will enable the alignment of multiple asynchronous media streams on the logical time axis, ensure the objective authenticity of the timing logic of dispatching instructions and on-site feedback information, and avoid the risk of causal logic reversal caused by network latency differences.

[0017] 3. Adopt a business role-based differentiated payload coding mechanism, dynamically configure the media plane audio coding format and payload type field according to the business role attributes parsed in the signaling plane, establish a mapping relationship between communication coding quality and business administrative level, maintain high-fidelity transmission of voice streams for core scheduling roles, and perform high-compression transcoding on voice streams for non-core listening roles, ensuring the clarity of key command information, reducing the bandwidth resource occupation of the power dedicated network by the multi-path parallel split architecture, and improving the carrying capacity and engineering stability of the communication network under narrowband constraints or high-concurrency sudden operating conditions. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the system logic architecture based on orthogonal signaling media splitting according to the present invention; Figure 2 This is a trend chart showing the latency performance of key processes under different concurrent session scales in this invention. Detailed Implementation

[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0020] A real-time intelligent speech-to-text system for power dispatching calls includes a service state mapping unit located in the signaling control plane, a media stream orthogonal splitting gateway located in the media transmission plane, and a protocol frame-level encapsulation unit embedded in the gateway's output port. Logically, this system is connected in series between the power dispatching switching network and the speech-to-text engine. Through cross-layer collaboration between signaling parsing and media splitting, it achieves independent source extraction and deterministic identity binding in multi-party call scenarios. In power dispatching communication networks, the mixing and bridging mode used in traditional switching equipment causes irreversible aliasing of multiple audio signals at the physical layer, making it impossible for the backend recognition system to separate overlapping speech features. The media stream orthogonal splitter gateway is configured to perform non-mixed channel separation logic. When the gateway intercepts raw Real-Time Transport Protocol (RTP) voice data packets sent by multiple communication terminals belonging to the same session, it does not perform digital superposition processing of multiple audio waveforms. Instead, it constructs a parallel logical transmission channel equal to the number of communication terminals based on a zero-copy memory mechanism. The gateway reads the Synchronization Source Identifier (SSRC) of each RTP data packet, copies it completely, and splits it into the corresponding unidirectional logical channel. This creates a physically isolated independent transmission environment for each voice data packet at the transport layer, ensuring that the output media stream retains the spectral characteristics and timing integrity of the original source.

[0021] To address the issue of inaccurate matching of the issuing entity in transcribed text due to the separation of signaling and media stream paths in existing VoIP architectures, the service status mapping unit accesses the signaling control plane of the dispatch switching network. It parses the Session Initiation Protocol (SIP) signaling messages during the call setup phase in real time. This unit extracts the physical address identifier of the calling terminal from the INVITE or INFO message body and retrieves a pre-defined power grid topology database to obtain the unique substation identification code, dispatch console seat number, and equipment bay number corresponding to that physical address. Based on this information, the service status mapping unit generates power grid service attribute codes and establishes a mapping between physical address identifiers and power grid service attribute codes in the signaling control plane. A session-level source identity mapping table; a protocol frame-level encapsulation unit is set at the output port of the logical transmission channel to execute a deterministic identity tag injection procedure when forwarding voice data packets. This unit retrieves the source identity mapping table, matches the power grid service attribute code corresponding to the communication terminal to which the current channel belongs, and encapsulates the code into a fixed-length binary tag using the header extension format defined by the RFC5285 standard. This tag is then written into the protocol header extension field of the original RTP voice data packet and sent packet by packet with each voice frame. The downstream transcription engine parses this extension field, and can parse out the unique owner of the current voice frame under any network latency or out-of-order conditions, thus realizing atomic binding between scheduling instructions and the identity of the issuing entity.

[0022] Considering that the power dispatch network covers heterogeneous access environments such as fiber optic, microwave, and satellite, the physical latency and jitter of data packets arriving at the switching node vary among different terminals. Directly relying on the sender's local clock would lead to a reversal of causal logic. The media stream orthogonal splitter gateway has built-in unified time base reshaping logic, which maintains a monotonically increasing virtual reference clock in memory for the current session. When the gateway's network interface card receives a raw RTP voice data packet sent by any communication terminal in kernel mode, it records the physical time of the data packet's arrival at the gateway. Before routing data packets to the logical transmission channel, the gateway rewrites the timestamp field in the RTP protocol header using a relative offset. The rewritten timestamp value... Calculate according to the following formula: ,in, The reference time for the virtual reference clock. The sampling rate of the current audio encoding. With a preset jitter buffer bias, this mechanism establishes a unified time measurement system based on the convergence time of the exchange nodes.

[0023] To address the bandwidth pressure caused by the multi-path splitting architecture in the narrowband environment of the power grid, the system stores a differentiated encoding strategy table based on service roles. This table defines the pass-through strategy for core roles and the transcoding strategy for non-core roles. The gateway matches the encoding strategy for each logical transmission channel based on the role field in the power grid service attribute code extracted by the service status mapping unit. For channels matching the pass-through strategy, the gateway maintains the original RTP voice data packet payload encoding format unchanged, such as G.711. For channels matching the transcoding strategy, the gateway calls the digital signal processing core to transcode the original RTP voice data packet payload into a high-compression format, such as G.72. The gateway adopts a 9-format and modifies the payload type field in the RTP protocol header. In addition, the gateway includes an adaptive payload suppression module that monitors the voice signal energy value in the logical transmission channel in real time. When the signal energy value of a certain channel is detected to be continuously lower than the preset background noise threshold, the gateway is controlled to stop sending voice payload data packets for that channel and send silent descriptor frames containing only the protocol header extension field at preset intervals to maintain the heartbeat connection of the service identity label. When the gateway detects that a communication terminal corresponding to a non-core role has obtained the session speaking right signaling, it triggers the logical transmission channel corresponding to that terminal to switch from the transcoding strategy to the transparent transmission strategy to ensure the clarity of key instruction information.

[0024] Example 1: In a high-concurrency power dispatching accident handling scenario involving a 500kV hub substation and a remote mountain wind farm, due to the communication links encompassing heterogeneous networks such as fiber optic private networks and satellite relays, dispatching commands initiated by each end and on-site reporting signals experience physical time delay differences and severe jitter when reaching the switching center. Furthermore, the interleaving of multiple signals leads to the physical loss of key voiceprint features in traditional mixing modes. To address the aforementioned issues of source feature aliasing and identity attribution loss, the service state mapping unit accesses the signaling control plane and parses the Session Initiation Protocol (SIP) signaling message, extracting information from the INVITE message body... The physical address of the calling terminal is extracted to generate a power grid service attribute code containing the substation's unique identification code and the dispatch console seat number. The media stream orthogonal split gateway then intercepts the original Real-Time Transport Protocol (RTP) voice data packets belonging to the session, executes unmixed channel separation logic, and uses a zero-copy memory mechanism to construct a parallel logical transmission channel equal to the number of terminals. The protocol frame-level encapsulation unit retrieves the aforementioned source identity mapping table, encapsulates the corresponding service attribute code into a fixed-length binary tag, and writes it into the RTP protocol header extension field, thereby achieving atomic binding between the voice payload and the calling entity's identity at the transport layer.

[0025] In addition to addressing the risk of temporal and causal logic reversal caused by heterogeneous access environments, the unified time base reshaping logic built into the media stream orthogonal splitter gateway maintains a monotonically increasing virtual reference clock in memory for the current session. When the gateway kernel receives any raw RTP voice data packet, it records its physical arrival time. And according to the formula Rewrite the timestamp field in the RTP header, where As the reference time and To improve the sampling rate, this mechanism establishes a globally unified time measurement system based on the convergence time of the switching nodes, eliminating timing errors caused by local clock asynchrony at the source and asymmetry in the transmission path. Under the objective constraint of limited bandwidth in the power grid, the system performs hierarchical processing on voice streams of different value densities according to a preset differentiated coding strategy table. The gateway identifies the core position of the dispatch console based on the role field in the service attribute code and maintains its voice stream in G.711 transparent transmission format. At the same time, it calls the digital signal processing core to transcode the non-core voice stream of the wind farm listening terminal into G.729 high compression ratio format and modify the load type field. Furthermore, when a non-core terminal is detected to be acquiring the right to speak, the channel strategy is switched in real time. This resource allocation mechanism based on semantic value ensures high-fidelity transmission of key instructions while reducing the overall bandwidth usage of the multi-path splitting architecture. Through the above-mentioned communication architecture reconstruction based on orthogonal splitting and deterministic signaling injection, this embodiment reshapes the murky mixed analog stream in traditional dispatch calls into multiple independent digital streams carrying identity tags and unified timing logic without modifying the existing terminal hardware. This ensures that the downstream transcription engine can still restore the semantic content and issuing body of each dispatch instruction under extremely high concurrency and high noise conditions.

[0026] Example 2: This example aims to verify the practical effectiveness of the proposed real-time intelligent speech-to-text system for power dispatching under complex operating conditions through engineering experiments. The experiment focuses on evaluating the system's performance in key performance indicators such as speech-to-text accuracy, identity recognition accuracy, and instruction timing consistency when facing typical dispatching communication challenges such as strong background noise, high packet loss rate transmission, and multiple concurrent callers. This demonstrates the engineering value of core mechanisms such as state-media orthogonal encapsulation and unified time-base reshaping. To reproduce the requirements of the power dispatching environment, a closed-loop test platform consisting of a physical simulation terminal group, a network impairment simulator, and a professional sound field synthesis system is constructed. On the signal input side, eight concurrent signal sources are configured, including four connected via a dedicated fiber optic network. The system includes a 500kV substation shift supervisor terminal, two 220kV substation terminals connected via microwave links, and two wind farm mobile terminals connected via low-orbit satellite links. This system covers typical heterogeneous access scenarios in power communication networks. Using sound field synthesis technology, background noise of transformer operation with a signal-to-noise ratio (SNR) of 15dB is superimposed on each original voice signal, and random frequency bursts of operation alarm tones are injected to simulate the real acoustic environment of a substation. A professional network impairment instrument is connected in series in the transmission link to inject packet loss (packet loss rate set to 5%) and jitter (jitter range set to 50ms-200ms) conforming to the Rayleigh fading model into the microwave and satellite links to reproduce communication link damage under severe weather conditions.

[0027] The experimental design introduced two parallel control groups for performance comparison. The control group adopted the traditional mixing and bridging mode based on softswitch MCU. All input voice streams were digitally mixed at the core switching node and output as a single RTP stream. No frame-level binding of signaling and media was performed. The back-end transcription system only recognized based on this mixed audio stream. The experimental group of this invention adopted the media stream orthogonal split gateway and service status mapping unit of this invention to perform non-mixed parallel logical channel separation for 8 concurrent signal sources, and enabled the protocol frame-level encapsulation unit to inject service fingerprint codes in real time. At the same time, the unified time base reshaping logic on the gateway side was activated. The experiment ran continuously for 24 hours. During this period, the high-pressure scheduling scenario of emergency circuit breaker control of the whole network was simulated. Each terminal interacted at a high frequency of 12 instructions per minute according to the preset script. The transcription engine uniformly adopted the current mainstream speaker-independent continuous speech recognition (ASR) model to process the two groups of output media streams. The definition and statistical results of key performance indicators (KPIs) are shown in Table 1 below.

[0028] Table 1: Performance Comparison Data under Different Operating Conditions

[0029] Regarding word error rate, the sample from this invention exhibits stability under complex conditions of strong noise and packet loss. In contrast, the control group, due to the mixing operation, experiences a sharp deterioration in the overall signal-to-noise ratio because background noise from each source is superimposed at the physical layer. Furthermore, single-channel packet loss disrupts the overall frame structure of the mixed stream, resulting in a word error rate as high as 48.7%. Conversely, this invention, through an orthogonal splitting mechanism, isolates the noise and packet loss effects of each source at the physical layer, enabling the transcription engine to identify each clean, independent stream, keeping the error rate at an extremely low level of 5.8%. This verifies the decisive advantage of the source-end entropy lossless preservation mechanism in terms of noise and interference resistance. Regarding identity attribution error rate, the control group… The group almost failed in scenarios with multiple people talking at once (error rate 65.3%). This is because traditional voiceprint recognition algorithms have difficulty separating independent identity features from overlapping speech. However, the sample group of this invention, relying on the deterministic business fingerprint code injected by the protocol frame-level encapsulation unit, achieved zero identity misjudgment (0.0%) under any degree of overlap. This shows that sinking the business identity from the soft association of the application layer to the hard binding of the transport layer solves the identity loss problem under the signaling and media separation architecture. In terms of the timing inversion rate, facing the heterogeneous network of fiber optic (low latency) and satellite (high latency), the control group frequently experienced logical errors caused by differences in physical arrival time (18.5%).

[0030] Example 3: This example combines Figures 1 to 2 A description of a real-time intelligent voice transcription system for power dispatching calls, such as... Figure 1As shown, after the signaling and media stream source outputs data in the scheduling and switching network, the service status mapping unit on the left receives SIP signaling messages, performs operations to parse the SIP signaling and extract the physical address, and generates power grid service attribute codes. It then establishes a session-level source identity mapping table. The media stream orthogonal splitting gateway on the right intercepts the original RTP voice packets. It integrates core logic for performing non-mixed channel separation and zero-copy memory, as well as unified time base reshaping logic for maintaining a virtual reference clock and rewriting the RTP timestamp field. It is also equipped with an adaptive load suppression module that monitors signal energy and sends silence descriptor frames. This outputs physically isolated and time-synchronized independent parallel logical transmission channels. The split independent logical channels enter the protocol frame-level encapsulation unit, which retrieves the service attribute codes and the aforementioned mapping table, writes the service attribute codes into the RTP header extension to form a deterministic identity label, and finally outputs an independent media stream carrying the identity label to the transcription engine interface. This interface receives the independent media stream and calls the acoustic model for processing.

[0031] like Figure 2 As shown, the horizontal axis represents the number of concurrent sessions, with values ​​ranging from 0 to 500 (in integers), and the vertical axis represents the processing time in milliseconds. The graph contains three main data curves, corresponding to the business status mapping response time (dashed line), the media stream processing latency (solid line), and the protocol encapsulation time (dotted line). All three show an approximately linear upward trend with the increase of the number of concurrent sessions. Under the same number of concurrent sessions, the business status mapping response time is always higher than the media stream processing latency, while the protocol encapsulation time remains at the lowest level.

[0032] Example 4: This example aims to address the potential risk of insufficient stability in transcription systems under network impairment conditions by proposing and verifying a forward error correction mechanism based on adaptive redundancy coding. In high-risk scheduling scenarios such as emergency network shutdowns, the instantaneous packet loss rate of microwave and satellite links may exceed the conventional design threshold of 5%, leading to frame loss in the audio stream and subsequent semantic breaks in the transcribed content. To solve this problem, this example integrates adaptive forward error correction logic into the media stream orthogonal splitting gateway and monitors the packet loss rate and round-trip latency of each logical transmission channel in real time. When the packet loss rate of a certain channel continuously exceeds a preset threshold, such as 3%, the system automatically activates the adaptive redundancy coding mode. In this mode, while the gateway copies and splits the original voice data packets, it generates a corresponding proportion of redundant data packets based on the dynamic changes in the packet loss rate. These redundant packets are encoded using low-density parity check codes or Reed-Solomon codes and interleaved with the original RTP data packets for transmission.

[0033] To verify the effectiveness of this mechanism, a test environment including a network impairment simulator was constructed. Random packet loss of up to 10% was injected into the simulated satellite link. Experimental results showed that the control group without adaptive forward error correction logic experienced a word error rate soaring to over 15%, with frequent keyword loss. In contrast, the sample group with this logic enabled recovered the vast majority of lost voice frames through the error correction decoding algorithm at the receiver, stabilizing the word error rate below 6% and ensuring the semantic integrity of scheduling instructions under high packet loss conditions. Furthermore, this logic implements differentiated redundancy strategies for voice streams with different service roles. For channels identified as having core scheduling roles, the system allocates higher redundancy, such as 50%, to maximize the reliable transmission of critical instructions. For channels with non-core roles, lower redundancy, such as 10%, is allocated to balance bandwidth consumption and transmission quality. This service-value-based adaptive error correction mechanism improves the system's survivability and service quality in extreme network environments without increasing the overall network load.

[0034] Example 5: To address the potential differences in acoustic environment and network topology that power dispatch communication systems may face when deployed in different provincial power grids, as well as the audio sampling baseline drift problem caused by different batches of hardware equipment, this example provides a standardized pre-deployment calibration procedure to ensure the system's performance consistency in any physical environment. After the media stream orthogonal split gateway is connected to the dispatch switching network, the procedure automatically starts the acoustic baseline calibration program. The program injects preset standard white noise and frequency scanning signals into each logical transmission channel through the gateway's built-in signal generator, and uses the protocol frame-level encapsulation unit to re-sample the signal samples transmitted through the physical link. The system calculates the signal-to-noise ratio, frequency response curve, and nonlinear distortion of the re-sampled signal, and automatically generates audio compensation filter coefficients for the physical site based on the preset acoustic model. These coefficients are then written into the gateway's digital signal processing core to offset the impact of environmental noise and hardware distortion on voice quality during subsequent real-time transcription.

[0035] To address the potential database mapping latency uncertainty during the generation of power grid business attribute codes, this embodiment constructs an offline mapping table filling and verification procedure. Before the system is officially launched, the business status mapping unit sends simulated call requests (INVITE) covering all substations, dispatch consoles, and equipment intervals across the entire network to the system in batches through a simulated signaling generator. The system records the processing time for each simulated request from signaling parsing to the generation of the final business fingerprint code, and verifies the consistency between the generated fingerprint code and the preset true value. If the mapping time for a certain type of business object exceeds the preset real-time threshold, such as 10ms, or if the fingerprint code has a mapping error, the system will automatically trigger the optimization and reorganization of the index structure of the power dispatch management database (TMS) or fine-tune the parameters of the hash algorithm of the mapping table until the mapping response of all business objects meets the real-time and accuracy requirements.

[0036] Example 6: This example addresses the challenges of acoustic parameter adaptation uncertainty and network topology heterogeneity that power dispatch communication systems may face during initial engineering deployment and subsequent large-scale expansion. It provides a standardized system initialization calibration and adaptive parameter optimization procedure to eliminate system performance drift caused by batch differences in hardware equipment, changes in the physical acoustic environment, and fluctuations in network link characteristics. This ensures the determinism and consistency of the state-media orthogonal encapsulation mechanism in any deployment scenario. After the system is first powered on and connected to the dispatch switching network, the media stream orthogonal splitting gateway executes the acoustic link baseline calibration process. The gateway's built-in signal generator sends signals to each physical link... The connected logical transmission channels are sequentially injected with a standard test signal sequence, which includes a frequency sweep signal covering 300Hz to 3400Hz and white noise pulses with a preset energy gradient. The gateway uses the retrieval function of the protocol frame-level encapsulation unit to capture the response signals fed back from the physical terminal and transmission link, and calculates the flatness of the frequency response curve, total harmonic distortion, and signal-to-noise ratio baseline of each channel. Based on the above measurements, the gateway calls the adaptive filter coefficient generation algorithm to calculate the audio compensation parameter matrix for the current physical environment, and writes the matrix into the register of the digital signal processing core to establish an acoustic standardization baseline for the deployment site.

[0037] In response to the inconsistency in transmission delay jitter caused by differences in network access methods (such as fiber optic, microwave, and satellite) at substations of different voltage levels, this embodiment constructs an active detection and modeling procedure for a network impairment feature database. During the system initialization phase, the gateway initiates a series of ICMP probe messages containing high-precision timestamps to each remote communication terminal, continuously monitoring and recording the round-trip time (RTT) distribution characteristics and packet loss rate fluctuation trends of each link under different load conditions. Based on the probe data, the system uses statistical algorithms to construct a jitter probability density model for each link and calculates the virtual reference clock base time required for unified time base resetting logic accordingly. and jitter buffer bias The initial value ensures that the unified time base remodeling logic has the ability to align timing for the current network environment at the moment of system startup.

[0038] Example 7: To ensure the physical accuracy of the preset parameters for the unified time base resetting logic, the system, at the initial stage of session establishment, executes the link state initialization calibration procedure based on the law of large numbers and the principle of statistical process control to determine the parameters in the formula. and The initial values ​​are the ICMP or RTCP probe channel established between the gateway and the remote communication terminal. The output is the statistical characteristic parameters of the link after convergence. The steps are as follows: The gateway continuously sends a set of probe data packets to the target terminal at a 20ms interval. The round-trip time (RTT) and one-way arrival time difference of each probe packet are recorded, and the mean one-way delay of the collected samples is calculated. With delay standard deviation Based on Chebyshev's inequality, the confidence interval is set, and the system will... Set as Ensure the calculated Including dejitter buffer margin, the physical time correction value for the arrival of the first valid voice packet at the gateway is locked to [value missing]. The procedure quantifies uncertain network states into deterministic computational boundary conditions, eliminating the risk of audio frame loss or accumulation caused by random parameter settings.

[0039] Physical arrival time for heterogeneous network environments Random jitter occurs when the media stream orthogonal splitter gateway executes a virtual isochronous clock alignment procedure. Based on the G / D / 1 queuing theory model, the randomly arriving physical data stream is reshaped into a deterministic logical data stream. The procedure execution logic is as follows: Input state: A circular jitter buffer is allocated in the gateway memory, the depth of which is determined by the calibration procedure. Decision and queuing process: The network interface card captures raw RTP packets, and the system records the moment of physical interruption. The data packet is written to the end of the buffer, and the virtual clock advances: the gateway maintains a monotonically increasing virtual counter. The incremental beat is locked to the audio sampling rate. Dequeueing and Overriding: Virtual Counter When a read interrupt is triggered, the system retrieves the audio frame from the head of the buffer. At this point, the formula... In, variables The physical arrival time of the original RTP voice data packet at the gateway, and the random arrival time at the non-physical interface, are calculated using the formula in the procedure. Reflecting the logical isochronous timing with the gateway as the reference frame and after jitter reduction processing, if buffer overflow or underflow is detected, the system performs frame dropping or interpolation operations and dynamically adjusts... The logical timeline is reset to ensure that the downstream transcription engine receives a uniform and clean bitstream. A metadata decoupling and heterogeneous protocol adaptation procedure is implemented between the protocol frame-level encapsulation unit and the transcription engine interface. The composite RTP media stream is deconstructed into independent audio payload streams and business metadata streams. The interface program reads RTP data packets carrying identity tags, extracts the power grid business attribute encoding from the header extension fields using bitmasking operations, and converts it into a standard JSON format metadata object. The system strips the RTP header, extracts the G.711 or G.729 format payload, stores it in a linear PCM circular buffer, and calls the transcription engine's RecognizeStream method via the gRPC or WebSocket standard interface. During the connection handshake phase, the JSON metadata object is sent first to establish the session context. Streaming continuously sends PCM audio data. This procedure achieves the integration of power business identity information and general speech transcription capabilities without intruding on or modifying the commercial ASR engine kernel, ensuring atomic-level binding of identity and voice.

[0040] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0041] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A real-time intelligent voice transcription system for power dispatching calls, characterized in that, include: The service status mapping unit is used to access the signaling control plane of the dispatching and switching network, parse the session initialization protocol SIP signaling message during the call setup phase and extract the physical address identifier of the calling terminal, map the physical address identifier to the power grid service attribute code according to the preset power grid topology database, and establish a session-level source identity mapping table in the signaling control plane that contains the correspondence between the physical address identifier and the power grid service attribute code. The media stream orthogonal splitting gateway is connected in series in the media transmission plane of the scheduling and switching network and communicates with the service state mapping unit. It is used to obtain the source identity mapping table and intercept the original Real-Time Transport Protocol (RTP) voice data packets sent by multiple communication terminals belonging to the same session. The gateway executes unmixed channel separation logic to build parallel logical transmission channels equal to the number of communication terminals, and uses a zero-copy memory mechanism to copy and split the original RTP voice data packets to the corresponding logical transmission channels, thereby building a physically isolated and time-synchronized independent transmission environment for each voice data packet. The protocol frame-level encapsulation unit is located at the output port of the logical transmission channel. Under the closed transmission constraints of the logical transmission channel, it retrieves the power grid service attribute code corresponding to the communication terminal to which the current channel belongs from the source identity mapping table, and writes the power grid service attribute code into the protocol header extension field of the original RTP voice data packet using the header extension format defined by the RFC5285 standard, forming an independent media stream output carrying a deterministic service identity label.

2. The real-time intelligent voice transcription system for power dispatching calls according to claim 1, characterized in that, When the media stream orthogonal splitter gateway intercepts voice streams sent simultaneously by multiple communication terminals, it maintains the independent encoding format and timing characteristics of each voice stream, does not perform digital mixing and overlay processing of multiple audio waveforms, and independently outputs a single audio stream in each logical transmission channel.

3. The real-time intelligent voice transcription system for power dispatching calls according to claim 1, characterized in that, The protocol frame-level encapsulation unit encapsulates the power grid service attributes into fixed-length binary tags using network byte order. These binary tags are sent frame by frame with each RTP voice data packet, ensuring that the downstream receiver can parse the unique owner of the current voice frame under any network latency.

4. The real-time intelligent voice transcription system for power dispatching calls according to claim 1, characterized in that, The media stream orthogonal splitter gateway also includes unified time base reshaping logic, which maintains a monotonically increasing virtual reference clock for the current session in memory; When the gateway receives a raw RTP voice data packet sent by any communication terminal in kernel mode of the network interface card, it records the physical arrival time and calculates the relative offset between the physical arrival time and the virtual reference clock. Before routing the original RTP voice data packets to the logical transmission channel, the gateway rewrites the timestamp field in the RTP protocol header using a relative offset. This rewritten timestamp field represents the logical alignment time of each communication terminal at the gateway.

5. A real-time intelligent voice transcription system for power dispatching calls according to claim 1, characterized in that, The media stream orthogonal splitter gateway stores a differentiated encoding strategy table based on business roles. This strategy table defines the pass-through strategy corresponding to core roles and the transcoding strategy corresponding to non-core roles. The gateway matches the encoding strategy for each logical transmission channel based on the role field in the power grid service attribute code extracted by the service status mapping unit; for logical transmission channels that match the pass-through strategy, the gateway keeps the payload encoding format of the original RTP voice data packets unchanged. For logical transmission channels that match the transcoding strategy, the gateway calls the digital signal processing core to transcode the payload of the original RTP voice data packets into a high compression ratio format and modifies the payload type field in the RTP protocol header.

6. A real-time intelligent voice transcription system for power dispatching calls according to claim 5, characterized in that, When the media stream orthogonal splitter gateway detects that a communication terminal corresponding to a non-core role has acquired the session speaking right signaling, it triggers the corresponding logical transmission channel to switch from the transcoding strategy to the pass-through strategy.

7. The real-time intelligent voice transcription system for power dispatching calls according to claim 1, characterized in that, The media stream orthogonal splitter gateway also includes an adaptive payload suppression module, which monitors the voice signal energy value in each logical transmission channel in real time. When the signal energy value of a certain logical transmission channel is detected to be continuously lower than the preset background noise threshold, the module controls the gateway to stop sending voice payload data packets for that channel and sends silent descriptor frames containing only protocol header extension fields at preset intervals to maintain the heartbeat connection of the service identity label.

8. A real-time intelligent voice transcription system for power dispatching calls according to claim 1, characterized in that, The data structure of the power grid business attribute encoding includes the substation unique identification code, dispatch console seat number, and equipment bay number. The business status mapping unit obtains the business attribute data by parsing the INVITE or INFO message body of the SIP protocol.

9. A real-time intelligent voice transcription system for power dispatching calls according to claim 4, characterized in that, The unified time base refactoring logic uses the following formula to calculate the timestamp value used to rewrite the RTP header. : ,in, The physical time when the original RTP voice data packet arrives at the gateway. The reference time for the virtual reference clock. The sampling rate of the current audio encoding. This is the preset jitter buffer bias.

10. A real-time intelligent voice transcription system for power dispatching calls according to claim 1, characterized in that, The system also includes a transcription engine interface located downstream of the logical transmission channel. This interface is used to parse the RTP header extension fields of independent media streams to extract power grid service attribute codes, and to call the corresponding acoustic model to perform speech recognition on the independent media streams based on the codes.

Citation Information

Patent Citations

  • Multimedia dispatching communication seat

    CN201130995Y