Telecommunications switch infrastructure for real-time audio interpretation and recording via artificial intelligence agents

The carrier-grade switch-type system embeds AI agents within the network infrastructure to deliver real-time multilingual speech conversion and other AI services with low latency and security, addressing the limitations of cloud-based solutions and enhancing network efficiency.

WO2026015881A1PCT designated stage Publication Date: 2026-01-15CUNNINGHAM CHERYL

Patent Information

Application Number
PCT/US2025/037422
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2025-07-11
Publication Date
2026-01-15

Smart Images

  • Figure US2025037422_15012026_PF_FP_ABST
    Figure US2025037422_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A telecommunications switch-type platform positioned between a private-branch exchange and a data network dynamically hosts artificial-intelligence agents that interpret, transcribe, translate and transliterate live call traffic while the session is in progress. Machine-readable instructions stored in local memory instantiate, migrate and terminate the agents on demand across dedicated processing resources such as field-programmable gate arrays, application-specific integrated circuits or system-on-chip devices, thereby maintaining end-to-end latency below 250 milliseconds. An integrated call-audio recorder captures bidirectional media streams for secure archiving without interrupting service. The platform exposes a network interface that passes both signalling and media traffic, enabling seamless deployment inside call-centre or enterprise environments and compatibility with softswitch architectures conforming to class-4 Computational Interpretation of Communication using Artificial Intelligence standards. The same functional stack is deliverable as a computer-implemented method and as a non-transitory machine-readable medium storing the instructions executed by the processing resources.
Need to check novelty before this filing date? Find Prior Art

Description

TELECOMMUNICATIONS SWITCH INFRASTRUCTURE FOR REAL-TIMEAUDIO INTERPRETATION AND RECORDING VIA ARTIFICIAL INTELLIGENCE AGENTSTECHNICAL FIELD:

[0001] The present disclosure relates to telecommunications infrastructure and, more particularly, to a carrier-grade switch-type system that embeds and distributes artificial-intelligence agent-driven services — including real-time multilingual speech conversion, transcription, and other network-based Al functions — across terrestrial and space-air-ground integrated communication networks.PRIORITY CLAIM

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 670,683, filed on July 12, 2024, the contents of which are incorporated herein by reference.BACKGROUND ART

[0003] The present disclosure relates to telecommunications infrastructure and, more particularly, to a carrier-grade switch-type system that embeds and distributes artificial-intelligence agent-driven services, including real-time multilingual speech conversion, transcription and other network-based Al functions across terrestrial and space-air-groundintegrated communication networks.

[0004] Global telecommunications networks are experiencing explosive growth in traffic driven by 5G / 6G radio, low-Earth-orbit satellites and massive loT deployments. Operators must raise capacity, lower latency and retire costly circuit-switched plant while introducing new revenue-bearing services under tight capital constraints.

[0005] At present, value-added functions such as real-time language translation or fraud analytics are usually delivered “over-the-top” from distant cloud farms. Each extra hop consumes bandwidth, exposes sensitive media to multiple trust domains and pushes end-to-end delay beyond the roughly 250 ms that users perceive in conversation. Moreover, mainstream practice still frames artificial-intelligence workloads primarily as a computing problem, whereas emerging evidence shows that many Al tasks are better solved as a networking function placed directly in the traffic path.DISCLOSURE OF INVENTION

[0006] Although soft-switches and service-based architectures are known, the art lacks a carrier-grade platform that embeds, routes and secures autonomous Al agents inside a tandem or class-4 switch so that any present or future AUAGI service — speech conversion, threat detection, edge analytics, large-language-model inference and the like — can be delivered in-network at sub-millisecond scale without upgrading subscriber equipment. The present disclosure addresses that unmet need.

[0007] In one embodiment, a carrier-grade switch-type infrastructure — for example, a class-4 soft-switch or functionally equivalent tandem node — hosts a microkernel or heterogeneous-agent runtime that instantiates, routes and migrates artificial-intelligence agents for each active communication session. An originating-side agent receives voice, video or data packets, invokes one or more Al services (speech-to-text, fraud detection, image analytics, intrusion prevention, large-language-model inference, etc.), and republishes the transformedpayload to target-side agents over a secure publish / subscribe message bus. Agents execute inside sandboxed runtime adapters on CPUs, FPGAs, ASICs or SoCs and may serialize or live-migrate across a distributed mesh of terrestrial and space-air-ground integrated communication links. Latency for at least speech-conversion use-cases is held below -200 ms end-to-end, enabling real-time, one-to-one, one-to-many, many-to-one and many-to-many topologies.

[0008] Integrated call-audio (or media-stream) recorders capture both raw and Al-modified streams together with time-aligned metadata for compliance, billing, audit and machine-learning refinement. A hierarchical blockchain wallet protects each agent’s axiom repository while a zero-trust security manager gates access to compute, storage and network resources. Because the Al-service layer is embedded at the switch, carriers can licence the platform universally without upgrading subscriber equipment, collapsing legacy hardware footprints, lowering power draw, and opening recurring revenue channels for any present or future AI / AGI workload delivered “in-network.”BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Claimed subject matter is particularly pointed out and distinctly claimed in the concluding portion of the specification. However, both as to organization and / or method of operation, together with objects, features, and / or advantages thereof, it may best be understood by reference to the following detailed description if read with the accompanying drawings in which:

[0010] Figure 1 illustrates a multi-layered telecommunications infrastructure (2500) designed to support real-time or near real-time computational language conversion using Al agents. At the top, the Geosynchronous Earth Orbit (GEO) satellite layer (2510) provides broad, high-latency global coverage. Below it, the Medium Earth Orbit (MEO) layer (2520) offers a balance between coverage and latency while the Low Earth Orbit (LEO) cubesats layer (2530) delivers low-latency, high-speed data transmission ideal for real-time communication. Thisspace-based network is integrated with a terrestrial layer (2540) encompassing loT consumer goods and gaming applications enabling connectivity for smart devices and interactive platforms. Supporting mobile and transport-based connectivity, the infrastructure includes aircraft (2532), vehicles and fleets (2542) and maritime systems (2544). End-user access is enabled through 5G / 6G smartphones and internet-enabled devices forming a cohesive global communication ecosystem capable of supporting seamless, Al-driven multilingual interaction across diverse domains and platforms.

[0011] Figure 2 depicts a telecommunications network (2200) that operates in a one-to-one configuration, enabling real-time or near real-time language conversion between individual users. In this implementation a Japanese speaker (330) communicates with an Arabic receiver (2620) through a network-enabled translation process. The system performs translation of foreign text to Japanese text (2602) and foreign text to Arabic text (2604) depending on the direction of communication. For audio-based interaction the translated text is further converted into spoken language audio files with foreign text translated to Japanese and converted to a Japanese .wav file (2612) for playback to the Japanese user and foreign text translated to Arabic and converted to an Arabic .wav file (2614) for playback to the Arabic user. This configuration allows seamless multilingual communication through automated translation and speech synthesis, enabling natural and efficient exchanges between users of different languages over the telecommunications network.

[0012] Figure 3 illustrates an example telecommunications network (300) operating in a one-to-many configuration where a single English speaker (310) communicates simultaneously with multiple recipients in different languages. The process begins with speech-to-text conversion (350) where the English audible speech is transcribed to text (352). The transcribed English text then passes through an interpretation and translation module (360) which translates the content into multiple target languages, including Japanese (2602), Arabic (2604), English refinement or repetition (2606), and Mandarin (2608). Following translation, the system utilizestext-to-speech synthesis (370) to generate spoken audio outputs in the respective languages: translated text + Japanese .wav file (302), translated text + Arabic .wav file (304) and translated text + Mandarin .wav file (306). These audio files are delivered to the respective receivers, including the Japanese receiver (2610), Arabic receiver (2620) and Mandarin speaker (320). This configuration demonstrates a multilingual broadcasting capability, allowing one speaker to engage with a diverse, multilingual audience through Al-powered translation and voice synthesis in real or near real time.

[0013] Figures 3A-3D illustrate a schematic block diagram of an example one-to-many telecommunications network configuration designed to support real-time multilingual voice translation using Al-enabled cloud services. The system includes various speakers English Actor (310), Arabic Actor (330), Mandarin Actor (320) and Japanese Actor (340) each connected via laptops (312, 314, 316, 318) equipped with HTML5 browsers, user-specific channel management agents and a backend processing system called JOSH which stores pre- and post-translation text and .wav files in the DOM. The English speech is first processed through a Speech-to-Text GA (322) for low-latency, streaming transcription and trickled as JSON data to databases. This transcription is then handled by a Language Translation GA (324, 326) to convert text into target languages like Arabic, Mandarin, and Japanese. Each translated text is synthesized using a Text-to-Speech GA producing natural-sounding audio output in the respective language formats (e.g., Arabic.wav + Arabic.txt). The system relies on HTTP POST requests to IBM Watson endpoints for speech recognition and synthesis, and stores results in graph databases for personal and regional LLMs. A publish-subscribe message bus enables real-time updates and all data exchanges use secure APIs with authorization tokens. This architecture enables scalable multilingual communication where one speaker can be understood simultaneously by multiple recipients in different languages.

[0014] Figure 4 illustrates a telecommunications network (400) configured to operate in a many-to-one configuration enabling multilingual inputs from various speakers to beinterpreted and delivered to a single receiver in their preferred language. In this setup, users such as a Mandarin speaker (320) a Japanese receiver (2610) and an Arabic receiver (2620) communicate toward a centralized English receiver (2630). Spoken input from the Mandarin speaker undergoes speech-to-text conversion (350) with Mandarin audible speech transcribed to text (358). The resulting text is processed through an interpretation and translation module (360) where it is translated into multiple languages including English (2606), Japanese (2602) Arabic (2604) and Mandarin (2608). Each translated text is then converted into an audio file using text-to-speech synthesis (370) producing outputs such as translated text + English .wav file (308) Japanese .wav file (302) and Arabic .wav file (304). These audio outputs are directed to their respective recipients ensuring that regardless of the source language, the final message is received in the listener’s native or selected language. This configuration supports converged multilingual communication toward a common recipient, streamlining global collaboration and multilingual conferencing.

[0015] Figure 5 presents a telecommunications network operating in a many-to-many configuration enabling real-time multilingual communication among speakers and receivers of different languages. Participants include an English speaker (310), Mandarin speaker (320), Japanese speaker (330) and Arabic speaker (340). Each speaker’s audible speech is first transcribed via speech-to-text modules Japanese (502), Mandarin (504), English (506) and Arabic (508). The transcribed text is then routed to a translation module which translates it into various target languages: to English (2606), to Mandarin (2608), to Japanese (2602) and to Arabic (2604). Each translated text is then processed through text-to-speech conversion resulting in audio files: Japanese .wav (302), Arabic .wav (304), Mandarin .wav (306) and English .wav (308). These outputs are delivered to the corresponding receivers English receiver (2630), Japanese receiver (2610), Arabic receiver (2620) and Mandarin receiver (2640) ensuring that each participant receives communications in their native or selected language. Thisconfiguration demonstrates a highly integrated Al-powered network capable of facilitating seamless, real-time multilingual dialogue across global participants.

[0016] Figure 6 illustrates a schematic block diagram (600) of an example Al-agent runtime adapter system, highlighting how Al agents perceive, process and act within a dynamic environment. At the core is the Al agent runtime environment (630) implemented within a wavefront array (SISD Single Instruction, Single Data) architecture which facilitates efficient sequential decision-making. The system begins with agent sensors (610) capturing inputs from the environment (608) interpreted as percepts (606) signals that inform the agent of “what the world is like now” (602). These inputs are processed using condition-action rules (606) to decide “what action I should do now” (604). The system maintains and updates internal states such as current state (612a) how the world evolves (612b) and what the agent’s actions do (612c). The outcomes are carried out via actuators completing the perception-decision-action loop. This runtime framework enables adaptive, context-aware responses from Al agents operating in real-time telecommunications or language processing environments.

[0017] Figure 7 compares two architectural models of operating systems: a monolithic kernel -based system (710) and a microkernel-based system (720). In the monolithic kernel architecture components such as application system calls, virtual file system (VFS), interprocess communication (IPC), file system, scheduler, virtual memory, device drivers and dispatcher all operate within the kernel mode providing direct and high-performance access to hardware. Both user mode and kernel mode interact closely, but this tightly coupled design can be less stable or secure if one component fails. In contrast, the microkernel architecture (720) separates key functionalities into isolated modules. Here applications, IPC, UNIX server, device drivers and file servers operate primarily in user mode while only essential services such as basic IPC, virtual memory management and scheduling remain in kernel mode. This separation enhances system stability, modularity and security by minimizing kernel responsibilities and isolating failures.

[0018] Figure 8 presents a schematic of a heterogeneous Al-agent runtime adapter (610) which may be implemented, either fully or partially within a uniprocessor architecture such as SISD (Single Instruction, Single Data). The system models the core perception-decision-action loop of an Al agent. Input is received from the environment (608) through sensors generating precepts that represent “what the world is like now” (602). The agent processes these precepts using condition-action rules (606) to determine “what action I should do now” (604). The chosen response is executed through actuators allowing the agent to interact with the environment. This runtime adapter supports adaptive, context-aware behavior, even within resource-constrained processing environments, by efficiently managing sensing, reasoning, and action cycles in a modular and unified system.

[0019] Figure 9 illustrates a model -reflexive Al-agent runtime adapter (620) which may be implemented either fully or partially within a wavefront array architecture such as SISD, MISD or similar computational structures. This advanced Al-agent framework expands on traditional perception-action models by incorporating self-awareness and predictive modeling. The agent receives input from the environment (608) as precepts (606) representing “what the world is like now” (602). Using condition-action rules (606) and internal state representations (612a) the agent determines “what action I should do now” (604). Additionally, the agent maintains models of how the world evolves (612b) and what my actions do (612c) enabling forward-looking, reflexive decision-making. Actuators then carry out the selected actions in the environment. This configuration allows the agent not only to react to current stimuli but also to adapt based on predicted outcomes, leading to more intelligent and context-aware behavior in real-time or near real-time operations.

[0020] Figure 10 illustrates a goal -type Al-agent runtime adapter (1000) which may be implemented, wholly or partially, within a wavefront-array architecture (e.g., SISD, etc.). This Al-agent framework is designed to support goal-driven decision-making. The agent interacts with the environment (608) by receiving precepts forming an understanding of “what the worldis like now” (602). It maintains an internal state (612a) and uses models to predict how the world evolves (612b) and what its actions do (612c). Before acting, the agent evaluates potential outcomes, such as “what it will be like if I do action A” (614) and compares them against its defined goals (616). Based on this analysis, the agent decides “what action I should do now” (604) and executes the selected action via actuators. This architecture enables the Al agent to plan purposefully, anticipate consequences and act intelligently toward achieving specific objectives in dynamic environments.

[0021] Figure 11 illustrates the architecture of a cybersecurity Al-agent system (2700) that integrates intelligent mobile agents, game theory, and distributed detection to defend against network attacks. A network security officer (2702) oversees operations, supported by an intrusion detection server (2704) and a mobile agent platform factory (2714). These platforms deploy mobile agents through migration mechanisms (2706) across a network mode (2712) that includes multiple nodes such as PCI, PC2, and PC3. The system uses an intrusion detection processor trained on the KDD dataset classifying network traffic into categories like normal, smurf, Neptune, back and multihop with detection volumes reaching over 143,000 events. The agents monitor the network using sniffer tools and detection engines and interact with attackers through a multi-step process: (1) detecting the attack (2) engaging in a non-cooperative game to analyze adversarial behavior (3) computing a risk value using Nash equilibrium (4) entering a cooperative game to coordinate defense (5) activating Working Group Agents (WGAs) (6) computing Shapley values to evaluate contribution, (7) generating and sending an attack report and (8) deploying Local Agents (LAs) to threatened nodes for direct intervention. This system enables real-time, intelligent and collaborative cyber defense across dynamic and distributed network environments.

[0022] Figure 12 illustrates a utility-type Al-agent runtime adapter (1200) which may be implemented, in whole or in part, within a wavefront array architecture such as SISD, MISD or MIMD. This agent model is designed to make intelligent, goal-aligned decisions based on utilityoptimization. The agent receives precepts from the environment (608) to assess “what the world is like now” (602) and maintains an internal state (612a). It uses models to understand how the world evolves (612b) and what its actions do (612c). The agent predicts “what it will be like if I do action A” (614) for various possible actions and evaluates each future state using a utility function (618) which determines “how happy I will be in such a state” (619). The action that yields the highest utility is selected and executed through actuators. This architecture enables Al agents to reason not only about outcomes but also about the desirability of those outcomes, making decisions that maximize satisfaction or benefit within complex, dynamic environments.

[0023] Figure 13 depicts an unsupervised and / or supervised learning Al-agent runtime adapter (1300) which may be implemented, in whole or in part, within a wavefront-array architecture (e.g., SISD, MISD, MIMD, etc.). This learning-oriented agent continuously improves its performance through interaction with the environment (608). The agent receives percepts from the environment via Critic Sensors (1302) which compare the agent’s behavior to a predefined performance standard and provide feedback. The Learning Element (1304) uses this feedback and defined learning goals to update the agent’s internal models and decision-making strategies. To support learning, a Problem Generator (1306) introduces new experiments or challenges, encouraging exploration and adaptation. The Performance Element (1308) applies learned behavior through effectors, producing actions that impact the environment, which in turn may change as a result. This adaptive feedback loop enables the agent to refine its behavior over time, whether through supervised learning (with feedback based on correct outcomes) or unsupervised learning (pattern discovery without explicit labels), enhancing autonomy and effectiveness.

[0024] Figure 14 illustrates an example multi-agent platform (1400) structured in a layered architecture to support intelligent task execution and communication across distributed components. The platform is divided into three functional layers: The Superior Layer (1402) the Intermediate Layer (1404) and the Reactive Layer (1406). The Superior Layer managescommunication by protocol and performs high-level planning according to tasks often handled by superior agents or super users. The intermediate layer focuses on task-based planning through message passing communication facilitating coordination among agents via structured messages, orders, and information exchanges. The reactive layer (1406) handles centralized planning and real-time action / perception where reactive agents operate locally within locality a, b, and c responding to environmental stimuli. Components such as component a, b, and c and connectors allow the platform to dynamically create new components or connectors enabling scalability and adaptability. Both normal users and super users interact with the platform by sending user messages, managing new connections or deploying new agents. This hierarchical, modular design allows the system to perform complex, coordinated tasks in a flexible and scalable multi-agent environment.

[0025] Figure 15 illustrates a system (1500) comprising an aggregator Al-agent architecture designed to process uncertain sensor data, generate predictions and support decision-making. Within the system, machine A (1510) operates in environment A (1520) and is monitored by a sensor agent (1502) which collects sensor data with uncertainty. This raw data is passed to the aggregator agent (1506) which combines and processes it into aggregated data subsequently stored in a historical database (1514). A predictor agent (1504) utilizes this historical and current aggregated data to generate predictive models, which are refined through a model trainer agent (1512). The trained models are then presented through a display model and visualized to users via the user interface agent (1516). Finally, insights and model outputs inform the decision maker agent (1508) enabling intelligent, data-driven decisions. This integrated agent-based system allows for robust handling of uncertain data, continuous learning and real-time interaction between human users and Al-driven processes.

[0026] Figure 16 depicts various aspects of a trusted platform module (TPM)-basedAl-agent system (1600) designed to enhance cybersecurity in embedded or networked environments. The system is built to detect and respond to threats such as hacking (1602),spoofing (1604) and data falsification (1606). Central to the architecture is the TRN chip (1620) which works in conjunction with crypto software (1622) and secure boot software (1624) to establish a secure execution environment. Critical cryptographic operations like sealing, signing and sealed-signing (1610) protect data and ensure integrity. The system monitors real CAN packets (1632) and uses feature extraction (1634) to identify patterns. These features are fed into a classifier (1636) that performs a normal vs. attack decision (1638). For accurate detection, the system is trained using labeled CAN packets (1640) processed through a DNN (Deep Neural Network) structure (1644) comprising input layers, hidden layers, and output layers (1642). This Al-agent architecture combines trusted hardware and deep learning to enable secure, intelligent and adaptive cyber threat detection in real time.

[0027] Figure 17 (1700) presents a comparative diagram of traditional PBX (Private Branch Exchange) and hosted VoIP PBX telecommunications approaches illustrating the transition from legacy systems to cloud-based telephony. On the left side, under "How it's done today (2025)" the setup begins at a Customer Office (1710) using analog phones (1712) and IP phones (IP 331) connected to an on-premises PBX system (1714). The PBX routes calls through a T-l line or POTS lines (1750) into the public switched telephone network (PSTN) (1720) with Monmouth voice switch linking to long-distance carriers like LD Carrier 1 and LD Carrier 5. Calls pass through a switch (1725) and are managed via a firewall (1716) and router (1718). On the right side, under "How it’s done after 2030" the setup shifts to a hosted VoIP PBX Solution (1760) where the same customer devices connect over the Internet with QoS (Quality of Service) to a cloud-hosted PBX system. This model eliminates the need for on-site PBX hardware allowing calls to route through cloud infrastructure connected to the Monmouth voice switch (1780) and long-distance carriers, preserving PSTN access but modernizing infrastructure. This shift enhances scalability, cost-efficiency and flexibility in business communications.

[0028] Figure 18 (1800) illustrates a call center use case in which CICAI (Cognitive Intelligent Communication Al) is deployed at the entry point to customer premises, enabling intelligent routing, recording and integration with cloud services. The call flow begins from the Public Switched Telephone Network (PSTN) (1804) or Plain Old Telephone Service (POTS) lines, reaching the customer (1802) via T1 / T3 or E1ZE3 trunks (1806). These lines terminate at the brick-and-mortar demarcation point (1808) where the call audio (1810) enters the system. The audio may be converted to VoIP call audio (1812) and processed by the CICAI route-switch processor (1816) which dynamically routes calls based on Al logic. Calls are further managed by a CTI / ACD server (1814) for intelligent call distribution and integrated with CRM systems hosted in the cloud (1824). Audio is also captured by an audio record server (1818) for compliance or analytics. Within the corporate LAN / WAN (1822) running TCP / IP over 802.1 the system connects to contact center employees (1820) using workstations supported by software applications (1826) and linked to a media gateway or PBX system (1828). This architecture blends traditional telephony with modern Al, cloud services, and IP networking to deliver efficient, scalable and intelligent call center operations.

[0029] Figure 19 depicts a CICAI-powered call flow architecture deployed at the entry point of customer premises, integrating traditional telephony, cloud Al and real-time communication protocols to streamline and personalize call center interactions. The process begins when a customer (1802) accesses a web portal (1904), which coordinates with the Public Switched Telephone Network (PSTN) (1902). The portal generates a unique SIP URI (1951) and auth token (1952), returning these along with IP and port details (1904). A signaling gateway (1906) initializes low-latency communication protocols (e.g., RTP / SRTP, WebRTC, JREAP C, websockets) via client initialization (1907). The system retrieves the customer’s estimated wait time (1905) and returns this information to the customer device (1906). Once ready, the IVR / CTI / ACD system with a built-in co-browser (1910) instructs the customer's device to initiate a call to the assigned SIP URI of a contact center employee (1920). The CICAI core(1912) facilitates intelligent routing, integrates with the CRM system (1916) and draws on Personalized, Regional and Generalized LLMs (1918) for Al-enhanced interaction. Call audio (1914) is recorded for compliance, and secure signaling and media transport is enforced through SRTP / DTLS (1916). Sessions are terminated securely with URI and RTP session cleanup (1921, 1922). The entire system complies with RFC 3986 for SIP URI authorization (1902), creating a seamless, secure and intelligent communication path between customers and call center agents.

[0030] Figure 20 is schematic block diagram illustrating an embodiment of an example system that includes one or more server computing devices and loT-type devices (102) communicating through a network infrastructure (122). The system comprises loT devices, a base station transceiver (108) and a local transceiver (112) which establish communication with various servers (116, 118, 120) via communication links (124). The network (122) may include wired and / or wireless links and can operate over standard Internet Protocol (IP)-type infrastructure enabling communication either directly or through intermediary transceivers. In some implementations network 122 may also represent cellular communication infrastructure involving a base station controller or mobile switching center to support mobile device connectivity. The servers (116, 118, 120) are flexible in function and may include update servers, backend servers, archive servers, location servers, navigation servers, crowdsourcing servers, or network-related servers, among others. This system architecture supports diverse real-time loT operations such as data processing, remote updates, positioning, and communication management in distributed environments.

[0031] Figure 21 presents a schematic block diagram (200) of an example Internet of Things (loT)-type device illustrating its core components and functional architecture. At the center is the Processor (210) responsible for executing instructions and coordinating operations. The processor interfaces with memory (230) which stores data and program logic, including software / firmware code (232). The device interacts with its environment through various sensors (250) that collect data and counters / timers (260) that support time-based operations and eventtracking. A display (240) provides user-facing feedback or system status. Communication with other devices or networks is enabled via a communications interface (220) allowing the loT device to send or receive data wirelessly or through wired connections. This architecture enables intelligent, connected operations suitable for various loT applications such as monitoring, automation or remote control.

[0032] Figure 22 illustrates a schematic diagram (1100) of an example computing environment implementation showcasing how multiple devices interact over a network and share resources for processing and data management. The system includes interconnected devices First Device (1102), Second Device (1104) and Third Device (1106) communicating via a network (1108). Each device comprises core computing components including a processor (1120) responsible for executing instructions and a memory hierarchy consisting of primary memory (1124) for fast-access data storage, secondary memory (1126) for long-term storage and general memory (1122). Data input and output operations are managed through an input / output interface (1132) while communication across devices and systems is handled by a communications interface (1130). Additionally, the system utilizes a computer-readable medium (1140) which can store instructions or data used by the processor. This modular architecture supports distributed computing, remote resource access and scalable processing across networked systems.

[0033] Reference is made in the following detailed description to accompanying drawings, which form a part hereof, wherein like numerals may designate like parts throughout that are corresponding and / or analogous. It will be appreciated that the figures have not necessarily been drawn to scale, such as for simplicity and / or clarity of illustration. For example, dimensions of some aspects may be exaggerated relative to others. Further, it is to be understood that other embodiments may be utilized. Furthermore, structural and / or other changes may be made without departing from claimed subject matter. References throughout this specification to “claimed subject matter” refer to subject matter intended to be covered by one or more claims,or any portion thereof, and are not necessarily intended to refer to a complete claim set, to a particular combination of claim sets (e.g., method claims, apparatus claims, etc.), or to a particular claim. It should also be noted that directions and / or references, for example, such as up, down, top, bottom, and so on, may be used to facilitate discussion of drawings and are not intended to restrict application of claimed subject matter. Therefore, the following detailed description is not to be taken to limit claimed subject matter and / or equivalents.DETAILED DESCRIPTION

[0034] In this disclosure, the phrases “one embodiment,” “an implementation,” and comparable wording signify that the accompanying feature exists in at least one embodiment; repetition of the phrase elsewhere does not require every embodiment to include all such features. Unless an explicit exclusion is stated, any disclosed structures or characteristics may be combined where technically compatible, and the contextual discussion herein governs the reasonable scope of each relative term.

[0035] Growing traffic, spectrum scarcity, and quality-of-service expectations drive continual capacity upgrades and cost reductions; accordingly, the embodiments deploy task-specific Al agents at in-network compute points (e.g., class-4 switches or edge servers) so that new features — multilingual speech conversion, fraud analytics, threat detection — are delivered without replacing subscriber equipment.

[0036] In one embodiment, digitally and / or analog-interconnected networks host data-processing units that couple to external platforms to provide source-to-target language conversion, interpretation, translation, transcription, or transliteration, within < 200 ms end-to-end latency, thereby collapsing legacy equipment footprints and mitigating trade barriers caused by language friction.

[0037] A “Computational Interpretation of Communication using Artificial-IntelligenceAlgorithms” (CICAI) network is a distributed, multi-agent fabric in which processing unitsexecute neural -network and rule-based instructions; each agent can instantiate, migrate, or retire on demand, allowing class-4 or equivalent switches to deliver secure, real-time multilingual and analytic services entirely in-network.

[0038] Leveraging this architecture, a single CICAI deployment can sustain one-to-one, one-to-many, many-to-one, and many-to-many conversation flows — e.g., a General-Assembly-style event in which every delegate hears and speaks in a native language while experiencing no more than a 200 ms conversational delay.

[0039] The platform pipelines interpretation, translation, transcription, and transliteration concurrently; during live dialogue it attaches diacritic-based tone annotations to text transcripts so that downstream sentiment analysis preserves nuance, while cached transcripts permit post-session auditing without loss of meaning.

[0040] After converting speech from a source language to a target language, the system renders both the target text and a phonetic guide — using the target alphabet — on the user interface, so an Arabic-speaking caller who receives Mandarin output also sees the Mandarin characters with pronunciation cues.

[0041] By instantiating Al agents inside the traffic path, the disclosed switch eliminates the 6-8 s interpreter lag typical of human specialists, enabling diplomats or business partners to exchange culturally nuanced statements in a smooth, conversational cadence.

[0042] The interpretation module cross-references user profiles and regional-etiquette tables to suppress phrasing that could offend a target culture, thereby safeguarding diplomatic exchanges while retaining technical accuracy.

[0043] The translation layer integrates with the interpretation engine so that domain-specific terminology, such as loanwords or specialized acronyms — is resolved through context-aware ontologies; thus Gulf-dialect Arabic engineering terms are rendered accurately for a Moroccan audience without loss of technical meaning.

[0044] To address these challenges, at least in part, it may be advantageous to gather content about individuals, groups, organizations, etc., with their permission, of course. In an embodiment, content may be collected, such as from Linkedln, for example, or from other sources. Such example profile content may provide a sense of an individual’s education level, literacy level, languages spoken, and so forth. With such profile content (e.g., data, information, etc.), dialogue may be structured to an elevated academic level or a modest academic level, for example, depending at least in part on the individual. Note that may not be an attempt to train the individual, but rather a CICAI may respond to individual(s) involved in terms and abstractions that may be more readily comprehended, for example.

[0045] In one embodiment, processing units detect speech onset, sample the microphone signal, convert it through an analogue-to-digital converter, and store packets as .wav files. Phoneme extraction converts each .wav file to a .txt representation, which is then interpreted, transcribed, translated, or transliterated for the target listener while Unicode code points preserve vowel and consonant nuances.

[0046] During transliteration the engine generates candidate phonemes to confirm that a spoken phrase has been captured as intended by the speaker. An Al agent also consults user profiles that include cultural preferences and adapts the text accordingly; for instance, a “y’all” uttered by a Texas caller becomes “you all” for a New York recipient.

[0047] When speech is detected, separate processing threads for interpretation, transcription, translation, and transliteration launch concurrently. If one thread requires interim values from another, the scheduler applies microsecond-scale waits that remain imperceptible to human users and preserve conversational flow.

[0048] Cultural nuance is preserved by appending diacritics to the generated characters as each word arrives from the source language. These marks encode tone, timbre, and emotion, allowing the receiving application to render speech or text in a manner that is calming, formal, or informal as dictated by the target culture.

[0049] In a teleconference, each participant may speak and hear a preferred language while parallel Al agents track speaker identity and timestamp every utterance. Separate transcripts — both in the original language and in each participant’s selected language — are produced, eliminating confusion when several voices overlap.

[0050] For example, when three participants speak simultaneously, the system segregates the audio streams, tags each stream with a speaker label, and writes a timeline such as “00:01.2 s: Participant A ...” so any listener can replay or search the dialogue later. Multiple parallel transcripts may be stored, one per participant, in both source and target languages.

[0051] Profile-driven Al agents can also reshape the teleconference display in real time. The application may enlarge the window of a high-priority speaker, adjust microphone gain, mute background noise, or let a user rewind the last few seconds of a selected participant’s stream for clarification.

[0052] Ubiquitous service is achieved by embedding interpretation, translation, transcription, and transliteration functions within class 4 telecommunications switches so that the existing network fabric itself delivers real-time or near-real-time language services worldwide, as illustrated in FIGS. 17-18.

[0053] In this specification, “Artificial Intelligence” denotes hardware or software that senses inputs, produces outputs, and autonomously selects actions to reach a desired state, while an “Al agent” is an individual instance of such intelligence capable of acting independently within the disclosed system.

[0054] In some embodiments, CICAI includes, for example, cell switched, frame switched, or packet switched fabrics, and may use Transparent Interconnection of Lots of Links, TRILL, whose configuration effort is minimal. By way of illustration, a metro Ethernet ring that employs TRILL can be extended to an intercontinental backbone without changing customer edge equipment.

[0055] Some embodiments provide a plug and play CICAI chassis containing a plurality of processing units that are interconnected by high-capacity fabric; when the chassis is powered and connected, the units automatically join the existing control plane and begin translating traffic with only trivial distance-dependent propagation delays.

[0056] The term “real time” in this disclosure means that the round-trip delay between a spoken source phrase and its audible or textual presentation to a human listener is less than two hundred fifty milliseconds, which corresponds to a quarter of a second and sits below the threshold at which most people perceive a pause in conversation.

[0057] The CICAI platform includes at least one processing unit that executes language-specific and subject-specific artificial intelligence algorithms, including but not limited to recurrent neural networks, transformer networks, and rule-based grammars; any of these algorithms can be combined within an embodiment where technically compatible.

[0058] By distributing those units across terrestrial, maritime, and orbital nodes, CICAI enables continuous multilingual communication across geographic regions and cultural groups so that a traveler in a remote village can converse naturally with a support specialist located on another continent.

[0059] In some embodiments, CICAI delivers the foregoing services by integrating into standard networking protocols collectively or severally.

[0060] (a) Address Resolution Protocol, ARP.

[0061] (b) Asynchronous Transfer Mode virtual-circuit upper-layer control plane and data plane using one or more virtual channel connections.

[0062] (c) Border Gateway Protocol, BGP

[0063] (d) Bluetooth core profiles including SDP, TCS, AVCTP, OBEX, Link Management Protocol, BNEP, and RFCOMM.

[0064] (e) Dynamic Resource Allocation Protocol, DRAP.

[0064] (f) Domain Name System, DNS.

[0065] (g) Dynamic Host Configuration Protocol, DHCP.

[0066] (h) File Transfer Protocol, FTP, which allows interpretation and translation of stored documents.

[0068] (i) Hierarchical Satellite Routing Protocol, HSRP.

[0067] (j) Hypertext Transfer Protocol, HTTP, which supports interpretation and translation of web page content.

[0068] (k) Internet Control Message Protocol, ICMP.

[0069] (1) Inter- Satellite Link cross links, ISL.

[0070] (m) LoRaWAN, a low-power wide-area networking standard for long-range sensor traffic.

[0073] (n) Open Shortest Path First, OSPF.

[0071] (o) Plesiochronous Digital Hierarchy, PDH.

[0072] (p) Routing Information Protocol, RIP.

[0073] (q) Satellite Grouping and Routing Protocol, SGRP.

[0074] (r) Simple Mail Transfer Protocol, SMTP, which enables interpretation and translation of electronic mail.

[0075] (s) Synchronous Optical Network or Synchronous Digital Hierarchy, SONET or SDH.

[0079] (t) Satellite Transport Protocol, STP

[0076] (u) Transmission Control Protocol and Internet Protocol suite, TCP IP.

[0077] (v) Telnet, also referred to as Plain Old Telephone Service.

[0078] (w) User Datagram Protocol, UDP

[0079] (x) Wireless interfaces including WiFi, WiMax, LTE, fourth-generation, fifth-generation, and sixth-generation cellular access.

[0084] The cited embodiments can be implemented wholly or partly within a space-air-ground integrated network in which unmanned aerial platforms or satellites relay traffic among remote islands and large regional land masses while maintaining persistent coverage.

[0080] A CICAI implementation that targets state-of-the-art space-air-ground integrated networks may adopt stacked reference architectures such as those that follow.

[0081] (a) Optical Channel Layer of the Optical Transport Network model, layer one.

[0082] (b) A hybrid radio-frequency and optical link-layer model that employs ten gigabits per second or higher laser uplinks to medium-earth or geostationary-orbit satellites with high-rate Ka-band or Ku-band downlinks.

[0083] (c) Data-link and network layers of the Asynchronous Transfer Mode model.

[0084] (d) Data-link and network layers of the Open Systems Interconnection model.

[0085] (e) Network-interface layer of the Transmission Control Protocol and Internet Protocol model.

[0086] (f) Link and application layers of low-power wide-area networking models.

[0087] (g) Optical-switched broadband or satellite-broadband networks that use multigigabit-per-second optical cross-links with ranges up to six thousand kilometres.

[0088] (h) Transparent Interconnection of Lots of Links data-link layer of the Open Systems Interconnection model.

[0089] An example infrastructure that corresponds to FIG. 1 includes interconnected satellite nodes, terrestrial gateways, and edge compute clusters; nevertheless the subject matter is not limited to that illustration and may be adapted to any topology that satisfies latency, security, and quality-of-service constraints.

[0090] In another embodiment, a CICAI network operates in one-to-one, one-to-many, many-to-one, or many-to-many configurations so that speakers of English, Japanese, Mandarin, or any other supported language can exchange communication in real time while the platform detects and propagates context, sentiment, intention, and emotion.

[0091] Some embodiments deploy CICAI across several parallel-computing models, as set out in subparagraphs through:

[0092] (a) Single instruction stream single data stream (SISD): a single core or node executes every task, while microkernel agents and heterogeneous agents run in user space to add interpretation and translation services without changing the underlying instruction flow.

[0093] (b) Multiple instruction single data (MISD): separate cores apply different operations to the same voice or text packet; a pipeline can apply speech detection, language identification, then sentiment scoring in strict order so that aerospace or drone telemetry remains deterministic.

[0094] (c) Multiple instruction multiple data (MIMD): many processing units work asynchronously on independent data blocks; shared-memory and distributed-memory variants enable large-scale simulation or high-speed carrier switching, and hypercube or mesh interconnects favour mesh networks that host CICAI agents.

[0095] (d) Single instruction multiple data (SIMD): identical operations execute across numerous data points at once; systolic arrays accelerate neural -network inference for live speech conversion while agents coordinate through inter-process messaging.

[0096] CICAI can mix SISD, MISD, MIMD and SIMD within one system; when dedicated SIMD hardware is unavailable, a SWAR approach reuses general-purpose registers so that agents still gain data-parallel speed-up without special micro-instructions.

[0097] A parallel multi-agent architecture permits runtime adapters to load on any suitable core, cell or node, whether MISD, SISD, MIMD or SIMD, so that hierarchical agent control can run under a monolithic or embedded operating system to match a given application.

[0098] Hardware implements secure boot to verify firmware, kernels and drivers against trusted databases; the operating system extends that validation into each runtime adapter, blocking un-signed executables and common exploit tools while preserving agent mobility.

[0099] One embodiment forms a wavefront array or a synchronous systolic array of hardware or software agents, each using stochastic methods common in large language models; the arrangement, shown in FIG. 6, supports machine learning, deep learning and workflow management for interpretation and translation tasks.

[0100] In another embodiment, CICAI employs microkernels that satisfy high-assurance standards such as Common Criteria EAL-4 to EAL-7; the configuration of FIG. 7 separateskernel and user space, supports IPC through a secure publish-subscribe bus and allows agents to migrate between systolic and wavefront arrays.

[0101] A HuroBOSS boot-loader agent monitors the inventory of software components expected to run inside the microkernel; if a component is missing or stalled, the agent schedules a start, restart or shutdown and records the event in a management information base.

[0102] During boot, an init process launched by the microkernel starts essential user-space services and daemons; if init detects a kernel anomaly it halts or restarts the system to protect service integrity for downstream CICAI agents.

[0103] To enhance fault tolerance, the microkernel can maintain redundant boot-loaders and init processes, while external tools and the HuroBOSS agent supervise health, restart failed threads and issue alerts.

[0104] Subsequent sections classify Al-agent types — homogeneous reflexive, heterogeneous reflexive, model-based reflexive and others — each optimised for a particular runtime adapter or array topology; details follow in paragraphs

[0074] and beyond.

[0105] Some embodiments instantiate Trusted Platform Module, TPM, Al agents to guard credentials, attest firmware integrity and seed cryptographic operations directly inside the CICAI chassis; two supported variants are:

[0106] (a) a discrete TPM chipset soldered to the class-4 switch motherboard that exposes FIPS 140-2 compliant random-number, key -generation and sealed-storage primitives; and

[0107] (b) a firmware TPM, for example Intel Platform Trust Technology or AMD fTPM, loaded at boot time so that even virtual switch instances running on cloud bare metal obtain measured-boot assurance without extra hardware.

[0108] The following paragraphs provide additional definitions used throughout this disclosure, thereby harmonising claim language with the description and ensuring that every term has clear antecedent basis.

[0109] In one embodiment an Al agent is a hardware or software entity whose own memory and code execute supervised, unsupervised or reinforcement-learning instructions on SISD, MISD, SIMD or MIMD hardware, and may interface with external systems through well-defined application-programming interfaces to perform tasks such as speech recognition, text generation or fraud analytics within the CICAI network.

[0110] Reinforcement learning, in the policy approach, allows a self-determinant Al agent to select its next action from a set of axioms without external commands, so a class-4 switch can adapt routing or codec selection in the field while maintaining carrier-grade availability.[OHl] Other embodiments employ value-seeking reinforcement learning in which the agent evaluates long-horizon utility — such as aggregate call-quality score over a ten-minute interval — rather than an immediate reward, thereby smoothing transient impairments and improving user experience during bursty traffic.

[0112] An Al agent can monitor real-time telemetry from a runtime adapter, compare the data with policy scripts and autonomously tune resources; for example, when packet-loss spikes on an edge link the agent shifts transcoding to a less congested path and annotates the incident in a secure log for post-event audit.

[0113] Behaviour-based classification of Al agents in the CICAI ecosystem is organised as follows:

[0114] (a) reactive;

[0115] (b) proactive;

[0116] (c) passive;

[0117] (d) fixed-environment;

[0118] (e) dynamic-environment;

[0119] (f) single-agent;

[0120] (g) plural-agent; and

[0121] (h) multi-agent systems.

[0122] A reactive Al agent responds to immediate sensor input — for example jitter telemetry from a media stream — by issuing corrective actions within microseconds, and its tight control loop is coded in a memory-safe language to prevent timing violations in the switch fabric.

[0123] A proactive Al agent stores goals, plans and environmental models so that, when operated in a dynamic environment such as a Low-Earth-Orbit backhaul, it can re-plan modulation schemes or route tables before a hand-over event degrades service.

[0124] Multi-agent systems coordinate several hardware- or software-based agents that share objectives — for example, a robotics fleet or a distributed intrusion-detection mesh — through concise grammars and context-aware ontologies, thereby achieving tasks like simultaneous language conversion and threat monitoring without human intervention.

[0125] In some embodiments an Al agent can take several concrete forms:

[0126] (a) a hardware or software entity that owns private or shared memory address space;

[0127] (b) a discrete processing core or other device that itself acts as the agent;

[0128] (c) a software construct that runs multiple threads on a processing unit; and

[0129] (d) a communication interface with its own code library so the agent can reprogram itself through dynamic planning algorithms.

[0130] An Al agent control unit can refresh its axiom repository by instructing the linked runtime adapter to supply new axioms, and the resulting knowledge abstraction may draw on a small language model, a large language model, or an encrypted vectorised hash table, chosen according to the inference power required.

[0131] The control unit may host a curated set of runtime adapters dedicated to tasks such as language-model ingestion, unstructured data processing, machine learning, deep learning, natural language processing, translation, or text to speech; it monitors adaptertelemetry, crafts new axioms in its native grammar, and orchestrates swarms of agents to reach stated goals without direct human command.

[0132] A dedicated axiom repository can reside in an encrypted hierarchical blockchain wallet so authorised agents can retrieve instructions quickly while relying on trusted platform agents to verify every entry; the same repository may hold both pre-written algorithms and instructions generated autonomously by the agents during operation.

[0133] The runtime adapter itself may be realised as a processor core, a field programmable gate array, an application specific integrated circuit, a system on chip, a virtual machine, or a software emulator, and it may operate under MISD, SISD, MIMD, SIMD, or SWAR arrangements so that agents gain access to suitable compute resources even when specialised hardware is unavailable.

[0134] A generalised large language model is a neural network trained on broad textual collections so that, when prompted, it predicts the next token and generates fluent language for tasks such as classification or open-ended text creation.

[0135] A personalised large language model captures the unique dialect, vocabulary, tone, and grammatical habits of an individual by learning from their own documents, transcripts, and related artefacts, thereby enabling context-aware responses that evolve as new material appears.

[0136] A personal Al agent can overlay a personalised language model on top of general models so that translation, interpretation, or question answering occurs in the user’s style, taking into account factors such as religious hermeneutics, dialect, profession, first language, and educational background, while the language abstraction layer queries several models in different languages and returns one coherent reply.

[0137] Real time experiments show that adding diacritics to transcripts preserves tone, timbre, intention, and emotion, which allows automated analysis to detect coded speech and to recruit additional agents for deeper scrutiny when needed.

[0138] A personal profile may store linguistic and cultural attributes such as regional pronunciation, grammar, and vocabulary variations so the system can tailor communication to the individual while recognising differences from a standard literary form.

[0139] Networking hardware and co-located software placed inside or next to a telecommunications switch convert speech from any caller into language, syntax and register that the listener understands in real time, defined here as no more than two hundred milliseconds of end-to-end delay; by erasing linguistic and educational barriers, the platform lets capital in large and small economies reach qualified labour simply by dialing a telephone number, thereby opening global commerce to participants who previously lacked access.

[0140] In one embodiment, digital and analogue carrier networks host data processing units that couple to external computing platforms so that interpretation, translation, transcription and transliteration occur during unilateral, bilateral or omnidirectional calls while keeping round-trip latency below two hundred fifty milliseconds, thus preserving normal conversational cadence for every sender and receiver combination.

[0141] The general characteristics of Al agents implemented in the system are as follows:

[0142] (a) Each agent assumes that its hosting runtime adapter already provides a cyber-secure execution environment.

[0143] (b) Agents share common services, including logging, security management, device adaptation and class loading, supplied by the runtime adapter.

[0144] (c) Agents participate in a limited, stochastic hierarchy that is defined by available algorithms, axioms and topology maps so that collaborative problem-solving remains coordinated.

[0145] (d) Every agent is a self-reflexive software clone that adapts state at runtime without altering baseline code.

[0146] (e) Each agent receives a universally unique identifier that, together with the adapter identifier, forms a unique agent name.

[0147] (f) The agent records its management status — controlling, subordinate or not applicable — and checkpoints its state so it can serialise to a new location if local resources become constrained.

[0148] (g) Named thread groups, which may execute in kernel or user space, isolate agent execution and apply operating-system controls.

[0149] (h) Shared variables support remote procedure calls, inter-process communication, queue management and voting among agents.

[0150] (i) The runtime adapter assigns each agent a priority that reflects its class and governing axioms.

[0151] (j) Upon instantiation the agent registers its identifier with the local adapter.

[0152] (k) Kernel threads can passivate an agent, preserving state, and later reactivate it without loss of service.

[0153] (1) When a request falls outside its jurisdiction, an agent may ignore, defer or reroute the axiom decision to a suitable peer.

[0154] Runtime adapters expose per-thread utilisation telemetry that a control unit uses to rebalance load across adapters, preventing resource exhaustion and maintaining deterministic performance.

[0155] Systolic and wavefront arrays of parallel agents accelerate tasks such as image analysis, speech recognition, neural -network inference and symmetric-key encryption, allowing the switch to process multimedia streams at line rate without sacrificing interpretation accuracy.

[0156] Large-language-model training data may be drawn, with permission, from public or non-proprietary documents produced by a cluster of related enterprises so that domain-specific lexicons — for example in construction — retain contextual meaning during automated translation.

[0157] A language-model abstraction layer reformats content requests to and from external Al platforms, for instance Google Gemini or Microsoft Copilot, and rewrites the returned data in phrasing and terminology appropriate to the target purpose before adding it to the training corpus.

[0158] Positioned at link-layer locations inside carrier networks, the CICAI switch delivers interpretation, translation, transcription and transliteration in real time for peer-to-peer, one-to-many, many-to-one and many-to-many traffic patterns without requiring additional over-the-top services.

[0159] Digital and analogue networks therefore provide computation-assisted source-to-target language conversion for any mix of senders and receivers, consolidating legacy equipment into a smaller and greener footprint while solving long-standing technical barriers to global communication.

[0160] In one embodiment a softswitch is a call-switching node implemented as software that runs on general-purpose computing hardware rather than on specialised exchange equipment. The softswitch connects telephone calls among callers, subscribers or other endpoints across a telecommunications network and, where appropriate, switches voice traffic using Voice over IP techniques; it may reside inside a class 4 telecommunications switch or on companion computing platforms that interface with that switch.

[0161] The following use case demonstrates a multilingual softswitch incorporating CICAI services and an integrated call-audio recorder, enabling compliance capture and machine-learning refinement of stored media.

[0162] In this configuration the CICAI infrastructure positions the call-audio recorder between a PBX or media gateway and a corporate local or wide area network, storing conversations so that deep-learning routines can later update personalised or anonymised language models derived from captured speech.

[0163] Another embodiment combines a class 4 software-defined networking switch with a class 1 international gateway that serves as both load balancer and quality-assurance platform; the class 4 softswitch does not terminate subscriber handsets directly, but instead interconnects peer class 4 switches and class 5 access switches so that rural users on legacy second-generation and third-generation radio networks still reach any destination through one or more intermediate class 4 routes.

[0164] A CICAI softswitch equipped with autonomous Al agents routes calls over Internet Protocol networks by processing digital signals, allowing carriers to deliver advanced services on commodity infrastructure while avoiding the cost of proprietary exchange hardware.

[0165] Embodiments embed CICAI functions inside existing telecommunication switches so that interpretation, transcription, translation and transliteration occur in real time; a multi-agent network of processing units executes computer-readable instructions within the switch, supporting language services for tens of thousands of concurrent users with nanosecond-scale internal latency.

[0166] Although the examples above focus on class 4 switches, CICAI components can likewise integrate with other switch classes or with stand-alone softswitches connected to fibre-optic or comparable networks.

[0167] Treating Al as a networking function rather than a purely computational workload yields sub-millisecond access times when CICAI executes inside a class 4 switch, keeping conversational linguistics tasks within natural speech pauses even under heavy traffic.

[0168] Because class 4 switches form the backbone of many regional, national and global networks, embedding CICAI into these nodes delivers language services to the public at scale; in one embodiment a two-hundred-gigabit backbone links a constellation of CICAI-enabled class 4 switches that perform language conversion inline while forwarding traffic toward larger carrier cores.

[0169] The system allows participants who speak different languages on opposite sides of a network, for example Mandarin and English, to converse seamlessly in one-to-one, one-to-many, many-to-one or many-to-many arrangements across fibre, copper or satellite links.

[0170] In some embodiments, CICAI class 4 switches form an integral part of the telephone network, and once a subscriber has stored profile preferences the switch remembers to provide real-time or near real-time interpretation, transcription, translation and transliteration on every outbound call, enabling that subscriber to converse naturally with friends in a foreign country while the platform handles language conversion automatically.

[0171] A service based architecture of CICAI registers each language or analytics micro-service with a discovery layer so that, when a user requests spoken or written translation, the appropriate service announces its endpoint and binds dynamically to peer services without the need for new interfaces, thereby shortening time to market for future capabilities.

[0172] In one use case, CICAI functions as a class 4 softswitch on a Voice over IP backbone, connecting large volumes of long-distance calls, providing intelligent routing, and inserting language interpretation, translation, transcription and transliteration directly at the junction where packet-switched and circuit-switched domains meet, which reduces latency, congestion and operating cost for carriers.

[0173] Public switched telephone network and private-branch-exchange infrastructure that has evolved over five decades is now being replaced by packet networks; CICAI supports that transition by embedding its language services in class 4 switches so operators can retire copper plant and still deliver high-quality voice even on legacy second-generation or third-generation radio links and on long-haul fibre that operates at eight-hundred-gigabit per second or more.

[0174] As illustrated in FIG. 17, CICAI can be deployed as a class 4 tandem switch in three representative locations of a telecommunications network: just inside the customer-premises demarcation point as a multilingual media gateway, inside the Internet coreas a software-as-a-service cloud application, or within a time-division-multiplex public switched network as a tandem softswitch, and similar functionality may also reside in subscriber-loop carriers that serve neighbourhoods.

[0175] To meet stringent latency goals, a linguistic service-provisioning subsystem runs on field-programmable gate arrays, application-specific integrated circuits, or systems on chip, each executing semiological and hermeneutical algorithms that personalise translations through a small language model filter applied to large language models while embedding diacritics that preserve tone, timbre and emotion for later synthesis and compliance replay.

[0176] In the example of FIG. 17, a call-centre employee working from home logs into a cloud contact-centre over a virtual private network and uses CICAI to fulfil live audio and text interactions; that centre may support languages including Chinese, Korean, Dutch, Turkish, Swedish, Indonesian, Filipino, Japanese, Ukrainian, Greek, Czech, Finnish, Romanian, Russian, Danish, Bulgarian, Malay, Slovak, Croatian, Classical Arabic, Tamil, English, Polish, German, Spanish, French, Italian, Hindi and Portuguese.

[0177] When a public switched telephone network call arrives at a private branch exchange and is routed to an interactive voice response unit, caller identification and menu choices are collected, and if the caller opts to speak with an agent the IVR passes those details to an automated call-distribution subsystem that prepares to pair the caller with a suitably qualified employee.

[0178] After IVR processing, the automated call distributor selects a routing strategy from the following options:

[0179] (a) Simultaneous call distribution, which alerts all available agents at once and assigns the call to the first responder, thereby minimising waiting time.

[0180] (b) Advanced skills based routing, which ranks agents by quantified metrics such as language proficiency, efficiency and response speed, then assigns the highest-scoring candidate.

[0181] (c) Business-hours routing, which connects the caller to agents currently on shift or diverts the call to voicemail when no agent is available.

[0182] (d) Hunting in groups of three, which polls three extensions for a set interval and, if unanswered, iterates through the next trio until an agent accepts.

[0183] (e) Prioritised ringing, which dials agents in a predefined order so that senior or specialised staff receive first opportunity to answer.

[0184] When an external caller finally connects to a call-centre employee, the IVR attaches contextual data to the call for agent visibility:

[0185] (a) the caller’s identifier, for example an account number;

[0186] (b) the last action completed in the IVR, such as a successful balance enquiry;

[0187] (c) a verification flag confirming that the caller provided required security credentials; and

[0188] (d) any additional parameters, including the caller’s selected language, that inform subsequent handling or translation choices.

[0189] When the agent receives the incoming call, a screen-pop user interface displays the caller data supplied by the interactive voice response unit and highlights the preferred language; CICAI immediately starts bidirectional interpretation and transcription so the agent hears and reads every utterance in the agent’s native language while the caller receives real time responses in the caller’s own language.

[0190] A session controller within CICAI inserts translation, interpretation, transcription and transliteration services inline with the media stream and logs a time-stamped record of every spoken phrase, the corresponding translated text, and any sentiment score generated by the analytics subsystem, thereby supporting later dispute resolution and model refinement.

[0191] To satisfy compliance requirements, the call-audio recorder writes two synchronised tracks — one in the original language and one in the translated language — alongwith a text transcript into long-term, write-once storage, ensuring evidentiary integrity for regulatory audits or legal proceedings.

[0192] The recorder also tags each media file with the following metadata:

[0193] (a) a globally unique session identifier;

[0194] (b) caller and agent identifiers drawn from secure directories;

[0195] (c) start-time and end-time stamps accurate to at least one millisecond;

[0196] (d) codec information for each leg of the call; and

[0197] (e) the language pair used for translation.

[0198] After the call ends, an offline analytics pipeline reviews speech, text and sentiment data, updates personalised or anonymised language models, and feeds improved parameters back into edge nodes so future sessions benefit from higher translation fidelity and lower latency.

[0199] Throughout the process, zero-trust security policies enforce mutual-TLS authentication among CICAI components, and an inline firewall filters malformed packets to prevent code-injection attacks that could corrupt live language services.

[0200] If network congestion or compute saturation threatens to push end-to-end delay above two hundred milliseconds, the orchestrator can invoke a graceful -degradation plan that switches from neural translation to cached phrase books or, as a last resort, forwards the call to a human interpreter while preserving the audit trail.

[0201] The quality-assurance module measures conversation-grade, word error rate, latency and user feedback to verify that every translated statement meets carrier-class standards for accuracy and timeliness before releasing aggregate metrics to the operations centre.

[0202] Key performance indicators tracked by the operations centre include:

[0203] (a) average round-trip latency per language pair;

[0204] (b) ninety -fifth-percentile word error rate;

[0205] (c) mean opinion score collected from post-call surveys;

[0206] (d) percentage of sessions requiring human fallback; and

[0207] (e) successful compliance-archive writes versus attempts.

[0208] Periodic regression tests replay archived calls through updated models to confirm that no material translation errors are introduced during software upgrades, and the results feed a continuous-integration pipeline that blocks deployment if any metric falls outside predefined acceptance ranges.

[0209] Persistent virtual-desktop infrastructure offers several key advantages:

[0210] (a) Customisation permits each call-center employee to store personal settings such as passwords, shortcuts or screen savers, giving every session the look and feel of a dedicated workstation.

[0211] (b) Usability improves because employees return to the same familiar environment at every log-on, reducing orientation time and training cost.

[0212] (c) Simple desktop management lets cloud or datacenter administrators maintain virtual machines with the same tools used for physical desktops, avoiding complex image re-engineering.

[0213] As already noted, CICAI functions as a multi-agent communications network in which one or more processing units execute computer-readable instructions that support interpretation, transcription, translation and transliteration for computational linguistics research and commercial workloads.

[0214] Embodiments may operate over digital or analogue carrier networks whose data-processing units couple to external platforms so that source-to-target language conversion occurs in real time or near real time for unilateral, bilateral and omnidirectional traffic patterns, details of which appear later with reference to FIG. 18 and FIG. 19.

[0215] FIG. 18 illustrates a call-center use case in which a CICAI route-switch processor sits between a PBX or media gateway and the corporate LAN or WAN, thereby insertingAl-driven language services at the customer-premises entry point without altering existing call-handling procedures.

[0216] Although FIG. 18 depicts the audio-record server as a separate element, other embodiments integrate that server into the same CICAI route-switch processor so the system records, stores, archives and retrieves call audio alongside language services, as further explained by the message flow in FIG. 19.

[0217] In general, an incoming call reaches the customer-premises demarcation point, commonly the location where a T1 or T3 cable enters the building; from there it passes to the PBX or media gateway, which assumes responsibility for telephone wiring and equipment on the subscriber side.

[0218] Once across that demarcation, the PBX prioritises the call and routes it, through a CTI or ACD server if present, to a call-center employee’s desktop system where the call may enter a queue; customer-relationship-management content and audio recording begin as needed to support personalisation and compliance.

[0219] The same PBX or gateway also bridges traditional T1 or T3 public-switched telephone network signalling with TCP IP traffic inside the premises, while the adjacent CICAI route-switch processor provides intelligent routing, interpretation, translation and transcription, together with integrated audio capture for historical records.

[0220] Deployment options include locating the CICAI system outside the customer LAN or WAN and communicating with PBX, CTI or audio-record servers over TCP IP, or installing the CICAI platform directly on the internal network so language services run alongside enterprise applications.

[0221] FIG. 19 presents an example message-flow diagram with operations numbered 1901 through 1922; embodiments may implement all, a subset or an expanded set of those messages, and individual steps may occur in the illustrated order or concurrently without departing from the claimed subject matter.

[0222] FIG. 19 illustrates the call setup message flow, and the following description list maps each numbered message 1901 through 1922 and each reference box 1951 through 1955 to its corresponding operation:

[0223] (a) Message 1901 : Request call to call center employee

[0224] (b) Box 1951 : Generate unique Session Initiation Protocol Uniform Resource Identifier for the customer call

[0225] (c) Message 1902: Authorize calls for the customer Uniform Resource Identifier under RFC 3966

[0226] (d) Box 1952: Generate a unique authorization token for the session

[0227] (e) Messages 1903 and 1904: Return the authorization token together with Internet Protocol and port assignments

[0228] (f) Box 1953: Initialize a low latency protocol client, for example Real Time Protocol, Secure Real Time Transport Protocol, Web Real Time Communication, Joint Range Extension Application Protocol Appendix C, or Websockets

[0229] (g) Message 1905: Retrieve the estimated wait time in queue and insert the CustomerlD into that queue

[0230] (h) Message 1906: Return the estimated wait time to the customer device

[0231] (i) Message 1907: Inform the customer of the estimated wait time, then place the call on hold

[0232] (j) Message 1908: Queue the customer call with CustomerlD, QueuelD and Uniform Resource Identifier values

[0233] (k) Box 1954: Hold the call until it reaches the top of the queue and alert the next available call center employee

[0234] (1) Message 1909: Transfer the call to the call center employee using CustomerlD and Uniform Resource Identifier

[0235] (m) Message 1910: Instruct the customer device to call the employee Session Initiation Protocol Uniform Resource Identifier

[0236] (n) Message 1911 : Screen-pop interface alerts the employee with Language ID, Customer ID and Customer Uniform Resource Identifier

[0237] (o) Message 1912: Employee desktop software signals the call audio recorder to start recording audio

[0238] (p) Message 1913: Employee desktop software requests customer profile data from the customer relationship management system

[0239] (q) Message 1913b: Employee desktop software signals CICAI to begin interpretation, transcription, translation or transliteration

[0240] (r) Message 1914: Employee desktop software returns customer profile data

[0241] (s) Message 1915: Real Time Protocol media flows

[0242] (t) Box 1955: Time division multiplexing or Internet Protocol PBX or Session Initiation Protocol with the media session established

[0243] (u) Message 1916: Secure Real Time Transport Protocol with Datagram Transport Layer Security is active

[0244] (v) Message 1917: Customer asks a question and the employee enlists an Al agent to query a large language model

[0245] (w) Message 1918: The Al agent submits the prompt, returns the response and the customer hangs up

[0246] (x) Message 1919: Employee desktop application signals CICAI to stop interpretation, transcription, translation or transliteration

[0247] (y) Message 1920: Employee desktop application signals CICAI to stop recording and to archive the audio

[0248] (z) Message 1921 : Real Time Protocol session ends and the application signals the media gateway to terminate the session

[0249] (aa) Message 1922: Secure Real Time Transport Protocol and Datagram Transport Layer Security session terminates and the Uniform Resource Identifier is released

[0250] In step 1902 the route switch processor authenticates the INVITE using mutual Transport Layer Security, validates the caller profile held in the subscriber database and reserves compute resources for interpretation, transcription and translation.

[0251] Subsequent signaling exchanges provide provisional responses, ring-back generation and final acknowledgements that together complete call setup within one hundred milliseconds, maintaining user-perceived delay inside carrier targets.

[0252] After the signaling handshake finishes, Real Time Transport Protocol media packets flow; the CICAI processor forks those packets simultaneously to the language-service pipeline and the compliance recorder without adding jitter or causing packet loss.

[0253] The language pipeline runs voice-activity detection, automatic-speech recognition, language identification and sentiment analysis in parallel threads so that each spoken phrase emerges as time stamped text ready for translation.

[0254] Because configuration variants exist, this disclosure lists three optional call-flow adjustments:

[0255] (a) direct media between PBX and agent desktop when translation is disabled, leaving only signaling through CICAI;

[0256] (b) split media where the caller leg passes through CICAI for translation while the agent leg remains direct; and

[0257] (c) fully mediated media in which both legs traverse CICAI, allowing consistent quality-of-service enforcement on ingress and egress.

[0258] During conversation a resource monitor tracks central-processing-unit, graphics-processing-unit and network utilisation; if any metric exceeds eighty percent for more than five seconds the orchestrator instantiates extra runtime adapters across available nodes.

[0259] The compliance recorder embeds cryptographic hashes every thirty seconds so later audits can verify file integrity, and hash checkpoints conform to the Privacy Enhancing Integrity Format described elsewhere in this document.

[0260] When either participant terminates the call, the PBX sends a BYE request that the route switch processor acknowledges and forwards; media flow stops, the recorder finalises storage, writes metadata and notifies the analytics pipeline.

[0261] The analytics pipeline then appends session metrics to the central dashboard, triggers model fine-tuning where needed and archives language data in accordance with retention policies that depend on jurisdiction, market vertical and regulatory requirements.

[0262] In one embodiment a permissioned blockchain stores user-profile metrics — education, linguistic background and professional credentials — harvested with consent from sites such as Linkedln; each block is encrypted so that only the rightful user and authorised CICAI agents can decrypt or append attributes, thereby ensuring tamper-evident provenance while allowing the switch to personalise language services instantly.

[0263] A hierarchical-wallet token architecture tracks digital-rights information for both virtual goods, for example newsletters and news feeds, and physical items such as hats and shirts; the marketplace that emerges helps users locate collaborators or products inside the CICAI ecosystem when conventional social-network searches have failed.

[0264] During call setup the system receives a unique user identifier, and the spawned Al agent presents a fractional-ownership proof of the corresponding token; dual two-factor authentication — agent-to-infrastructure and user-to-agent — completes in under one hundred milliseconds even when cipher suites differ across subsystems.

[0265] A fleet of CICAI class-4 switches distributed worldwide replicates profile and token ledgers in near real time so any participating node can route traffic and deliver language services with local point-of-presence latency, effectively letting “every machine knoweverything about everybody” from a routing standpoint while respecting the encryption boundaries of each user’s data.

[0266] Each Al agent is provisioned with its predefined token at spawn time and carries that credential wherever it migrates, assuring continuous policy enforcement throughout the agent life cycle.

[0267] Architecturally a CICAI class-4 switch operates as a tandem node that assists class-5 and class-3 exchanges; the same hardware and software stack can attach, according to deployment preference, to customer-facing packet networks, non-customer-facing time-division links or internal IP cores.

[0268] To achieve non-stop service the switch employs watchdog timers, hot-swappable modules and self-healing control-plane routines that isolate or restart faulty elements without dropping active sessions.

[0269] The foregoing embodiments may be realised, individually or together, within the devices, systems and infrastructures shown in FIG. 20, FIG. 21 and FIG. 22, or in equivalent configurations.

[0270] The World Wide Web continues to expand through daily additions of text files, images, audio files, video files, web pages and even machine-generated measurements of physical phenomena; increasingly these stored signals are created and transmitted by Intemet-of-Things devices identifiable by unique IP addresses — examples include automobile sensors, biochip transponders, heart-monitor implants, thermostats, kitchen appliances, solar-panel arrays, home gateways and smart locks — that embed computing resources to acquire, process and send content across heterogeneous networks.

[0271] Improving communication among such loT devices involves processing the queries they generate, optimising routing and ensuring that each request reaches an appropriate CICAI service with deterministic latency despite constrained bandwidth.

[0272] A query generator running inside each loT node forms outbound requests by combining a unique device identifier, an optional location tag and one of three query types:

[0273] (a) immediate push of urgent telemetry such as cardiac arrhythmia data;

[0274] (b) periodic status updates that include temperature, humidity or battery level; and

[0275] (c) on-demand replies to CICAI agents that poll the node for configuration details.

[0276] A distributed indexing service stores every incoming request in a sharded key-value database, appends a monotonic timestamp with microsecond precision and records the cryptographic hash of the payload so later validation can prove chain-of-custody integrity.

[0277] When a CICAI edge cluster receives a high-priority query it checks a local cache before forwarding the request upstream, which reduces backhaul bandwidth and lowers median round-trip time below fifty milliseconds for frequently accessed sensor data.

[0278] Transport-layer security relies on Datagram Transport Layer Security version one-three and uses cipher suites that meet or exceed National Institute of Standards and Technology level one approvals; perfect forward secrecy is maintained through ephemeral elliptic-curve key exchange.

[0279] The system supports three adaptive compression modes selected by a link-quality estimator:

[0280] (a) no compression when packet-loss is below one percent and bandwidth is ample;

[0281] (b) lightweight run-length encoding when bandwidth tightens yet central-processing-unit load remains low; and

[0282] (c) context-adaptive binary arithmetic coding when both bandwidth and power are constrained, for example during satellite passes.

[0283] At the far edge of the network a miniature inference engine executes a distilled transformer model that classifies sensor events; if the confidence score exceeds eighty-five percent the node sends only a label instead of raw data, saving energy and radio airtime.

[0284] An anomaly-detection agent inside the CICAI switch inspects inbound sensor labels and raises an alert when observed frequency diverges from a rolling twenty-four-hour baseline by more than three standard deviations, thereby flagging potential device failure or malicious activity.

[0285] To extend battery life, sensor firmware schedules deep-sleep intervals and keeps the radio off except when a wake-word interrupt or a scheduled transmission window occurs; typical duty-cycle is under two percent, achieving multi-year operation on a coin cell.

[0286] Each loT endpoint installs new firmware through the following secure update sequence:

[0287] (a) publish a signed manifest that lists version, size and hash of the new image;

[0288] (b) verify the manifest signature against the vendor certificate stored in the Trusted Platform Module;

[0289] (c) download the image in encrypted chunks, each authenticated by chunk-level hashes;

[0290] (d) write the update to an alternate partition while the current firmware keeps running;

[0291] (e) perform a pointer swap and reboot into the new image; and

[0292] (f) roll back automatically if a health-check heartbeat fails within sixty seconds.

[0293] FIG. 23 depicts a representative topology in which thousands of loT endpoints forward compressed, classified messages through regional CICAI edge clusters to a central analytics cloud, while a parallel compliance path stores encrypted telemetry and translation logs in an immutable archive for regulatory review.

[0294] A “network” encompasses any arrangement of interconnected links and nodes — such as segments of the Internet, local-area networks (LANs), wide-area networks (WANs), wirelinetrunks, wireless links, or any combination thereof — that transports digital signals. The term expressly includes present or future derivatives and improvements, for example cloud object stores, network-attached- storage (NAS) arrays, storage-area networks (SANs), and comparable device-readable media deployed at the core, edge, or user premises.

[0295] A “sub-network” is a logically distinct portion of a larger network — for instance a LAN segment, an optical ring, or an ad-hoc mesh — that exchanges packets or frames among its nodes over wired or wireless links while remaining interoperable with neighbouring segments. Devices may communicate “transparently” across intermediary routers or bridges, meaning the endpoints need not identify the intervening hops.

[0296] A “private network” denotes a bounded set of nodes authorised to exchange traffic with one another without routine re-routing through public gateways. The private network can be self-contained or realised as an isolated slice (for example, a virtual private cloud) within a larger infrastructure so that external devices may originate outbound traffic but cannot initiate inbound sessions unless expressly permitted.

[0297] The term “Internet” refers to the decentralised, packet-switched aggregation of interoperable networks that comply with any current or future version of the Internet Protocol (IP); constituent elements include LANs, WANs, wireless access networks, and long-haul public backbones. The subset of the Internet that conforms to the Hypertext Transfer Protocol (HTTP) is commonly known as the World Wide Web (Web).

[0298] Although the embodiments are illustrated with reference to the Internet or the Web, the disclosed subject matter applies equally to any present or future network architecture capable of storing, retrieving, or conveying hypermedia or other digital content.

[0299] An “electronic file” or “electronic document” is a logically associated set of stored memory states or transmitted physical signals — regardless of syntax, container format, or physical storage location — that collectively represent content such as source code, text, images, audio, or video.

[0300] Markup languages, including without limitation any version of Hypertext Markup Language (HTML) or Extensible Markup Language (XML), may describe digital content in a form suitable for storage as an electronic document (for example, a Web page) or rendering by a browser, server, or other computing node.

[0301] A “Web page” is an electronic document retrievable over a network by means of a uniform resource locator (URL), and a “Web site” is an electronic association of multiple such pages; developers may embed executable code — such as JavaScript — within these pages to generate or manipulate the displayed content.

[0302] Terms such as “entry,” “document,” “electronic document,” “content,” “digital content,” and “item” refer to signals or memory states expressed in a digital format; when these signals are rendered in human-perceptible form — such as displayed text, audible speech, or haptic feedback — the user is said to consume the digital content.

[0303] An electronic document may comprise multiple components — e.g., text segments, image objects, or attribute sets — each embodied as physical signals or memory states; memory-resident representations are tangible, whereas in-flight signals become tangible only when converted, for example, into pixels on a display.

[0304] The term “parameter” denotes descriptive information that characterises a collection of signal samples — such as an electronic document or file — and that exists as physical signals or memory states. Representative parameters include: time of image capture; latitude and longitude of the capture device; author or contributor names; collection identifier; creation technique and purpose; creation timestamp; logical storage path; and any encoding format or protocol specification (for example a markup language) that renders the content substantially compliant with a designated standard.

[0305] “Signal packets” and “signal frames,” collectively referred to as transmissions, are physical signals exchanged between network nodes, each node comprising one or more computing or network devices that may share a local address space. The term transmissionimposes no directionality: a packet may flow to or from any endpoint, and either endpoint may initiate the exchange under a push type or pull type paradigm, the distinction resting solely on which end of the communication path starts the transfer.

[0306] A signal packet or frame may traverse a communication path that includes Internet or Web segments; for instance, a site can send traffic through an access node to the Internet and onward through gateways and servers until the packet reaches a target site on a local network. Routing follows the destination address and current path availability, and because not every interoperable network is publicly reachable, certain segments may be unavailable at a given time.

[0307] A “network protocol” is a set of signalling conventions that enable devices in a network to communicate, and may be described according to the layered Open Systems Interconnection (OSI) model. Expressions such as “between” encompass “among” where appropriate, and “compatible with” or “comply with” expressly includes substantial compatibility or compliance.

[0308] A network protocol stack comprises multiple layers: the physical layer specifies how symbols are conveyed as signals over a medium such as copper, fibre or a wireless air interface; higher layers add functions such as addressing, session control and permission management that become available when devices communicate using the corresponding layer of the stack.

[0309] In some embodiments a network or sub-network conveys signal packets or frames using any present or future version of one or more protocol stacks, including ARCNET, AppleTalk, ATM, Bluetooth, DECnet, Ethernet, FDDI, Frame Relay, HIPPI, IEEE 1394, IEEE 802.11, IEEE-488, Internet Protocol Suite, IPX, Myrinet, the OSI Protocol Suite, QsNet, RS-232, SPX, System Network Architecture, Token Ring, USB and X.25; additional transport protocols may include TCP / IP, UDP, DECnet, NetBEUI and AppleTalk. Versions of the Internet Protocol may include IPv4, IPv6 or successors thereof.

[0310] A wireless network couples devices to the wider network and may operate as a stand-alone, ad hoc, mesh, WLAN or cellular arrangement. It can incorporate technologies such as Long-Term Evolution (LTE), WLAN, wireless-router mesh, or second- through fifth-generation (2G-5G) cellular standards to provide wide-area coverage for client and gateway devices with varying mobility profiles.

[0311] Wireless communications within such a network may employ air interfaces or access technologies including GSM, UMTS, GPRS, EDGE, LTE, LTE-Advanced, WCDMA, Bluetooth, ultra-wideband, and IEEE 802.11b / g / n, as well as any future mechanism that permits radio-frequency or analogous signal exchange between devices, networks or sub-networks.

[0312] As illustrated in FIG. 22, one embodiment comprises a local network that includes device 1104 and medium 1140, together with another network such as computing or communications network 1108. Network 1108 may furnish links, processes, services or applications — implemented over wired or wireless connections, telephone or telecommunications systems, Wi-Fi, Wi-MAX, the Internet, a LAN or a WAN — that enable communication signals between client device 1102 and server device 1106.

[0313] Devices depicted in FIG. 22 may function as client computing devices, server computing devices or hybrids. The term “computing device” refers to a processor and a memory coupled by a communication bus; that structure is sufficient to avoid treatment under 35 U.S.C. § 112(f) . Should § 112(f) nonetheless be deemed to apply, the corresponding structure, material or acts are intended to include the elements illustrated in FIGS. 1-20 and their associated textual descriptions.

[0314] FIG. 22 shows an embodiment in which first device 1102, second device 1104 and third device 1106 each present the same graphical user interface through a browser, thin client or native application, so that handsets, wearables, satellite terminals and virtual desktops expose identical controls. Processor 1120 cooperates with memory 1122 (including primarymemory 1124 and secondary memory 1126) over bus 1115 to deliver the interface together with back-end services, allowing differing hardware to support a common user workflow.

[0315] As used here, “device” embraces any digital apparatus capable of executing or participating in networked processes; illustrative examples range from desktop and notebook computers to smart televisions, smartphones, wearables and Internet-of-Things sensors. Any described process may run entirely within a single device, be partitioned across multiple devices or migrate among them, enabling coverage from minimalist keypads with monochrome screens to full-featured machines equipped with touch-sensitive three-dimensional colour displays, inertial sensors and on-board global-positioning receivers.

[0316] A device can communicate with a wireless network through protocols such as GSM, EDGE, IEEE 802.11 variants or WiMAX, and may rely on subscriber-identity modules in removable, embedded or purely electronic form. Each device can carry one or more network addresses, including telephone numbers, Internet Protocol addresses or comparable identifiers, assigned by cellular, wired or Internet service providers, and may shift between wired, wireless or hybrid connections without departing from the disclosed teachings.

[0317] A device may host any current or future operating system, including desktop platforms such as Windows, macOS or Linux and mobile platforms such as Android, iOS or Windows Mobile, and can run applications from email clients to cloud-native micro-services. Such operating-system diversity allows the same mechanisms to function across vendor-specific kernels and to remain effective when an application migrates to a hosted environment.

[0318] In FIG. 22, first computing device 1102 can provide executable instructions stored as physical memory states and transmit those instructions or resulting data to second computing device 1104 over network 1108; the link may be wired, optical, radio or virtual. Code mobility of this kind permits on-premises installations, cloud deployments and edge-computing nodes to share the same architecture without requiring separate implementations.

[0319] Still with reference to FIG. 22, memory 1122 represents any non-transitory storage mechanism, including random-access memory, magnetic or optical disks, solid-state drives or combinations thereof, and holds the machine-readable instructions that realise the claimed functions. The range of suitable media supports deployment across legacy and modem hardware.

[0320] Processor 1120 retrieves and executes instructions from memory 1122, while a memory controller accesses device-readable medium 1140 for additional code or data, establishing a concrete execution path between stored content and outbound network signals. Executing instructions generate outbound signals for transmission and may store resulting states back into memory or another storage layer.

[0321] Memory 1122 may also store electronic files or documents logically associated with one or more users even when the physical blocks reside at different locations within the tangible memory illustrated in FIG. 22. Such logical association covers distributed-storage implementations, including sharded databases and cloud object stores, that organise data across multiple physical devices while presenting a unified namespace.

[0322] Algorithmic descriptions and symbolic representations act as concise notations for physical signal-processing operations; an algorithm is a self-consi stent sequence that manipulates electrical or magnetic signals to achieve a defined result, such as compressing video frames or executing an inference model.

[0323] Terms like “bit,” “value,” “parameter” and “number” are convenient labels for underlying physical signals or states, each representing a measurable quantity expressed in tangible form. Anchoring these labels to physical artefacts enables a practitioner to implement the described techniques without undue experimentation.

[0324] Operation of a memory device — such as memory 1122 illustrated in FIG. 22 — entails a tangible state change; for example, flipping a stored value from logical 1 to logical 0 transforms the physical article itself. Depending on technology, the transformation may storeor release electric charge, re-orient magnetic domains, induce a crystalline-to-amorphous phase shift, or invoke quantum-mechanical effects in qubits; each mechanism alters the substrate that embodies the data and therefore qualifies as a non-transitory physical transformation.

[0325] As further shown in FIG. 22, processor 1120 comprises one or more digital-logic circuits — controllers, microprocessors, microcontrollers, application-specific integrated circuits, digital-signal processors, programmable-logic devices or field-programmable gate arrays — that execute instructions fetched from memory. Execution generates, manipulates and stores electrical signals representing program states, thereby implementing portions of the disclosed computing procedures without departing from these teachings.

[0326] FIG. 22 also depicts component 1132 within device 1104; this component interfaces with input and output peripherals so that electrical signals can flow bidirectionally between device 1104 and external equipment. Representative input peripherals include mice, styli, trackballs, keyboards and speech-capture microphones, while representative output peripherals include displays and printers capable of presenting visual or audio stimuli to a user.

[0327] The foregoing description presents particular amounts, systems and configurations solely to facilitate understanding; skilled artisans will recognise that numerous modifications, substitutions and equivalents may be adopted without departing from the scope of the appended claims, which are intended to encompass all subject matter that falls within their express language and legally permissible equivalents.

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising:(a) a telecommunications switch-type infrastructure that includes at least one processing unit and non-transitory memory;(b) machine-readable instructions stored in the memory which, when executed by the at least one processing unit, instantiate at least one artificial-intelligence agent configured to perform one or more computational-linguistics operations selected from interpretation, transcription, translation and transliteration on digital signals representative of a telecommunication session; and(c) a network interface configured to couple the infrastructure between a private-branch exchange (PBX) and a network selected from a local-area network (LAN) and a wide-area network (WAN) such that the computational-linguistics operations occur in real time or near real time while the telecommunication session is in progress.

2. The apparatus of claim 1, wherein the infrastructure is a softswitch that conforms to a class-4 architecture designated as Computational Interpretation of Communication using Artificial Intelligence (CICAI).

3. The apparatus of claim 1, further comprising a call-audio recorder integrated with the infrastructure and operative to record, store and retrieve audio content of the telecommunication session.

4. The apparatus of claim 1, wherein the network interface is positioned between the PBX and the LAN or WAN such that bidirectional signalling and media traffic of the telecommunication session traverse the infrastructure.

5. The apparatus of claim 1, wherein the infrastructure is deployed within a call centre on a call-centre side of a premises demarcation point.

6. The apparatus of claim 1, wherein the at least one artificial-intelligence agent is instantiated, migrated and terminated by the infrastructure in response to events of the telecommunication session and performs the computational-linguistics operations in real time or near real time.

7. The apparatus of claim 1, wherein the at least one processing unit comprises at least one of a field-programmable gate array, an application-specific integrated circuit and a system-on-chip dedicated, at least in part, to the computational-linguistics operations.

8. The apparatus of claim 3, wherein the at least one processing unit cooperates with the call-audio recorder to archive recorded audio content in secure storage.

9. The apparatus of claim 6, wherein a first artificial-intelligence agent represents an originating caller and a second artificial-intelligence agent represents a destination caller.

10. The apparatus of claim 4, wherein the infrastructure facilitates communication between a PBX / media gateway and a corporate LAN or WAN.

11. A computer-implemented method comprising:(a) receiving, at a telecommunications switch-type infrastructure that includes at least one processing unit, digital signals representative of a telecommunication session between user devices;(b) instantiating, by execution of machine-readable instructions, at least one artificial-intelligence agent;(c) performing, by the at least one artificial-intelligence agent, one or more computational-linguistics operations selected from interpretation, transcription, translation and transliteration on the digital signals in real time or near real time; and(d) transmitting transformed digital signals toward a destination device via a network interface that couples the infrastructure between a PBX and a network selected from a LAN and a WAN.

12. The method of claim 11, wherein the telecommunications switch-type infrastructure is a softswitch that conforms to the CICAI class-4 architecture.

13. The method of claim 11, further comprising recording call-audio data in a call-audio recorder integrated with the infrastructure and archiving the call-audio data in secure storage.

14. The method of claim 11, further comprising connecting the infrastructure between the PBX and the LAN or WAN.

15. The method of claim 14, further comprising facilitating bidirectional signalling and media communication between the PBX and the LAN or WAN through the infrastructure.

16. The method of claim 11, wherein the infrastructure is deployed within a call centre on a call-centre side of a premises demarcation point.

17. The method of claim 11, wherein end-to-end latency of the telecommunication session is maintained below 250 milliseconds.

18. The method of claim 11, wherein a first artificial-intelligence agent represents an originating caller and a second artificial-intelligence agent represents a destination caller.

19. The method of claim 11, wherein the at least one processing unit comprises at least one of a field-programmable gate array, an application-specific integrated circuit and a system-on-chip dedicated, at least in part, to the computational-linguistics operations.

20. The method of claim 19, further comprising dynamically migrating at least one artificial-intelligence agent between processing units in response to a change in session load.

21. A non-transitory machine-readable medium storing instructions that, when executed by at least one processing unit within a telecommunications switch-type infrastructure, cause the at least one processing unit to perform the method of claim 11.

22. The non-transitory machine-readable medium of claim 21, wherein executing the instructions causes the infrastructure to operate as a softswitch that conforms to the CICAI class-4 architecture.

23. The non-transitory machine-readable medium of claim 21, wherein executing the instructions causes the infrastructure to record call-audio content via an integrated call-audio recorder.

24. The non-transitory machine-readable medium of claim 21, wherein executing the instructions causes the infrastructure to connect between the PBX and the LAN or WAN.

25. The non-transitory machine-readable medium of claim 24, wherein executing the instructions causes the infrastructure to position itself such that bidirectional signalling and media traffic of a telecommunication session traverse the infrastructure.

26. The non-transitory machine-readable medium of claim 24, wherein executing the instructions causes the infrastructure to facilitate communication between a PBX / media gateway and a corporate LAN or WAN.

27. The non-transitory machine-readable medium of claim 21, wherein executing the instructions maintains end-to-end latency of the telecommunication session below 250 milliseconds.

28. The non-transitory machine-readable medium of claim 21, wherein executing the instructions instantiates a first artificial-intelligence agent to represent an originating caller and a second artificial-intelligence agent to represent a destination caller.

29. The non-transitory machine-readable medium of claim 21, wherein the telecommunications switch-type infrastructure is deployed within a call centre on a call-centre side of a premises demarcation point.

30. The non-transitory machine-readable medium of claim 21, wherein executing the instructions dynamically migrates at least one artificial-intelligence agent between processing units in response to session load.

Citation Information

Patent Citations

  • Transcription and analysis of meeting recordings

    US11315569B1

  • Performing artificial intelligence sign language translation services in a video relay service environment

    US20190130176A1

  • Quality control configuration for machine interpretation sessions

    US20190279621A1

  • Artificial-intelligence powered skill management systems and methods

    US20210344800A1

  • Latency-as-a-service (LAAS) platform

    US20220086846A1

Cited By

  • Context sharing method and system based on multiple agents and medium

    CN122086972A