An AI-powered interactive voice-activated sign system
Patent Information
- Application Number
- DE202025104311
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2035-07-31
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD OF THE INVENTION
[0001] This disclosure relates to an AI-supported interactive drawing system with voice control, in particular the "Sound2Sketch" system, for voice-controlled diagram creation for disabled users, especially blind people and people with motor disabilities. The system enables visually impaired and motor-impaired users to create diagrams and flowcharts using voice commands. Visually impaired users can independently draw flowcharts, UML diagrams, etc. using voice commands and continuous interpretive feedback from the system. In case of errors, users can return to the previous diagram version using the undo and delete functions. BACKGROUND OF THE INVENTION
[0002] People with visual or motor impairments have difficulty drawing or interpreting diagrams, a necessary skill in education, assessments, and professional life. Traditional drawing tools, whether digital or physical, rely heavily on visual input and manual dexterity, making them inaccessible to many. Furthermore, many existing accessibility tools lack real-time interactivity, contextual feedback, or support for complex structures such as flowcharts and graphs. This leads to a lack of independence, reduced confidence, and exclusion from important learning or work activities, preventing users from receiving immediate guidance or correcting errors during the process.
[0003] There is an unmet need for an advanced voice-activated interactive drawing system that allows the user to draw using voice commands and provides continuous interpretive feedback so that they can navigate back in case of a correction.
[0004] In view of the foregoing discussion, the present invention provides an AI-powered, interactive, voice-guided drawing system called Sound2Sketch that addresses the above-mentioned drawbacks of existing systems by enabling voice-guided, interactive diagramming with continuous feedback, thus making visual communication truly inclusive. Summary of the invention
[0005] This disclosure concerns an AI-powered interactive, voice-activated drawing system called Sound2Sketch. This voice-activated diagramming system is specifically designed for users with visual and motor impairments. The system integrates speech recognition technology, natural language processing, artificial intelligence, and real-time auditory feedback to enable users to create complex diagrams, including flowcharts, UML diagrams, and graphs, using only voice commands. The system operates with a closed-loop feedback mechanism that continuously provides auditory descriptions of the current diagram state, allowing users to make corrections and navigate back in case of errors.
[0006] The present disclosure aims to provide an AI-powered, interactive, voice-driven signing system. The system includes: a speech recognition engine that captures user speech and converts the user's spoken voice commands into text transcripts. The speech recognition engine is connected to a user input module that allows the user to capture speech via one or more microphones; a natural language processing module connected to the speech recognition engine that receives the text transcripts from the speech recognition engine and interprets the contextual meaning of the spoken voice commands; a code generation module connected to the natural language module that includes a trained language model that converts the interpreted voice commands into structured graph code;a code parser unit that receives the structural diagram code from the code generation module and converts it into visual diagram elements; a canvas connected to the code parser unit that displays the visual diagram elements as a visual diagram; a text-to-speech feedback module that generates audible descriptions of the current state of the visual diagram; a user interface module that provides voice-activated interaction controls for diagram modification; and a version control module configured to manage the version history of diagram changes and to enable the restoration of previous diagram states; wherein the system is configured to operate in a closed feedback loop in which the user receives audible feedback about the current state of the visual diagram after each voice command.and wherein the system is configured to enable users with visual and motor disabilities to create diagrams through voice commands without the need for visual or manual input.;
[0007] One objective of the present disclosure is to provide an AI-powered, interactive, voice-activated sign system.
[0008] Another objective of the present disclosure is to provide an accessible diagramming system that enables visually impaired and motor-impaired users to independently create technical diagrams through natural language interaction, thus eliminating the barriers posed by conventional visual drawing tools.
[0009] Another objective of the present disclosure is to implement a real-time feedback system that provides comprehensive auditory descriptions of diagram states, allowing users to understand and modify their creations in a non-visual manner while maintaining complete control over the drawing process.
[0010] Another objective of the present disclosure is to provide an intelligent system that can interpret contextual voice commands and convert them into structured diagram code, providing both guided and free interaction modes to accommodate different user preferences and skill levels.
[0011] To further clarify the advantages and features of the present disclosure, the invention will be explained in more detail with reference to specific embodiments illustrated in the accompanying drawings. These drawings illustrate only typical embodiments of the invention and are therefore not to be considered as limiting its scope. The invention will be described and explained in more detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE CHARACTERS
[0012] These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings, in which like characters represent like parts throughout. Fig. 1 shows a block diagram of an AI-based interactive, voice-activated character system according to an embodiment of the present disclosure. Fig. 2 is a diagram showing the working framework of the proposed system according to an embodiment of the present disclosure.
[0013] Those skilled in the art will also appreciate that the elements in the drawings are shown for convenience and are not necessarily to scale. For example, the flowcharts illustrate the method by key steps to enhance understanding of aspects of the present disclosure. Furthermore, with respect to device construction, one or more components of the device may be represented in the drawings by conventional symbols. The drawings may show only the specific details relevant to understanding embodiments of the present disclosure in order not to clutter the drawings with details that would be readily apparent to those skilled in the art from the present description. DETAILED DESCRIPTION:
[0014] To facilitate understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings and will be clearly described. However, the scope of the invention is not limited thereby. Changes and further modifications to the illustrated system, as well as further applications of the principles of the invention, are possible, as would normally occur to one skilled in the art to which the invention pertains.
[0015] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are exemplary and explanatory of the invention and are not intended to be limiting thereof.
[0016] References in this specification to "one aspect," "another aspect," or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, the language "in one embodiment," "in another embodiment," and similar language throughout this specification may or may not refer to the same embodiment.
[0017] The terms "comprises," "comprising," or other variations thereof are intended to cover non-exclusive inclusion, such that a process or method comprising a list of steps may include not only those steps, but also additional steps not expressly listed or inherent in that process or method. Likewise, the statement "comprises" for one or more devices, subsystems, elements, structures, or components does not exclude, without further limitation, the existence of other devices, subsystems, elements, structures, components, or additional devices, subsystems, elements, structures, or components.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains. The systems, methods, and examples provided herein are for illustrative purposes only and should not be considered limiting.
[0019] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0020] The functional units described in this specification are referred to as devices. A device may be implemented in programmable hardware devices such as processors, digital signal processors, central processing units, field-programmable gate arrays, programmable array logic systems, programmable logic devices, cloud processing systems, or the like. The devices may also be implemented in software for execution by various types of processors. An identified device may contain executable code and may consist, for example, of one or more physical or logical blocks of computer instructions, which may be organized, for example, as an object, procedure, function, or other construct.However, the executable file of an identified device does not have to be physically present on the device, but can consist of different instructions stored in different locations which, when logically linked, form the device and fulfill its purpose.
[0021] The executable code of a device or module can consist of one or more instructions and can even be distributed across multiple code segments, different applications, and multiple storage devices. Likewise, operational data can be identified and represented within the device and presented in any form and data structure. The operational data can be captured as a single data set or distributed across different storage devices and can be represented, at least in part, as electronic signals in a system or network.
[0022] References in this specification to "a selected embodiment," "an embodiment," or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosed subject matter. Therefore, the phrases "a selected embodiment," "in an embodiment," or "in an embodiment" in various places in this specification do not necessarily refer to the same embodiment.
[0023] Furthermore, the described features, structures, or characteristics may be combined in any manner in one or more embodiments. The following description contains numerous specific details to provide a thorough understanding of embodiments of the disclosed subject matter. However, those skilled in the art will recognize that the disclosed subject matter may be practiced without one or more of the specific details, or with different methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail in order not to obscure aspects of the disclosed subject matter.
[0024] According to the exemplary embodiments, the disclosed computer programs or modules may be executed in a variety of ways, for example, as an application in the memory of a device or as a hosted application on a server that communicates with the device application or browser using various standard protocols such as TCP / IP, HTTP, XML, SOAP, REST, JSON, and other suitable protocols. The disclosed computer programs may be written in exemplary programming languages that execute from the memory of the device or from a hosted server, such as BASIC, COBOL, C, C++, Java, Pascal, or scripting languages such as JavaScript, Python, Ruby, PHP, Perl, or other suitable programming languages.
[0025] Some of the disclosed embodiments involve or otherwise involve the transmission of data over a network, for example, the delivery of various inputs or files over the network. The network may include, for example, the Internet, wide area networks (WANs), local area networks (LANs), analog or digital wired and wireless telephone networks (e.g., PSTN, Integrated Services Digital Network (ISDN), cellular networks, and Digital Subscriber Line (xDSL)), radio, television, cable, satellite, and / or other transmission or tunneling mechanisms for transmitting data. The network may include multiple networks or subnetworks, each containing, for example, a wired or wireless data path. The network may include a circuit-switched voice network, a packet-switched data network, or another network for transmitting electronic communications.For example, the network may include Internet Protocol (IP) or Asynchronous Transfer Mode (ATM) networks that support voice, such as VoIP, Voice over ATM, or other comparable protocols for voice data communication. In one implementation, the network includes a cellular network configured for the exchange of text or SMS messages.
[0026] Examples of the network include a Personal Area Network (PAN), a Storage Area Network (SAN), a Home Area Network (HAN), a Campus Area Network (CAN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a Virtual Private Network (VPN), an Enterprise Private Network (EPN), the Internet, a Global Area Network (GAN), etc.
[0027] Fig. 1 shows a block diagram of an AI-based interactive, voice-activated character system (100) according to an embodiment of the present disclosure.
[0028] According to Fig. 1, the system (100) comprises: a user input module (102) which is triggered by holding down the space bar and starts the microphone recording when the space bar is pressed, allowing the user to issue commands and instructions verbally without manual input; a speech recognition engine (104) for recording the user's speech when the user input module (102) is triggered, which stores the recorded speech and converts the spoken voice commands into text transcripts; a natural language processing module (106) connected to the speech recognition engine (104), which receives the text transcripts from the speech recognition engine (104) and interprets the contextual meaning of the spoken voice commands; a code generation module (108) connected to the natural language processing module (106),which comprises a trained language model and converts the interpreted voice commands into structured diagram code; a code parser unit (110) that receives the structured diagram code from the code generation module (108) and converts it into visual diagram elements. A drawing surface (112) connected to the code parser unit (110) and displays the visual diagram elements as a visual diagram. A text-to-speech feedback module (114) as an S2S agent generates acoustic descriptions of the current state of the visual diagram. The S2S agent generates a description by drawing inferences from the same model that was trained in the code generation module (108) after backward learning. This generated summary is displayed to the user. A user interface module (116) that provides voice-activated and keyboard-controlled interaction controls for diagram changes so that users accept the change,reject the current step or restart the process completely. A progress control unit (118) that allows the user to navigate to the next step of the drawing process, return to the previous step, or start over using voice-based commands and keyboard shortcuts. And a version control module (120) configured to manage the version history of diagram changes and enable the restoration of previous diagram states; wherein the system (100) is configured to operate in a closed feedback loop in which the user receives audible feedback about the current state of the visual diagram after each voice command.
[0029] In one embodiment, the speech recognition engine (104) comprises: a browser-integrated speech recognition system (104a) configured to provide real-time, low-latency speech-to-text conversion without requiring backend processing.
[0030] In one embodiment, the code generation module (108) is configured to: utilize few-shot prompting with a pre-trained large language model, convert natural language transcriptions of voice commands into Mermaid.js code syntax, and generate structured diagram code based on the contextual interpretation of user voice commands.
[0031] In one embodiment, the code parser unit (110) with drawing surface (112) renders the Mermaid code into diagram blocks, wherein the code parser unit (110) acts as a link between graphical output and text representation, wherein the code parser unit (110) interprets the Mermaid code generated by the code generation unit (108) based on the syntax and semantics specified by the Mermaid.js code syntax and converts it into structured visual elements, including nodes, connectors, decision points and labels.
[0032] In one embodiment, the user interface module (116) with version control module (120) facilitates drawing modification, wherein the diagram can be updated in real time by changing the underlying code in response to further voice instructions from the user or via keyboard shortcuts.
[0033] In one embodiment, the user input module (102) and the user interface module (116) comprise: a spacebar activation mechanism (122) configured to activate the speech recognition engine when pressed and deactivate it when released; keyboard shortcut controls configured to enable navigation and modification of diagrams; and voice-activated commands configured to enable diagram editing operations.
[0034] In one embodiment, the text-to-speech feedback module (114) is configured as an S2S agent to describe the current state of the visual diagram, including the number of nodes, node names, flow directions, decision points, and component interactions, by interpreting the current diagram code using the few-shot-large language model previously used after reverse learning; and to provide real-time audible confirmation of each diagram change; and to enable users to understand diagram content in a non-visual manner, wherein the user is allowed to modify the diagram.
[0035] In one embodiment, the version control module (120) is configured to: automatically save diagram states after each change, maintain the complete version history of all diagram changes, enable restoration of previous diagram states through voice commands or keyboard shortcuts, and prevent data loss in case of connection problems.
[0036] In one embodiment, the system is configured to operate in multiple modes, including: a guided mode in which the S2S agent prompts the user at each step with available diagram options; and a free mode in which the user has the freedom to provide voice commands without guided prompts, wherein in the guided mode, the text-to-speech feedback module (114) is configured as the S2S agent to prompt the user to select from available diagram types, including flowcharts, graphs, and UML diagrams; the agent is configured to prompt the user at each step with available diagram elements; and the system is configured to create diagrams through an interactive voice dialogue between the user and the S2S agent.
[0037] In one embodiment, the system (100) further comprises: error correction features configured to allow users to change certain diagram elements without having to recreate the entire diagram; step-by-step interaction controls configured to mimic natural sketching and erasing actions; and a backward learning feature configured to allow the system to interpret and describe complete diagrams in the current context.
[0038] In one embodiment, the user interface module (114) is connected to a progress control section (114a) consisting of voice-based cues such as "Next", "Previous", and "Restart" triggered via the same spacebar activation mechanism, along with keyboard shortcuts such as "←" for "Previous", "→" for "Next", and "r" for "Restart" configured to enable navigation and modification of the diagram; and voice-activated commands configured to enable diagram editing operations.
[0039] In one embodiment, the user input module (102), the speech recognition engine (104), the natural language processing module (106), the code generation module (108), the code parser unit (110), the drawing surface (112), the text-to-speech feedback module (114), the user interface module (116), the progress control unit (118), and the version control module (120) may be implemented in programmable hardware devices such as processors, digital signal processors, central processing units, field-programmable gate arrays, programmable array logic, programmable logic devices, cloud processing systems, or the like.
[0040] The present invention addresses the significant challenges faced by people with visual and motor impairments when creating technical diagrams and flowcharts. Conventional drawing tools rely heavily on visual feedback and manual editing, making them inaccessible to users who cannot see or physically manipulate the drawing surfaces. The present invention overcomes these limitations through a comprehensive voice-controlled system that converts spoken commands into visual diagrams while providing continuous auditory feedback on the current state of the diagram.
[0041] The system architecture consists of several interconnected components that work seamlessly together to enable an accessible drawing experience. The speech recognition engine forms the foundation of user interaction. It captures spoken commands and converts them into text transcriptions using the browser's built-in speech recognition technology. This approach ensures real-time processing with minimal latency, which is crucial for a natural conversational flow during diagramming. The natural language processing module interprets the contextual meaning of spoken commands and understands user intent, even when commands are expressed in different natural language forms.
[0042] The code generation module represents the intelligent core of the system. It uses advanced language models trained using few-shot prompting techniques to convert interpreted language commands into structured diagram code. This module specifically generates the Mermaid.js code syntax, which provides a standardized format for representing various diagram types. Choosing few-shot prompting over fine-tuning or external API approaches ensures an optimal balance between accuracy, scalability, and system independence while preserving data privacy and reducing computational overhead.
[0043] The code parser serves as a bridge between text representation and visual output. It interprets the generated Mermaid.js code and converts it into structured visual elements. This component performs the complex task of translating abstract code into concrete visual components such as nodes, connectors, decision points, and labels. The canvas receives these visual elements and renders them as complete diagrams, providing real-time updates as users continue to enter voice commands.
[0044] The S2S agent provides important accessibility features by generating comprehensive auditory descriptions of the current diagram state. This component describes various aspects of the diagram, including the number of nodes, their names, flow directions, decision points, and interactions between components. This auditory feedback creates a mental model of the visual diagram for users who cannot see the actual visual output, allowing them to effectively understand and adjust their creations. This step is extremely important because it increases awareness of the representation and keeps the person on track, rather than losing track and distorting the entire representation.
[0045] The system operates in two different modes to accommodate different user preferences and experience levels. Guided mode provides structured interaction, where the system presents available options to users at each step, effectively replicating the drag-and-drop experience through "sound-and-drop" interactions. In this mode, the system guides the user through diagramming by suggesting available elements and asking for specific choices. This makes the process intuitive even for users with no diagramming experience. Free mode gives experienced users complete freedom to express their intentions naturally without guided prompts. This enables more efficient diagramming once users become familiar with the system.
[0046] The version control module ensures data integrity and provides important editing capabilities by storing a complete version history of all diagram changes. This component automatically saves diagram states after each change, preventing data loss due to connectivity issues or system interruptions. Users can revert to previous states via voice commands or keyboard shortcuts, making incremental corrections, similar to the natural process of sketching and erasing on paper.
[0047] The user interface module offers various interaction options to ensure accessibility and usability. Progress checks through voice commands such as "Next," "Previous," "Restart," and keyboard shortcuts provide additional control options for users with keyboard skills, while voice-activated commands enable completely hands-free operation for users with motor impairments.
[0048] Fig. 2 is a diagram showing the working framework of the proposed system according to an embodiment of the present disclosure.
[0049] Fig. Figure 2 shows the AI-powered, speech-guided drawing system. Users, especially visually impaired users, can enter voice commands that are interpreted contextually. These interpreted commands are then converted into code by trained ML models. The generated code is passed to markup languages, where diagrams are created using JavaScript libraries. The system features a feedback loop that makes it robust. The LLM's backward learning ensures that the system can interpret the entire diagram in the current context. The system is easily accessible and fully voice-controlled, operating via keyboard shortcuts.
[0050] In one implementation, the user presses the spacebar on the keyboard, causing the microphone to capture the user's voice command. Upon releasing the spacebar, the microphone is turned off and processing of the captured voice command begins. The voice instruction given by the user for the current step is added, and the diagram on the canvas is modified accordingly. This occurs before user review to ensure that the current state is preserved and no loss occurs in the event of connection issues. With the version history feature, the system ensures that updating, resetting, and restoring diagrams is manageable and easy. After receiving updates, the system's text-to-speech agent explains the current state of the diagram to the user, points them to the drawn diagram, and allows them to reset the diagram, for example, delete it.
[0051] In one embodiment, the system operates in two modes: guided and free mode. Guided mode is similar to the drag-and-drop model. The system agent interacts with the user, prompting them with available options at each step, similar to drag-and-drop tools. This reduces user effort, as they only need to respond to the agent's requests and questions, after which the diagram is drawn.
[0052] Fig. Figure 2 shows how the S2S agent interacts with the user in the system's guided mode, guiding them through the available options at each step and prompting them to make their selections. In free mode, however, the system allows the user to freely draw diagrams. The user has complete freedom to initiate and follow along. Pressing the spacebar activates the microphone to record voice commands, which are then interpreted and the drawing area adjusted. The system's free mode allows the user to explore the scope of the drawing.
[0053] According to Fig.2, the system includes a speech conversion engine configured to convert the voice command captured from the user into a text transcript. The speech conversion engine leverages a browser-integrated speech recognition system, such as the WebSpeech API, enabling instant feedback and interactivity without a high computational load on the server side, while maintaining responsiveness and accessibility. The system further includes a Mermaid conversion engine or code generation module configured to convert the English commands into Mermaid code using trained large language models. The conversion engine uses the few-shot prompting technique for code generation and is configured to run a pre-trained general-purpose LLM through a limited number of example instances encoded in the prompt.This ensures minimal overhead, high adaptability to different input methods, and improved data control, while enabling accurate translation of language instructions into Mermaid syntax without retraining. The system also includes a code parser unit that converts the generated code into diagram blocks. The resulting Mermaid.js codes are converted into visual diagram blocks by the code parser, serving as a link between graphical output and textual representation. Based on the syntax and semantics defined in the Mermaid.js standard, the parser interprets the Mermaid code generated by the Mermaid Conversion Engine and transforms it into structured visual elements such as nodes, connectors, decision points, and labels. The visual representation performed by the code parser is dynamic and modular.This allows the system to update diagrams in real time upon user instructions by modifying the underlying code. This feature gives the system a real-time feedback loop by enabling voice-guided interaction with immediate audible confirmations, thus ensuring accessibility. The system ensures that users, especially those with visual impairments, can understand and improve their diagrams using non-visual means. The system also includes a drawing canvas that updates. All changes are saved to maintain version history. Versions are retained up to the latest state, even in the event of a connection loss. This ensures robustness and reliability. The system also features a feedback system that operates through a closed-loop interaction mechanism, enabling relearning, interactive corrections, and diagram interpretation in real time.After creating visual diagram elements based on the user's voice commands, the system enters a reflection phase in which the text-to-speech feedback module, the S2S agent, provides detailed acoustic descriptions of the current diagram state. This feedback enables users, especially those with visual impairments, to understand the representation by providing information about the number of nodes, their labels, directional flow patterns, decision points, and component interactions. The closed-loop feedback mechanism allows users to issue commands and receive meaningful, structured descriptions of their diagram output. This creates an interactive dialogue between user and system. The system offers intuitive control options via keyboard and voice commands to facilitate user interaction.Users utilize the various modes of the Progress Control panel; they access the controls either through voice commands such as "Next," "Previous," or "Restart," activated by the spacebar, or through keyboard shortcuts such as the "Back" arrow to go back or "R" to erase and restart the entire drawing. This step-by-step interaction method mimics the natural process of sketching, erasing, and redrawing, allowing users to edit specific parts of their diagrams without having to recreate the entire design.The feedback mechanism promotes reverse learning by providing detailed information about user changes and enabling immediate correction of errors, making the entire diagramming process interactive, reversible, and voice-guided, while ensuring accessibility, enhancing user understanding, and building confidence in diagramming.
[0054] The present invention provides an AI-assisted, interactive, voice-guided drawing system. The system is a robust tool for users with visual impairments and limited motor skills, enabling users to express themselves and draw. The system leverages technologies such as Gen AI, LLMs, speech, and feedback. The system allows users to create diagrams using natural language commands without relying on visual input or manual interaction. The system features an interactive, guided mode in which users can simply follow the instructions of the system agent, which is equivalent to drag-and-drop drawing tools. Furthermore, the system provides a free mode in which users can draw freely. They can create the diagram step by step by expressing their drawing wishes in commands.The system also features a continuous audio feedback mechanism that verbally communicates the current state of the diagram, allowing users, especially those with visual impairments, to track progress. The system offers a correction and reversal function, activated via voice commands, allowing users to instantly change, undo, or delete diagram elements.
[0055] The drawings and the foregoing description illustrate examples of embodiments. Those skilled in the art will recognize that one or more of the described elements may well be combined to form a single functional element. Alternatively, certain elements may be separated into multiple functional elements. Elements of one embodiment may be added to another embodiment. For example, the order of the processes described herein may be changed and is not limited to the manner described herein. Furthermore, the actions of a flowchart need not be performed in the order shown; nor do all actions need to be performed. Also, actions that are not dependent on other actions may be performed in parallel with the other actions. The scope of the embodiments is in no way limited by these specific examples.Numerous variations, whether explicitly stated in the specification or not, such as differences in structure, dimensions, and use of materials, are possible. The scope of the embodiments is at least as broad as indicated in the following claims.
[0056] Advantages, further benefits, and solutions to problems have been described above with reference to specific embodiments. However, the advantages, advantages, solutions to problems, and any components that may effect or enhance an advantage, advantage, or solution are not to be construed as critical, required, or essential features or components of any or all of the claims. REFERENCES 100 An AI-powered Interactive Voice-Controlled Drawing. 102 User input module 104 Speech recognition module 106 Natural Language Processing Module 108 Code generation module 110 Code parser unit 112 Drawing canvas unit 114 S2S Agent (Text-To-Speech Feedback Module) 116 User Interface Module 118 Progress control unit 120 Version control module 202 System 204 Guided Mode 206 Free Model 208 Capturing user input with a microphone. 208a Language Conversion Engine 208b Mermaid Conversion Engine 208c Code Parser 208d drawing area 208e Backward Learning 210 Current Status of the Diagram 212 users accepted changes 214 Save 216 User Rejects Changes 216a User rejection diagram
Claims
[1] An AI-based interactive voice-activated sign system (100) comprising: a user input module (102) that is triggered when the space bar is held down, the user input module (102) allowing the microphone to begin recording when the space bar is pressed, allowing the user to give commands and instructions verbally without the need for manual input; a speech recognition engine (104) for recording user speech when the user input module (102) is triggered, the speech recognition engine (104) configured to store the recorded speech and convert the spoken voice commands into text transcripts; a natural language processing module (106) connected to the speech recognition engine (104) and configured to receive the text transcripts from the speech recognition engine (104) and interpret the contextual meaning of the spoken voice commands; a code generation module (108) connected to the natural language processing module (106) and comprising a trained language model configured to convert the interpreted language commands into structured diagram code; a code parser unit (110) configured to receive the structural diagram code from the code generation module (108) and convert the structured diagram code into visual diagram elements; a drawing area unit (112) connected to the code parser unit (110) and configured to display the visual diagram elements as a visual diagram; a text-to-speech feedback module (114) as an S2S agent configured to generate acoustic descriptions of a current state of the visual diagram, wherein the S2S agent generates a description by drawing conclusions from the same model trained in the code generation module (108) after reverse learning, and wherein this generated summary is then displayed to the user; a user interface module (116) configured to provide voice-activated and keyboard shortcut interaction controls for the diagram change, allowing users to either accept the change, reject the current step, or perform a complete restart; a progress control unit (118) configured to allow the user to navigate to the next step of the drawing process, return to the previous step, or start again from the beginning using voice-based commands and keyboard shortcuts; and a version control module (120) configured to manage the version history of diagram changes and to enable reversion to previous diagram states; wherein the system (100) is configured to operate in a closed feedback loop in which the user receives audible feedback about the current state of the visual diagram after each voice command. [2] The system of claim 1, wherein the speech recognition engine (104) comprises: a browser-integrated speech recognition system (104a) configured to provide real-time, low-latency speech-to-text conversion without requiring backend processing. [3] The system of claim 1, wherein the code generation module (108) is configured to use few-shot prompting with a pre-trained large language model, convert transcriptions of natural language voice commands into Mermaid.js code syntax, and generate structured diagram code based on the contextual interpretation of user voice commands. [4] The system of claim 1, wherein the code parser unit (110) having a drawing surface (112) converts the Mermaid code into diagram blocks, the code parser unit (110) acting as a link between graphical output and textual representation, the code parser unit (110) interpreting the Mermaid code generated by the code generation unit (108) based on the syntax and semantics specified by the Mermaid.js code syntax and converting it into structured visual elements, including nodes, connectors, decision points, and labels. [5] The system of claim 1, wherein the user interface module (116) with version control module (120) facilitates drawing modification, wherein the diagram can be updated in real time by changing the underlying code in response to further voice instructions from the user or via keyboard shortcuts. [6] The system of claim 1, wherein the user input module (102) and the user interface module (116) comprise: a spacebar activation mechanism (122) configured to activate the speech recognition engine when pressed and deactivate it when released; keyboard shortcut controls configured to enable navigation and modification of diagrams; and voice-activated commands configured to enable diagram editing operations. [7] The system of claim 1, wherein the text-to-speech feedback module (114) is configured as an S2S agent to describe the current state of the visual diagram, including the number of nodes, node names, flow directions, decision points, and component interactions, by interpreting the current diagram code using the few-shot-large language model previously used after reverse learning; and to provide real-time audible confirmation of each diagram change; and to enable users to understand diagram contents in a non-visual manner, allowing the user to modify the diagram. [8] The system of claim 1, wherein the version control module (120) is configured to: automatically save diagram states after each change; maintain the complete version history of all diagram changes; enable reverting to previous diagram states through voice commands or keyboard shortcuts; and prevent data loss in case of connection problems. [9] The system of claim 1, wherein the system is configured to operate in multiple modes, including: a guided mode in which the S2S agent prompts the user at each step with available diagram options; and a free mode in which the user has the freedom to provide voice commands without guided prompts, wherein in the guided mode, the text-to-speech feedback module (114) is configured as the S2S agent to prompt the user to select from available diagram types, including flowcharts, graphs, and UML diagrams; the agent is configured to prompt the user at each step with available diagram elements; and the system is configured to create diagrams through an interactive voice dialogue between the user and the S2S agent. [10] The system of claim 1, wherein the system (100) further comprises: Error correction features configured to allow users to change specific diagram elements without having to recreate the entire diagram; step-by-step interaction controls configured to mimic natural sketching and erasing actions; and a backward learning feature configured to allow the system to interpret and describe complete diagrams in the current context.