Method special for long text voice recording scene and digital screen system
By combining a digital screen system with AI-powered intelligent script analysis and quick operation units, the problems of fragmented interaction and poor marking experience in long text voice recording scenarios have been solved, achieving efficient and seamless recording control and an immersive working experience, thus improving recording quality and user experience.
Patent Information
- Application Number
- CN202511544630.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-01-30
AI Technical Summary
Existing technologies suffer from problems such as fragmented interaction, low efficiency, poor tagging experience, high cognitive load, and cumbersome script preparation for multiple people in long text voice recording scenarios. In particular, in fields such as radio drama production, audiobook narration, and teaching courseware recording, existing equipment cannot provide an integrated and seamless workflow solution.
Employing a digital screen system, combined with an AI-powered intelligent script analysis engine, a pressure-sensitive stylus, and quick operation units, it enables intelligent document analysis, high-fidelity handwritten annotation, and seamless recording control. Uninterrupted operation is achieved through a Bluetooth shortcut keyboard and foot pedal, while gesture recognition and intelligent ink layers provide an immersive working experience.
It improves recording efficiency and continuity, reduces post-editing workload, enhances script preparation efficiency and recording quality, provides a highly focused work experience, and reduces user attention shifts and occupational health risks.
Smart Images

Figure CN121433604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and more particularly to the combination of computer peripherals and dedicated software systems. Specifically, it relates to a method and a digital screen system specifically designed for long text voice recording scenarios. Background Technology
[0002] Digital pen displays, as a type of pen input device that integrates a display panel and electromagnetic induction or capacitive technology, have traditionally been widely used in fields such as digital painting, industrial design, electronic signatures, and educational presentations. Their core advantage lies in providing a natural "what you see is what you get" writing and drawing experience.
[0003] However, in the professional and growing field of long-text speech recording, such as radio drama production, audiobook narration, film and television dubbing, and teaching material recording, practitioners (such as voice actors, announcers, and teachers) face a unique and complex workflow. Currently, the mainstream working methods in this field heavily rely on general-purpose computer hardware. Typically, staff use one or more ordinary computer monitors (such as Dell, AOC, etc.) to display electronic documents (such as scripts, lecture notes), while using a keyboard and mouse to perform a series of key operations, including but not limited to: page scrolling, text turning, keyword highlighting, content marking, and control of standalone recording software.
[0004] This workflow based on general-purpose devices has many inherent and insurmountable drawbacks, which seriously affect work efficiency, recording quality, and the work experience of practitioners:
[0005] Interactive Disruptions and Inefficiency: During recording, the core task for voice professionals is maintaining the continuity of voice and emotion. However, when it's necessary to turn pages or scroll through the text, they must interrupt their speech and move their hands from a comfortable position to the keyboard or mouse. This seemingly simple action actually disrupts breath control, ruins carefully crafted emotions, and causes unnatural pauses in the recording, increasing the difficulty and workload of post-production editing. For long monologues or intense dialogues requiring precise emotional control, this interruption is fatal.
[0006] The marking and annotation experience is poor: Professional voice actors are accustomed to making extensive personalized markings on paper manuscripts, such as circling stressed syllables, marking emotional transitions, underlining words and phrases requiring attention, and recording the director's revisions. Although existing document editing software or PDF readers also offer electronic marking functions, the detailed marking and underlining operations using a mouse are cumbersome, inefficient, and highly inconsistent with long-established human writing habits. This non-intuitive marking method fails to provide users with the freedom and immediacy of pen-and-paper interaction, thus inhibiting the creative inspiration of voice actors.
[0007] Cognitive Load and Distraction: In a typical recording environment, a user's attention needs to frequently switch between multiple focal points: the text in front of them, the recording software interface on the monitor (observing waveforms and timecodes), the mouse or keyboard in their hand, and external recording control devices. This multitasking state easily leads to cognitive overload, increasing the probability of errors, such as missing lines, misreading markings, or misoperating the recording software. The ideal working state should allow the user to be fully immersed in the text and emotional expression.
[0008] The complexity of multi-person script processing: In projects involving multiple characters, such as radio dramas and multi-person audiobooks, the initial script preparation work is particularly arduous. Staff members need to manually read through the entire script, manually distinguish the lines and narration of different characters, and highlight or mark the characters they are playing. This process is time-consuming and labor-intensive, and is very easy to make mistakes due to negligence, especially in complex scripts with many characters and frequent overlapping dialogues.
[0009] Currently, although there are software patents on the market involving long text processing or speech synthesis, they mainly focus on the automatic generation of text to speech or backend data processing. They do not provide solutions for the front-end human-computer interaction needs of voice workers in the specific scenario of "recording," nor do they propose a dedicated device concept that deeply integrates the hardware form of digital screen with the recording workflow.
[0010] Therefore, the industry urgently needs a new solution that can overcome all the above-mentioned drawbacks, namely a dedicated hardware device and its supporting system that deeply integrates "intelligent text parsing, high-fidelity handwritten marking, seamless recording control and integrated display", so as to liberate voice workers from tedious, inefficient and fragmented operations and return them to the core of creation and performance. Summary of the Invention
[0011] Technical problems to be solved
[0012] The purpose of this invention is to overcome the shortcomings of existing technologies in long text speech recording, such as fragmented workflows, low interactive efficiency, poor annotation experience, high cognitive load, and cumbersome script preparation for multiple users. This invention aims to provide a method and digital screen system specifically designed for long text speech recording scenarios, achieving a highly focused and immersive work experience through an integrated "script preparation-recording" process.
[0013] Technical solution
[0014] To achieve the above objectives, the present invention provides a method specifically for long text speech recording scenarios. The method is executed on a digital screen system comprising a digital screen hardware body (1), a pressure-sensitive stylus (2), and a shortcut operation unit (3).
[0015] The specific steps of this method are as follows: First, the computer receives and imports the electronic document (step S1). Next, the computer calls a core AI intelligent script parsing engine to perform natural language processing (NLP) analysis on the imported electronic document (step S2). The core task of this step is to automatically identify one or more character information contained in the document and format and render all or part of the text content of the document according to the identified character information. Then, the electronic document formatted and rendered by the AI engine is displayed on the high-resolution screen of the digital screen hardware body (1) (step S3). In order to achieve a natural marking experience, the computer will virtually build an intelligent ink layer on the displayed electronic document (step S4). This ink layer is specifically used to receive and record all handwritten annotation information generated by the user through the pressure-sensitive stylus (2) in real time without damaging the original document data. Finally, in order to achieve uninterrupted recording control, the computer will continuously listen to and receive control commands from the shortcut operation unit (3) (step S5). These instructions are clearly divided into two categories: the first control instructions, which are used to control page scrolling or page turning in electronic documents; and the second control instructions, which are used to directly control the start or stop of the recording process.
[0016] As a preferred solution, the AI intelligent script parsing engine assigns a unique, visually exclusive color to each identified character information when performing formatted rendering, and uses this color to color all the dialogue text corresponding to that character, thereby visually distinguishing the content of different characters.
[0017] As another preferred option, the AI-powered script analysis engine further integrates a sentiment analysis and annotation module. This module performs in-depth sentiment analysis on the dialogue text and, based on the analysis results, automatically generates suggestive annotations at specific dialogue locations to guide the performers' interpretation, such as "(urgently)" or "(reluctantly)".
[0018] At the hardware interaction level, as a preferred solution, the shortcut operation unit (3) is designed as a combination comprising a wireless Bluetooth shortcut keyboard (3a) and a foot pedal (3b) connected via a USB interface. The first control command for page turning or scrolling is triggered by the physical buttons on the Bluetooth shortcut keyboard (3a), while the second control command for controlling the start and stop of recording is triggered by the stepping action of the foot pedal (3b), thus completely freeing the user's hands.
[0019] To achieve more precise document navigation, as a preferred solution, the Bluetooth shortcut keypad (3a) integrates a scroll knob. When the first control command is to control page scrolling, the command is triggered by the user rotating this knob, thereby enabling precise and smooth scrolling of long documents line by line or even pixel by pixel.
[0020] At the software level, as a preferred solution, the intelligent ink layer is designed as a logically transparent layer. The computer system allows users to hide all annotations on this ink layer with a single click at any time, facilitating a clean reading of the original manuscript; it also supports merging these annotations with the original electronic document and exporting them as a new document file with handwritten markings.
[0021] To further enhance the smoothness of the interaction, as a preferred option, this method also includes a gesture recognition step. The computer recognizes single-finger or two-finger touch swipe actions performed by the user on the surface of the digital display hardware (1) and converts these specific gesture patterns into control commands for page scrolling or view zooming, as a supplement to keyboard shortcuts.
[0022] To ensure the integrity of the work results, as a preferred option, this method also includes a file association step. The computer automatically and logically associates the audio files generated by the recording process controlled by the second control command with the currently processed electronic document project and stores them together for easy subsequent management and traceability.
[0023] Furthermore, as a preferred embodiment, the quick operation unit (3) is also equipped with a dedicated marking mode switching key. The computer receives a third control command triggered by this key to quickly and conveniently switch between the "writing mode" and "erasing mode" of the pressure-sensitive stylus (2) without having to select it in the software interface.
[0024] The present invention also provides a digital screen system specifically designed for long text voice recording scenarios. The system includes: a digital screen hardware body (1), a pressure-sensitive stylus (2), a shortcut operation unit (3), and a computer processing unit. This computer processing unit is specifically configured to run the various modules of the above-described method, including: an AI intelligent script parsing engine, an intelligent ink layer module, and a recording control module, working together to achieve the objectives of the present invention.
[0025] Beneficial effects
[0026] Compared with existing technologies, this invention, through deep integration and co-design of software and hardware, brings one or more of the following significant beneficial effects:
[0027] Significantly improved recording efficiency and continuity: This invention integrates all core interactive behaviors, including script preparation (AI analysis), annotation (stylus), narration, and control (shortcut keys, pedals), into a closed-loop system. Users can control recording via foot pedal and turn pages by lightly pressing shortcut keys with their fingertips, without interrupting their voice to operate the keyboard and mouse. This ensures absolute continuity in the recording process and seamless emotional expression, greatly reducing the workload of post-editing.
[0028] Intelligent and automated script preparation process: The built-in AI intelligent script analysis engine can automatically complete the most time-consuming and error-prone tasks of character segmentation and dialogue color-coding, and can provide preliminary emotional annotation suggestions. This greatly reduces the burden of early preparation for multi-person script projects, improves the efficiency and accuracy of script preparation, and allows users to enter the core stages of character analysis and performance more quickly.
[0029] A highly focused and immersive work experience: Full-screen document display, a natural pen-and-paper marking experience, and a design that simplifies control operations to muscle memory (such as foot pedals and finger presses) minimizes the user's attention switching between documents, software, and control devices. This integrated interaction method creates a "flow" environment, helping users better immerse themselves in the document content, invest their emotions and state of mind, thereby improving the final recording quality.
[0030] Human-centered design and occupational health considerations: The high-precision pressure-sensitive pen provides an annotation experience comparable to real pen and paper, perfectly aligning with the long-established work habits of professionals. Combined with a foot pedal, it completely frees the user's hands, avoiding shoulder, neck, and wrist strain that may result from frequent keyboard and mouse operation, thus contributing to the long-term occupational health of professionals.
[0031] Specialized Functionality and Powerful Performance: This system is deeply optimized for the specific scenario of long text voice recording, solving the core pain point of general-purpose computer equipment being "usable but not user-friendly" in this field. The design of every hardware control and the implementation of every software function are directly aimed at the specific needs of this scenario, providing a level of professionalism and ease of use that general solutions cannot match. Attached Figure Description
[0032] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of the overall structure of a digital screen system according to an embodiment of the present invention.
[0034] Figure 2 This is a schematic diagram of the Bluetooth shortcut key layout in a digital screen system according to an embodiment of the present invention.
[0035] Figure 3 This is a flowchart of a working method according to an embodiment of the present invention. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0037] Example 1: System Structure and Method Details
[0038] Please see Figure 1 This embodiment provides a digital screen system specifically designed for long text voice recording scenarios. Logically, the system consists of hardware and software components that are deeply coupled and work together.
[0039] Hardware system composition
[0040] The hardware system mainly includes: the digital screen hardware body (1), the pressure-sensitive stylus (2), and the quick operation unit (3).
[0041] Digital display hardware (1): This is the core display and input platform of the system. In this embodiment, it is a high-resolution (e.g., 2.5K or 4K level) LCD screen that integrates high refresh rate and low latency pen input sensing technology to ensure that the handwriting of the pressure-sensitive stylus (2) can be captured and displayed quickly and accurately. In order to adapt to long-term reading and annotation work, the screen surface has been treated with anti-glare etched glass, which can effectively reduce the reflection of ambient light and provide a paper-like visual experience and writing damping feel.
[0042] Pressure-sensitive stylus (2): This is a passive electromagnetic induction or active capacitive stylus with high-precision pressure sensing capabilities (e.g., 8192 levels or higher). This means that the system can accurately sense subtle changes in pressure applied by the user to the pen tip and convert them into attributes such as the thickness and density of the digital handwriting, thus highly simulating the experience of writing with a real pen or pencil on paper. The pen body usually also has one or two side buttons with customizable functions.
[0043] Quick Operation Unit (3): This unit is key to achieving "uninterrupted" interaction in this invention. In this embodiment, it consists of two independent physical devices:
[0044] Bluetooth Shortcut Keyboard (3a): This is a small, wireless keyboard device that can be placed in a convenient location for the user. See also Figure 2 Its dial layout is meticulously designed, containing only a few core buttons crucial to the recording process. These include:
[0045] Previous / Next Page keys: Used for large-scale page switching.
[0046] Scroll knob: Used for fine and smooth scrolling through a document, especially suitable for scenarios where you need to go back to the previous text or align with a specific line.
[0047] Marking mode switch key: Used to switch between "writing mode" and "erasing mode" with one click, eliminating the hassle of selecting tools on the screen.
[0048] Color selection button: Allows users to quickly cycle through several preset brush colors (such as red, blue, and black) for easy marking of different types.
[0049] One-button voice control switch: used as a supplement or replacement for the foot pedal to control the start and stop of recording.
[0050] Foot pedal (3b): This is a device that connects to a computer via a USB-C or USB-A interface, resembling a musical instrument's sustain pedal. Its core function is preset to control the "start / pause" and "stop" of recording. Transferring recording control to the feet is one of the core design features of this invention that frees the user's hands. Throughout the entire presentation, the user's hands can remain in a most comfortable position, or be used to operate the shortcut keyboard and stylus, without needing to move their arms to click the mouse.
[0051] Software system composition
[0052] The software system is a dedicated document preparation software that runs on a computer connected to a digital display. Its core modules include:
[0053] Document rendering engine: This engine is specifically optimized for displaying long text. It can efficiently parse and render common document formats (such as .docx, .txt, .pdf) and provides rich layout adjustment options, such as font, font size, line spacing, background color (e.g., eye-friendly beige or night mode), ensuring optimal readability of documents on digital screens.
[0054] AI-powered script analysis engine: This is the core of the system's intelligent script preparation. This engine deeply integrates Natural Language Processing (NLP) technology and contains several key sub-modules:
[0055] Character Recognition Module: This module is activated when a user imports a script-formatted manuscript. It scans and analyzes the entire text using a hybrid model based on rules (such as recognizing the "character name: dialogue" format) and machine learning, automatically identifying and extracting all different character names (such as "Zhang San" and "Li Si") as well as narration.
[0056] Automatic formatting and coloring module: After character recognition is complete, this module assigns a unique, high-contrast color to each identified character and automatically renders all dialogue text corresponding to that character in that color. Simultaneously, it performs uniform formatting on the entire text, such as uniform indentation and adjusting spacing, making the structure of the entire script clear at a glance. As described in claim 2, the specific computer implementation of this process is as follows: The engine first traverses the manuscript, constructing a set containing all unique character names. Then, the system creates a data structure (such as a hash table or dictionary) that maps each unique character name to a predefined color value (e.g., {"Zhang San":"#FF0000","Li Si":"#0000FF","Narrator":"#000000"}). Finally, when rendering text line by line or paragraph by paragraph, the manuscript rendering engine first determines which character the text block belongs to, then queries the mapping table to obtain the corresponding color value, and applies it as the display color of the text block.
[0057] Sentiment Analysis and Annotation Module: As described in claim 3, this module aims to provide intelligent assistance to the user's interpretation. Specifically, for each line of dialogue, the system uses it as input and passes it to a deep learning model (such as a text classification model based on the Transformer architecture) pre-trained with a large sentiment corpus. This model analyzes the semantics and context of the dialogue and outputs a probability distribution containing multiple preset sentiment categories (such as "joy," "sadness," "anger," "urgency," "reluctance," etc.). The system selects the sentiment tag with the highest probability and converts it into human-readable annotation text (e.g., the tag "urgent" is converted into the string "
urgently
[0058] Intelligent Ink Layer Module: As described in claim 6, this is the key to achieving non-destructive, free annotation. In terms of technical implementation, when rendering a document, the graphics rendering engine creates a logically independent, completely transparent overlay layer (Layer B) above the document content layer (Layer A) and assigns it a higher rendering priority (Z-index). All writing and drawing operations performed by the user using the pressure-sensitive stylus (2) will generate handwriting data (including coordinates, pressure, tilt, etc.) that will be drawn in real time on this transparent Layer B in the form of vector graphics. Because the two layers are separate, any operation on Layer B will not affect the original document data of the underlying Layer A. When the user triggers the "Hide Annotation" function, the system only needs to set the visibility attribute of Layer B to false. When the user selects "Export Document with Annotations", the system will perform a layer merging rendering in the background: drawing the contents of Layer A and Layer B together on a temporary canvas, and then exporting the merged complete image as a PDF or image file.
[0059] Recording Control Module: This module acts as a bridge between hardware operation and recording software. It deeply integrates control capabilities for mainstream recording software, typically achieved by calling the application programming interfaces (APIs) provided by these software programs or by simulating system-level media control shortcuts. As described in claim 4, when the user presses the foot pedal (3b), the computer receives a USB input signal. The system's device driver interprets this signal as a specific event. Upon detecting this event, the recording control module immediately calls the corresponding API function to send a "start recording" or "stop recording" command to the recording software.
[0060] Gesture and Stroke Recognition Module: To provide richer interaction dimensions, this module is responsible for handling touch input from the digital display hardware (1). As described in claim 7, its implementation principle is as follows: The system's touch event listener continuously captures touch point data on the screen, including the number of touch points, coordinates, movement speed, and direction. When a finger is detected continuously sliding vertically on the screen, the event handler determines it as a "scrolling" gesture and calculates the appropriate scrolling distance accordingly, updating the document view. When the ALI detects that the distance between two fingers is continuously increasing or decreasing, it determines it as a "zooming" gesture and calculates the corresponding zoom ratio to zoom in or out of the document view. In addition, this module can also be configured to recognize specific strokes, such as quickly drawing a horizontal line with a pen to represent deletion.
[0061] Detailed Explanation of Work Methods and Procedures
[0062] Please see Figure 3Based on the above system configuration, the working method of the present invention is as follows:
[0063] Power on and connect: The user connects the digital display hardware (1) to the computer via a data cable (such as HDMI / USB-C) and ensures that the Bluetooth shortcut keyboard (3a) and foot pedal (3b) are also successfully connected.
[0064] Importing a manuscript: Users launch the dedicated manuscript preparation software, select and import an electronic manuscript (e.g., a Word document for a radio drama).
[0065] AI-powered automated processing: Upon importing the script, the AI-powered intelligent script analysis engine immediately executes its functions. It first identifies the characters and determines the location of the lines, then invokes the automatic formatting and coloring module to render different colors for the lines of different characters. Next, the sentiment analysis and annotation module activates, generating performance prompts for some lines. The entire process typically takes only a few seconds. Finally, a script with distinct colors, a clear structure, and intelligent annotations is presented to the user.
[0066] Personalized annotation: Users can use a pressure-sensitive stylus (2) to make personalized annotations directly on the digital screen, just like on a paper manuscript. For example, they can highlight their lines, take notes in blank spaces, and mark breathing points that require special attention. If modifications are needed, the marking mode switch key on the Bluetooth shortcut key (3a) can be pressed to switch the pen to eraser mode. As described in claim 9, the implementation mechanism is as follows: the system maintains a state variable, such as currentTool="pen". When the marking mode switch key is pressed, the system will perform a logical judgment. If the current state is "pen", it will be changed to "eraser", and vice versa. All subsequent input data from the stylus will be dispatched to different processing functions (drawing or erasing strokes) according to the value of this state variable.
[0067] Entering Recording Mode: After preparing the text, the user clicks the "Full Screen Recording" button on the software interface. At this time, the software interface will hide all unnecessary elements, leaving only the maximized text content, creating an immersive environment without distractions.
[0068] Start and control recording: After the user adjusts the microphone and posture, they lightly step on the foot pedal (3b). The recording control module receives the instruction, and the background recording software immediately starts recording.
[0069] Narration and Seamless Page Turning: The user begins narration. During narration, as the content of a page is about to end, the user's finger naturally rests on the Bluetooth shortcut keypad (3a), and gently presses the "Next Page" key. The text immediately and seamlessly switches to the next page. Throughout the process, the user's voice and breath remain completely continuous, and the narration is uninterrupted. If fine-tuning of the display position is required, it can be achieved by rotating the scroll knob. As described in claim 5, the principle of the scroll knob achieving fine scrolling is as follows: the encoder inside the knob generates a series of pulse signals when rotated. After receiving these signals, the computer decodes them into scrolling events with direction and amplitude. The software accumulates these tiny scrolling events and adjusts the display coordinates of the text content in real time and smoothly, thereby producing a smooth scrolling effect, rather than the jumping feeling of traditional page turning.
[0070] Pause and Re-encode: If a mistake is made during the narration, the user can simply press the foot pedal (3b) again to pause the recording. The user can use the stylus to mark the location of the mistake, then scroll the text view to the point where they need to restart, compose themselves, and press the foot pedal again to resume recording from the break point.
[0071] Processing, Saving, and Exporting: After all recording is complete, the user presses the foot pedal to stop recording. At this point, some simple editing or re-recording can be performed. Finally, select "Complete Recording." As described in claim 8, the system will automatically perform file association operations at this time. Specifically, when creating a project, the system generates a metadata file (e.g., XML or JSON format) such as a .project file, which records the path to the original script file. Whenever a recording task is completed and an audio file (e.g., take_01.wav) is generated, the system appends the complete path of this audio file to the .project file.<audio_files> Under this node, all related scripts and audio files will be automatically loaded the next time the user opens the project. The user can then package the script with personal handwritten annotations (achieved through the export function of the Smart Ink Layer module) and all associated audio files into a single folder and submit it to the post-production team.
[0072] Specific application scenario examples
[0073] Voice actress Xiao Wang is recording her role for a radio drama. She's using the system described in this invention. First, she connects her digital screen to her studio computer and opens the dedicated script preparation software. She drags the Word script file sent by the director directly into the software window. Almost instantly, the AI engine completes the processing. The script is automatically formatted, all the lines of her character "Li Si" are highlighted in blue, the lines of her antagonist "Zhang San" are in red, and the narration is in default black. Simultaneously, emotional annotations such as "[urgently]" and "[reluctantly]" appear above some key lines, giving her a first impression of the character's emotions. Next, Xiao Wang picks up a pressure-sensitive stylus and begins to carefully read the script. She directly circles what she considers to be stressed words on the screen with a virtual red pen, underlines long sentences requiring a single, continuous stroke with a yellow highlighter, and writes her understanding and notes about the character's inner thoughts in the margins. The whole process is as natural and smooth as writing on paper. After completing the script preparation, she puts on headphones and enters full-screen recording mode. She lightly pressed the foot pedal with her right foot, and the recording began. Throughout the performance, her hands remained completely relaxed, or occasionally she used her left hand to operate the Bluetooth shortcut keyboard beside her. When she was almost finished reciting a page of lines, she would press the "next page" button in advance, and the screen content would seamlessly continue, ensuring that her performance's emotion and breath were never interrupted. When she was not satisfied with her performance of a particular line, she would press the foot pedal again to pause, use a pen to locate the position on the screen, and then press the foot pedal again to seamlessly restart recording from the previous position. After all the recordings were completed, she exported the final version of the script, which included AI intelligent color separation and her personal handwritten annotations, along with all the audio files, into a project folder with one click and sent it to the director via the internet. The entire workflow was several times more efficient than the previous method of using a regular monitor and keyboard and mouse, and the integrity and quality of the recordings were also much higher.
[0074] Example 2: Integrated Mobile Terminal Form Factor
[0075] The core innovation of this invention lies in optimizing the workflow and interactive experience of long text voice recording through deep integration of software and hardware. This idea is not limited to the specific product form of "external digital display" in the above embodiments.
[0076] In another embodiment, the present invention can be implemented as a highly integrated, all-in-one mobile terminal device. Its specific configuration is as follows:
[0077] Built-in computing unit: The device integrates a low-power, high-performance ARM architecture processor, a solid-state drive (SSD), and sufficient RAM. This means it is a complete computer in itself, without the need for any external host.
[0078] Lightweight Custom Operating System: Runs a highly optimized operating system based on Linux or Android with extensive customization. This system removes all functions and applications unrelated to recording, directly launching into a dedicated recording software interface upon boot, ensuring system stability, efficiency, and simplicity.
[0079] Built-in high-quality recording module: The device integrates an array of one or more MEMS microphones, equipped with a professional audio codec (CODEC) and preamplifier circuitry, enabling direct broadcast-quality audio recording without the need for an external sound card and microphone. It also retains Bluetooth audio functionality, supporting connection to wireless monitoring headphones.
[0080] Built-in battery and wireless connectivity: The device features a high-capacity built-in battery, supporting hours of offline work and greatly enhancing portability. Users can record in any quiet environment (such as at home or in a hotel). Additionally, built-in Wi-Fi and Bluetooth modules allow for easy reception of documents from the cloud or local network, as well as wireless export of completed recordings and annotations.
[0081] Interface: Retains a USB-C interface for charging and connecting external devices, such as the aforementioned foot pedal.
[0082] The advantages of this integrated mobile terminal form factor are:
[0083] Ultimate portability and mobility: Completely free from dependence on traditional desktop or laptop computers, enabling professional recording anytime, anywhere.
[0084] Higher system integration and efficiency: The integrated hardware and software design makes the system respond faster and the user experience smoother.
[0085] A cleaner work environment: The wireless design minimizes the constraints of cables, making the workspace more concise and organized.
[0086] This embodiment powerfully demonstrates that the scope of protection of the present invention should not be limited to its specific physical implementation, but should cover all systems and methods that adopt the core innovative concept of the present invention—namely, using AI to process documents, combining natural annotation with a stylus, and achieving uninterrupted recording control through dedicated and convenient hardware.
[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method specific to long text speech recording scenarios, characterized in that, The method is executed based on a digital screen system, which comprises a digital screen hardware body (1), a pressure-sensitive stylus (2) and a shortcut operation unit (3), and comprises the following steps: S1: receiving and importing an electronic manuscript by a computer; S2: calling an AI intelligent script analysis engine by the computer to perform natural language processing on the electronic manuscript to automatically identify one or more role information in the manuscript, and to format and render the text content of the electronic manuscript according to the role information; S3: displaying the formatted and rendered electronic manuscript on the digital screen hardware body (1); S4: establishing a smart ink layer on the displayed electronic manuscript by the computer to receive and record the annotation information generated by the pressure-sensitive stylus (2) in real time; S5: receiving the first control instruction from the shortcut operation unit (3) to control the page scrolling or page turning of the electronic manuscript, and receiving the second control instruction from the shortcut operation unit (3) to control the start or stop of the recording process.
2. The method of claim 1, wherein, The specific implementation of the S2 step of format rendering is: Step 2.1: the AI intelligent script analysis engine traverses the manuscript to extract all unique role name strings; Step 2.2: the computer creates a role-color mapping table data structure and assigns a preset color value to each unique role name string extracted in step 2.1; Step 2.3: the computer's manuscript rendering engine renders the manuscript, judges each monologue text, obtains its corresponding role name, queries the mapping table to obtain the corresponding color value, and finally applies the color value to the font color attribute of the monologue text for display.
3. The method according to claim 1 or 2, characterized in that, The S2 step also includes emotion analysis and annotation, which is implemented as follows: Step 3.1: the AI intelligent script analysis engine transmits the recognized monologue text as input to a pre-trained emotion classification neural network model; Step 3.2: the model calculates the input monologue text and outputs one or more preset emotion label confidence scores; Step 3.3: the computer selects the preset emotion label with the highest confidence score and generates the corresponding text annotation according to the label; Step 3.4: the computer inserts the generated text annotation into the corresponding monologue text and displays it together.
4. The method of claim 1, wherein, The shortcut operation unit (3) comprises a Bluetooth shortcut keyboard (3a) and a foot pedal (3b); the computer processes the instructions by the following steps: Step 4.1: the device driver of the computer listens to the hardware input events from the Bluetooth shortcut keyboard (3a) and the USB signals from the foot pedal (3b) respectively; Step 4.2: when a specific key hardware input event is listened to, it is interpreted as the first control instruction, and the page control function of the system is called; Step 4.3: when the USB signal of the foot pedal is listened to, it is interpreted as the second control instruction, and the application programming interface of the recording software is called to switch the recording state.
5. The method of claim 4, wherein, The Bluetooth shortcut keyboard (3a) is integrated with a scroll knob, which realizes fine scrolling through the following steps: Step 5.1: The computer receives continuous rotation input events generated by the scroll knob, each event containing a direction and an incremental value; Step 5.2: The computer converts the incremental value of the received rotation event into a scrolling distance in pixel units; Step 5.3: The computer updates the vertical coordinate of the view port of the electronic document display area in real time according to the scrolling distance in pixel units, thereby realizing the visual effect of smooth scrolling.
6. The method of claim 1, wherein, The implementation and operation steps of the smart ink layer are as follows: Step 6.1: The computer creates a transparent overlay layer with a higher Z-index value in the graphics rendering pipeline above the rendered electronic document layer as the smart ink layer; Step 6.2: The coordinate and pressure data generated by the pressure-sensitive stylus (2) are drawn as vector strokes on the transparent overlay layer in real time; Step 6.3: When a hiding instruction is received, the computer sets the visibility attribute of the transparent overlay layer to "off"; Step 6.4: When an export instruction is received, the computer performs a layer merging operation to render the electronic document layer and the transparent overlay layer into an off-screen buffer, generating a single image or document containing all visual information, and saving it as a file.
7. The method of claim 1, wherein, The method further includes a gesture recognition step, which is implemented as follows: Step 7.1: The touch event processor of the computer continuously monitors the touch input of the digital screen hardware body (1) and recognizes the number of touch points and motion trajectories; Step 7.2: When the processor recognizes a single touch point moving continuously in the vertical direction, it is determined to be a scrolling operation, and the scrolling position of the document view is adjusted accordingly; Step 7.3: When the processor recognizes a continuous change in the distance between two touch points, it is determined to be a zooming operation, and a zooming scale factor is calculated accordingly to zoom in or out the document view.
8. The method of claim 1, wherein, The method further includes a file association step, which is implemented as follows: Step 8.1: The computer creates a project metadata file for the current task, which records the file path of the electronic document; Step 8.2: When the second control instruction controls the generation of an audio recording file, the computer obtains the storage path of the audio recording file; Step 8.3: The computer writes the obtained audio recording file path as a new record into the project metadata file, thereby logically binding the audio recording file and the electronic document under the same project.
9. The method of claim 1, wherein, The shortcut operation unit (3) is also provided with a marking mode switching key, and the computer responds through the following steps: Step 9.1: The computer maintains a tool state variable in memory to define the current function of the stylus, which can be "writing" or "erasing"; Step 9.2: The computer listens to the key press events of the marking mode switching key; Step 9.3: Once the key event is detected, a state switching function is triggered, which checks the current value of the tool state variable and switches it to another preset value, i.e. from "writing" to "erasing", or from "erasing" back to "writing"; Step 9.4: Subsequent input received by the pressure-sensitive stylus (2) will be interpreted as an operation of drawing or deleting handwriting according to the value of the switched tool state variable.
10. A digital signage system specialized for long text voice recording scenarios, characterized by, Comprise: A digital screen hardware body (1) for displaying electronic documents; A pressure-sensitive stylus (2) for handwritten input on the digital screen hardware body (1) to generate annotation information; A shortcut operation unit (3) for generating and sending first and second control instructions for page navigation and recording control; And a computer processing unit configured to: Run an AI intelligent script analysis engine for processing imported electronic documents to automatically identify role information and format render text content; Run a smart ink layer module for receiving and recording annotation information generated by the pressure-sensitive stylus (2) on top of the electronic documents rendered on the digital screen hardware body (1); Run a recording control module for responding to the first and second control instructions from the shortcut operation unit (3) to perform page navigation and recording start-stop operations without interrupting the recording subject process.