SYSTEM FOR CONVERTING VOICE INTO COMMANDS
A local BERT-based voice command system for embedded devices processes commands efficiently, addressing the challenge of limited resources by using a tokenizer, text classifier, and parsers, achieving fast and network-independent device control.
Patent Information
- Application Number
- BR102025005719
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-19
- Filing Date
- 2025-03-25
- Publication Date
- 2026-07-07
AI Technical Summary
Implementing voice commands in embedded systems with limited computing resources and network access is challenging due to the need for large cloud-based language models.
A local voice command system utilizing a compact BERT-based text classifier and parsers to process voice commands, including a tokenizer, text classifier, string parser, and time parser, enabling command execution without network dependence.
Enables intuitive device control with millisecond response times and reduced dependency on external resources, ensuring seamless operation in embedded systems.
Smart Images

Figure 00000000_0000_ABST
Description
1 / 21 SYSTEM FOR CONVERTING VOICE INTO COMMANDS BACKGROUND
[0001] Implementing voice commands for electronic devices with limited computing resources utilizes large cloud-based language models to process text. Implementing voice commands in embedded systems, which generally have limited computing resources and sometimes lack network access to access cloud resources, is difficult. SUMMARY
[0002] A computer-implemented method for deriving voice commands involves receiving a voice request to execute a command to a device. A text string is obtained corresponding to the voice request. The text string is tokenized to generate text string tokens. The text string tokens are fed into a trained text classifier to identify command classes corresponding to the voice request, the command classes including a search class command, a timing class command, and other commands. Information is extracted from the text string tokens via a string parser for a search class command. Timing information is extracted from the text string tokens via a timing parser, and the resulting commands and extracted information are provided to the device for execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] FIG. 1 is a block diagram of an enhanced voice interaction system for command according to an example embodiment.
[0004] FIG. 2 is a block diagram illustrating the parameterization of an example network architecture for a Petition 870250023302, dated 03 / 25 / 2025, page 19 / 49 2 / 21 Text classifier model according to an example modality.
[0005] FIG. 3 is a block diagram that illustrates a parameterization of a network architecture for the string parser model according to an example embodiment.
[0006] FIG. 4 is a flowchart that illustrates a speech processing method for commands according to an example modality.
[0007] FIG. 5 is a schematic block diagram of a computer system for implementing one or more example modes. DETAILED DESCRIPTION
[0008] The following description refers to the accompanying drawings which form part hereof, and in which specific embodiments that can be practiced are shown by way of illustration. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention, and it should be understood that other embodiments may be used and that structural, logical and electrical changes may be made without departing from the scope of the present invention. The following description of examples of embodiments should therefore not be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0009] Voice-based interaction allows users to issue commands without traditional input devices such as a mouse, keyboard, or touchscreen. Natural Language Processing (NLP) is crucial for understanding user input, accommodating diverse expressions and nuances. AI models analyze the semantics of commands converted into text, interpreting intent and extracting relevant information.
[0010] A typical AI model used for text processing is called a word bag model, which uses text vectors. Petition 870250023302, dated 03 / 25 / 2025, page 20 / 49 3 / 21 representing text as collections of words with frequencies and models based on word embedding. Another type of model for text processing is Bidirectional Encoder Transformer Representations (BERT), which uses bidirectional attention to capture contextual information from preceding and subsequent words. Compact versions of the base BERT model, which consists of 110 million parameters, can be obtained by varying the number of Transformer layers (L), the size of the representation vector (H), and the number of attention headers (A). An example of a compact model, called Small BERT (L = 2, H = 128, and A = 2), comprises 4.4 million parameters. An increase in the application of preprocessing techniques, such as the elimination of empty words and punctuation, has resulted in a decrease in model training accuracy.This suggests the need for careful consideration of preprocessing techniques, as they can potentially increase the model's error rate.
[0011] An enhanced voice-command interaction system for embedded devices allows users to control devices intuitively and conveniently. Inference processes can be performed locally, ensuring results on the order of milliseconds and avoiding dependence on external resources.
[0012] FIG. 1 is a block diagram of an enhanced voice interaction system 100. A high-level description of the system 100 is provided, followed by further details of each of the elements of the system 100.
[0013] System 100 receives speech 110 which is recognized by speech recognizer 115 to provide text 120 corresponding to speech 110. In one example, speech recognizer 115 may run locally on system 100, which may be a local device, or remotely, such as via networked cloud computing resources. Petition 870250023302, dated 03 / 25 / 2025, page 21 / 49 4 / 21
[0014] The recognized text 120 is preprocessed in a preprocessing tokenizer 125. The tokenizer 125 divides the text 120 into smaller units, such as words or portions of words, and maps the units to unique token identifiers (IDs). The token IDs are used as input 130 to a text classifier 135.
[0015] The text classifier 135, in one example, is a model with BERT architecture and is used for intentional classification of text commands into command classes. The text classifier 135 can be trained for specific commands to be executed by the local device. The specific commands can be a limited set of commands corresponding to commands acceptable and executable by the local device. The text is then classified into mapped classes 140.
[0016] The mapped classes are provided to a decision block 145 that routes commands that include searching via 150 to a string parser 155. The string parser 155 can be a BERT-based token classification model, for example, to identify relevant text segments to be used to execute the search command.
[0017] Example of text input with relevant segments (search terms or search string) of text to be searched includes: "Look for healthy recipes for dinner." Search for movie reviews online and find information about popular people online.
[0018] Decision block 145 will route commands involving time, such as the time of a meeting, via 160 to a time parser 165. In an example, time parser 165 might use Regex rules and receive a list of words extracted from input 130. The rules are applied to map the words. Petition 870250023302, dated 03 / 25 / 2025, page 22 / 49 5 / 21 for different types of labels. A first label is not related to time, a second label can include a pattern, such as multiple times. AM, PM, relative days like tomorrow, absolute days identified by date, and even months are additional labels that can be mapped.
[0019] Examples of text input and the corresponding output of time parser 165 include: Set a timer for 30 seconds and 25 minutes. Output: ([30,25], [“sec,“min]) “Set an alarm for tomorrow at 4 PM: Output: ([1,4,0,1]), [“relative_day,“hour,“min,“period])
[0020] Decision block 145 will route, via 170, commands that are neither search commands nor time-related directly for execution in commands 175 to a device 180, such as an interactive display. Device 180 may incorporate system components 100 that can be incorporated into device 180 and not require a network connection to access additional processing resources to convert text into commands 175. In one example, device 180 may also incorporate speech recognition 115 or rely on a fast network connection to perform speech recognition.
[0021] If string parser 155 or time parser 165 is invoked, information will be extracted to parameterize the command or provide information to correctly execute command 175.
[0022] Text classifier 135 is optimized to run as an embedded template on the local device. An embedded template is a template that runs on a local device without the need for a network connection. Text classifier 135 can be significantly reduced in size by limiting the number of command classes that are executable by the local device once identified. The use of string parser 150 and time parser 165 eliminates the need for a network connection. Petition 870250023302, dated 03 / 25 / 2025, page 23 / 49 6 / 21 the text classifier 135 needs to be trained to extract additional information for classes of commands related to search and time.
[0023] More details about the components of system 100 are now provided. The speech recognizer 115 is responsible for locally transcribing the audio of speech 110 received from a user of system 100. The tokenizer 125 is a preprocessing block in which text is encoded in a process called tokenization to be used in the models. The text classifier 135 is a model trained with the intent of classifying text commands, returning which command should be executed by system 100. The string parser 155 is a trained token classification model that identifies the most relevant part of the input text 130. It is used in internet search commands, returning the relevant part of that text to feed the search engine. (e.g., “search for the latest sports news” returns “latest sports news”, “search for funny videos on the web” returns “funny videos”).Finally, the 165 time parser is used in commands containing time information, to extract months, days, hours, minutes, and seconds. The 165 time parser does not need to be a trained model, but uses a parsing solution with a Regex pattern, followed by a business rule. The 175 commands are returned, along with extra information for execution by the 100 system.
[0024] The speech recognizer 115 is responsible for transcribing the audio and is run in one instance by the Android® SpeechRecognizer with a Google® voice engine. Transcription can also be performed with other tools such as Microsoft® Azure® Speech to Text or OpenAI™'s Whisper System.
[0025] System 100 has minimal dependency on the audio transcription tool used, therefore it allows Petition 870250023302, dated 03 / 25 / 2025, page 24 / 49 7 / 21 seamless integration with other speech recognition solutions, such as OpenAI's Whisper, which includes versions optimized for embedded systems, featuring smaller variants adapted for local execution.
[0026] Preprocessing by the 125 tokenizer (tokenization) divides the text into smaller units and maps them to unique tokens (IDs). As a midpoint between words and characters, subword units retain linguistic meaning (as morphemes) while alleviating vocabulary gaps, even with a relatively small vocabulary. The tokenizer can be selected based on a trade-off between precision and complexity. In one example, a Fast WordPiece tokenizer is used and has a complexity of O(n) (where n is the input length) and is based on the WordPiece tokenizer. A vocabulary of 30,522 tokens can be used.
[0027] Given a set T of texts and a vocabulary V, where V ⊆ T, the vocabulary V contains a list of unique words (or subwords) w where each is associated with an index i. Therefore, a tokenizer represents a text T in a token vector X[ with predefined size d, where X^ is a token composed of the index j of an associated token wj ⊆ T. Tokenization is performed through the following steps: (1) Divide the text into sub-words; (2) Map each subword to its associated token ID. Subwords outside the vocabulary are represented by the token [UNK]; (3) Add the special tokens [CLS] at the beginning and [SEP] at the end of the sequence. (4) Correct the number of tokens by means of truncation and padding, using the special token [PAD].
[0028] The token vector is then converted into a dictionary D composed of three arrays of equal size, following the keys: “input mask (im), “input type ids Petition 870250023302, dated 03 / 25 / 2025, page 25 / 49 8 / 21 (itid) and “input word ids (iwid), to match the BERT standard input. iwid contains the token vector; im is composed of a binary array, indicating with 0 the tokens obtained by padding and with 1 the valid tokens; and itid is fully padded with zeros, since its original use is not suitable for this application.
[0029] FIG. 2 is a block diagram illustrating the parameterization of an example network architecture 200 for the text classifier model 135. The text classifier 135, in an example, is a model with a BERT architecture for intentional classification of text commands. The main objective is to determine which command should be executed by the system. The model architecture includes the following components: (1) Input layer 210: Receives a standard dictionary D, comprising arrays of predefined size d. These arrays are obtained through the tokenization process. (2) BERT 220 Encoder: Applies the BERT encoder, which consists of L transformer encoder blocks, H hidden units (or embedding size) and A attention headers. (3) Output representation: Uses BERT pooled output, representing an embedding vector for all input tokens. Then, a 230 dropout layer with a rate r is applied to mitigate overfitting. (4) The final fully connected layer 240 includes a softmax layer, with the number of neurons c equal to the total number of mapped classes.
[0030] In one example, device 180 is an interactive display, such as a large smart whiteboard with a touch screen and an input mechanism 185 to initiate the reception of a voice command. These whiteboards are commonly used in conference rooms for meetings. Other devices that accept and execute a set of commands Petition 870250023302, dated 03 / 25 / 2025, page 26 / 49 9 / 21 can also use the 100 system. The input mechanism can be an icon or menu selection on the screen or on a remote control device, such as a smartphone, tablet, keyboard or laptop, wirelessly coupled to the whiteboard which, after selection, results in an activation signal to enter speech reception mode 110 by the 100 system.
[0031] The network parameters (d,L,H,A,r) for the text classifier model 135 can be determined through hyperparameter optimization. In one example, a dataset for text classifier 135 comprises short English texts, each representing a single command associated with a specific class. The inputs are tuples where the texts are in string format and the classes are integers, for example, shut down the system (0); start external audio recording (47); please set the system to dark mode (33); please navigate to system settings (27); enter the whiteboard (50); machine, turn on the wireless network (20).
[0032] In one example, two data augmentation techniques can be employed to increase the dataset volume and intraclass variability: (i) using synonyms to refer to the device, including device, monitor, system, machine, and equipment; consequently, any reference to equipment was initially labeled as {device}, to be randomly replaced later by one of the five variations mentioned; (ii) randomly inserting the word please at the beginning or end of randomly selected phrases within each class. It is worth noting that a special class called unrec (unrecognized) was developed to denote phrases that cannot be mapped to one of the pre-established commands, either due to ambiguous meaning or non-existent commands.
[0033] To mitigate model bias in inferring classes with unequal sample sizes, the data can be balanced Petition 870250023302, dated 03 / 25 / 2025, page 27 / 49 10 / 21 so that each class contains an equal number of samples. 180 samples were chosen per class, except for the search and unrec classes, which had 217 and 595 commands, respectively. Consequently, the resulting dataset consisted of 12,332 sentences, mapped into 66 classes.
[0034] Given that the distinction between uppercase and lowercase letters does not affect the model's accuracy in this scenario, and the removal of common empty words such as on, off, up, down, and others — frequently used in the classification of longer texts — would lead to the loss of crucial information, it was decided, as part of a data preprocessing strategy, to simply convert all texts to lowercase and retain all words, including empty words.
[0035] The training parameters for the text classifier model 135 are: Optimizer: AdamW1, Learning rate: 10-5, Weight decay: 0.01, 1st moment decay rate (beta1): 0.9, 2nd moment decay rate (beta2): 0.999, Constant for stability (epsilon): 10-7, Batch size: 16, Loss function: Sparse categorical cross-entropy, Stopping criterion: 5 epochs without loss reduction.
[0036] The TextClassifier model training was conducted in two stages: initially, a higher learning rate of 0.003 was used, with all model weights frozen except the output layer; subsequently, all model weights were unfrozen to facilitate full tuning.
[0037] After training, the model can be converted to the .tflite format to ensure compatibility for mobile deployment. Additionally, post-training quantization can be applied using float16 to reduce model size and decrease inference processing time. The inclusion of a comprehensive standard vocabulary in the BERT solution allows the Petition 870250023302, dated 03 / 25 / 2025, page 28 / 49 11 / 21 text classifier model 135 handle unseen words during training.
[0038] The 155 string parser is a token classification model in a BERT-based example and is designed to identify the most relevant segments of the 130 input text. The 155 string parser shares architectural similarities with the 135 text classifier, with differences only in the encoder output and the model output layer. The BERT output is now a sequence output, represented as a matrix, where each row corresponds to the embedding of each input token. The model's softmax layer, on the other hand, produces a (2, d) format output, corresponding to the size of the input tokens. This output consists of zeros and ones, indicating whether the corresponding input token is part of the relevant text or not, respectively.
[0039] In one example, the following procedure is followed to extract the relevant part of the text: (1) Preprocessing of tokenizer 125 to obtain input dictionary 130 and, consequently, the list of tokens of the key “insert word ids; (2) Feeding the trained string parser model 155 with the input dictionary 130 to obtain the token sorting list.
[0040] A string parser dataset 155 in an example consists of internet search commands, each associated with a segment of interest from this text to serve as input for the search engine, for example, input text search funny videos on the web output funny videos; input text search the latest sports news - output the latest sports news; input text browse example.xyz - output example.xyz.
[0041] For training string analyzer model Petition 870250023302, dated 03 / 25 / 2025, page 29 / 49 12 / 21 of 155 characters, an auxiliary encoding function converts the output text into a sort vector composed of zeros and ones. These values indicate whether each token belongs to the segment of interest based on the input text. The input text is converted to lowercase, and empty words are retained. The resulting dataset comprises 812 sentences in one example.
[0042] FIG. 3 is a block diagram illustrating a network architecture parameterization 300 for the string parser model 155 and includes an input layer 310, a BERT encoder 320, a dropout layer 330, and a fully connected layer 340. The training parameters remain consistent with those of the text classifier 135, except for the stopping criterion, now defined as 8 epochs in a non-loss reduction example. Training can be conducted in two phases, following the same protocol used for text classifier model training 135.
[0043] After training is complete, the model can be converted to .tflite format, followed by posttraining quantization using float16 to reduce the model size.
[0044] The 165 time parser is designed to extract temporal information (months, days, hours, minutes, and seconds) from commands. The 165 time parser adapts to variations in language, recognizing period expressions (e.g., morning, afternoon, evening, night) equivalent to AM / PM, common time expressions (noon, midnight, half past, quarter), and days of the week (including relative references such as tomorrow). Instead of relying on a trained model, the 165 time parser employs a Regex-based parsing approach coupled with specific business rules.
[0045] The Regex module receives a list of words extracted from the input text, followed by mapping each word to one of seven tags using Regex rules (without Petition 870250023302, dated 03 / 25 / 2025, page 30 / 49 13 / 21 distinction between uppercase and lowercase) in the form of business rules, where the labels in an example might include: (0) Other - anything that is not related to time; (1) Time pattern (e.g., 5:00, “11:30); (2) AM period (e.g., “am, “morning); (3) PM Period (e.g., “pm, “afternoon); (4) Relative days (e.g., “Monday, “tomorrow); (5) Absolute days (e.g., “12, “25); (6) Months (e.g., “January, “December).
[0046] Business rules leverage word lists and tags to extract the desired temporal information or flag an error under specific conditions. The function performs the following steps: (1) Initial check: Checks for inconsistencies and out-of-context entries. Flags as invalid if the tag list does not have enough information to characterize a time command (type 0), has multiple unwanted tags (e.g., multiple relative days), or contains a phrase without a time pattern (tag 1). (2) Information extraction: Iterates over lists of labels and words to populate an output structure with values representing month, absolute day, relative day, time of day, hour, and minute. (3) Information validation: Ensures that the extracted information matches the expected time pattern. Flags are invalid if month > 12, absolute day > 31, hour h 12 with day period, hour h 24 or minute h 60.
[0047] FIG. 4 is a flowchart illustrating a speech processing method 400 for commands. The method 400 begins in operation 410 by receiving a voice request to execute a command for a device. Operation 420 obtains a text string corresponding to the voice request. The string Petition 870250023302, dated 03 / 25 / 2025, page 31 / 49 14 / 21 of the text characters are tokenized in operation 430 to generate text string tokens.
[0048] Text string tokens are used as input to a text classifier in operation 440. The text classifier has been trained to identify command classes corresponding to the voice request. The command classes in one example include a search class command, a time class command, and other commands. The text classifier classifies an intent from the text string and includes a softmax output layer with a number of neurons equal to the number of command classes. The text classifier can be trained with an equal number of samples for each class and can be post-trained using quantization to reduce the model size.
[0049] Operation 450 extracts information from text string tokens using a string parser for a search class command. The string parser can be a model trained to find search terms in speech, including commands, and is subsequently trained using quantization to reduce the model size.
[0050] Time information is extracted in operation 460 from text string tokens using a time parser. Time information can be referred to as temporal data related to a time-related command, such as a meeting request at a specific time. The time parser may include a Regex module with rules to map words extracted from text string tokens to temporal information corresponding to the command. The commands, along with the extracted information, are provided to the device in operation 470 for execution.
[0051] In one example, method 400 is incorporated into the device, which can be an interactive screen after receiving the data. Petition 870250023302, dated 03 / 25 / 2025, page 32 / 49 15 / 21 of an activation signal from user input via the interactive screen.
[0052] FIG. 5 is a schematic block diagram of a 500 computer system for performing speech processing for commands and for implementing methods, models, and algorithms according to the example modes. Not all components need to be used in multiple modes.
[0053] An example of a computing device in the form of a computer 500 may include a processing unit 502, memory 503, removable storage 510, and non-removable storage 512. Although the example computing device is illustrated and described as a computer 500, the computing device may be in different forms in different embodiments. For example, the computing device may be a smartphone, a tablet, a smartwatch, a smart storage device (SSD), or another computing device including the same or similar elements, as illustrated and described in relation to FIG. 5. Devices such as smartphones, tablets, and smartwatches are generally collectively referred to as mobile devices or user equipment.
[0054] Although the various data storage elements are illustrated as part of the 500 computer, storage may also or alternatively include cloud-based storage accessible via a network such as the Internet or server-based storage. Note also that an SSD may include a processor on which the analyzer can run, allowing the transfer of analyzed and filtered data via I / O channels between the SSD and main memory.
[0055] Memory 503 may include volatile memory 514 and non-volatile memory 508. Computer 500 may include - or have access to a computing environment that includes - a variety of Petition 870250023302, dated 03 / 25 / 2025, pp. 33 / 49 16 / 21 computer-readable media, such as volatile memory 514 and non-volatile memory 508, removable storage 510 and non-removable storage 512. Computer storage includes random access memory (RAM), read-only memory (ROM), programmable and erasable read-only memory (EPROM) or electrically programmable and erasable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVDs) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other means capable of storing computer-readable instructions.
[0056] Computer 500 may include or have access to a computing environment that includes input interface 506, output interface 504, and a communication interface 516. Output interface 504 may include a display device, such as a touch screen, which may also serve as an input device. Input interface 506 may include one or more touch screens, touchpads, mice, keyboards, cameras, one or more device-specific buttons, one or more sensors integrated or coupled via wired or wireless data connections to computer 500, and other input devices. The computer may operate in a network environment using a communication connection to connect to one or more remote computers, such as database servers. The remote computer may include a personal computer (PC), server, router, network PC, peer device, or other common data stream network switch or similar.The communication connection may include a Local Area Network (LAN), a Wide Area Network (WAN), cellular, Wi-Fi, Bluetooth, or other networks. Depending on the embodiment, the various components of the 500 computer are connected to a 520 system bus. Petition 870250023302, dated 03 / 25 / 2025, pp. 34 / 49 17 / 21
[0057] Computer-readable instructions stored on a computer-readable medium are executable by the processing unit 502 of the computer 500 as a program 518. The program 518 in some embodiments comprises software to implement one or more methods described herein. A hard disk, a CD-ROM, and RAM are some examples of items that include a non-transient computer-readable medium as a storage device. The terms computer-readable medium, machine-readable medium, and storage device do not include signals or carrier waves insofar as signals and carrier waves are considered too transient. Storage may also include network storage, such as a storage area network (SAN). The computer program 518 can be used to cause the processing unit 502 to execute one or more methods or algorithms described herein. Examples:
[0058] A computer-implemented method for deriving voice commands involves receiving a voice request to execute a command to a device. A text string is obtained corresponding to the voice request. The text string is tokenized to generate text string tokens. The text string tokens are fed into a trained text classifier to identify command classes corresponding to the voice request, the command classes including a search class command, a timing class command, and other commands. Information is extracted from the text string tokens via a string parser for a search class command. Timing information is extracted from the text string tokens via a timing parser, and the resulting commands and extracted information are provided to the device for Petition 870250023302, dated 03 / 25 / 2025, pp. 35 / 49 18 / 21 execution.
[0059] 2. The method from example 1 where the method is incorporated into the device.
[0060] 3. The method of either of examples 1 to 2 where the device includes an interactive screen.
[0061] 4. The method in example 3, where the method is executed after receiving an activation signal from user input via the interactive screen.
[0062] 5. The method of any of examples 1 to 4 in which the text classifier includes a softmax output layer with a number of neurons equal to a number of command classes.
[0063] 6. The method of any of examples 1 to 5 in which the text classifier classifies an intent from the text string.
[0064] 7. The method in example 6 where the text classifier is trained with an equal number of samples for each class and is subsequently trained using quantization to reduce the model size.
[0065] 8. The method of any of examples 1 to 7 in which the string parser includes a model trained to find search terms from speech, including commands, and is subsequently trained using quantization to reduce the model size.
[0066] 9. The method of any of examples 1 to 8 where the time parser includes a Regex module with rules for mapping words extracted from text string tokens to temporal information corresponding to the command.
[0067] 10. A machine-readable storage device has instructions for execution by a machine's processor to cause the processor to perform operations to accomplish any of the methods in examples 1 through 9.
[0068] 11. A device includes a processor and a memory device coupled to the processor and having a Petition 870250023302, dated 03 / 25 / 2025, pp. 36 / 49 19 / 21 program stored in it for execution by the processor to perform operations to carry out any of the methods from examples 1 to 9.
[0069] The functions or algorithms described herein may be implemented in software in one embodiment. The software may consist of computer-executable instructions stored on computer-readable media or on a computer-readable storage device, such as one or more non-transient memories or other type of hardware-based storage device, local or networked. Furthermore, such functions correspond to modules, which may be software, hardware, firmware, or any combination thereof. Several functions may be implemented in one or more modules, as desired, and the embodiments described are merely examples. The software may run on a digital signal processor, ASIC, microprocessor, or other type of processor operating in a computer system, such as a personal computer, server, or other computer system, transforming such computer system into a specifically programmed machine.
[0070] Functionality can be configured to perform an operation using, for example, software, hardware, firmware, or similar. For example, the phrase "configured to" can refer to a logic circuit structure of a hardware element that must implement the associated functionality. The phrase "configured to" can also refer to a logic circuit structure of a hardware element that must implement the coding design of the associated firmware or software functionality. The term "module" refers to a structural element that can be implemented using any suitable hardware (e.g., a processor, among others), software (e.g., an application, among others), firmware, or any combination of hardware, software, and firmware. The term "logic" encompasses any functionality to perform a task. Petition 870250023302, dated 03 / 25 / 2025, pp. 37 / 49 20 / 21 example, each operation illustrated in the flowcharts corresponds to the logic for performing that operation. An operation can be performed using software, hardware, firmware, or similar. The terms component, "system," and similar terms can refer to entities related to computers, running hardware and software, firmware, or a combination thereof. A component can be a process running on a processor, an object, an executable, a program, a function, a subroutine, a computer, or a combination of software and hardware. The term "processor" can refer to a hardware component, such as a processing unit of a computer system.
[0071] Furthermore, the claimed matter may be implemented as a method, apparatus, or article of manufacture using standard programming and engineering techniques to produce software, firmware, hardware, or any combination thereof to control a computing device to implement the disclosed matter. The term “article of manufacture,” as used herein, is intended to encompass a computer program accessible from any computer-readable storage device or media. Computer-readable storage media may include, but is not limited to, magnetic storage devices, for example, hard disk, floppy disk, magnetic tape, optical disc, compact disc (CD), digital versatile disc (DVD), smart cards, flash memory devices, among others.In contrast, computer-readable media, i.e., non-storage media, may additionally include communication media, such as transmission media for wireless signals and the like.
[0072] Although some modalities have been described in detail above, other modifications are possible. For example, the logical flows represented in the figures do not require the specific order shown, or sequential order, to achieve the Petition 870250023302, dated 03 / 25 / 2025, pp. 38 / 49 21 / 21 desirable results. Other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to or removed from the described systems. Other embodiments may be within the scope of the following claims. Petition 870250023302, dated 03 / 25 / 2025, pp. 39 / 49
Claims
1 / 5 CLAIMS 1. A computer-implemented method, characterized in that it comprises: receiving a voice request to execute a command to a device; obtaining a text string corresponding to the voice request; tokenizing the text string to generate text string tokens; inserting the text string tokens into a trained text classifier to identify command classes corresponding to the voice request, the command classes including a search class command, a time class command, and other commands; extracting information from the text string tokens by means of a string parser for a search class command; extracting time information from the text string tokens by means of a time parser; and providing the resulting commands and extracted information to the device for execution.
2. Method according to claim 1, characterized in that the method is incorporated into the device.
3. Method according to claim 1, characterized in that the device comprises an interactive screen.
4. Method, according to claim 3, characterized in that the method is performed after receiving an activation signal from user input via the interactive screen.
5. Method, according to claim 1, characterized in that the text classifier includes a softmax output layer with a number of neurons equal to a number of command classes. Petition 870250023302, dated 03 / 25 / 2025, pp. 40 / 49 2 / 5 6. Method, according to claim 1, characterized in that the text classifier classifies an intent of the text string.
7. Method, according to claim 6, characterized in that the text classifier is trained with an equal number of samples for each class and is post-trained using quantization to reduce the model size.
8. Method, according to claim 1, characterized in that the string parser comprises a model trained to find search terms from speech, including commands, and is post-trained using quantization to reduce the model size.
9. Method, according to claim 1, characterized in that the time parser comprises a Regex module with rules for mapping words extracted from text string tokens to temporal information corresponding to the command.
10. Machine-readable storage device, characterized in that it has instructions for execution by a machine processor to cause the processor to perform operations to carry out a method, the operations comprising: receiving a voice request to execute a command to a device; obtaining a text string corresponding to the voice request; tokenizing the text string to generate text string tokens; inserting the text string tokens into a trained text classifier to identify command classes corresponding to the voice request, the command classes including a search class command, a time class command, and other commands; Petition 870250023302, dated 03 / 25 / 2025, p.41 / 49 3 / 5 extract information from text string tokens using a string parser for a search class command; extract time information from text string tokens using a time parser; and provide resulting commands and extracted information to the device for execution.
11. Device according to claim 10, characterized in that the method is incorporated in the device and in that the device comprises an interactive screen.
12. Device according to claim 11, characterized in that the method is performed after receiving an activation signal from user input via the interactive screen.
13. Device according to claim 9, characterized in that the text classifier includes a softmax output layer with a number of neurons equal to a number of command classes and classifies an intent from the text string.
14. Device according to claim 13, characterized in that the text classifier is trained with an equal number of samples for each class and is post-trained using quantization to reduce the model size.
15. Device according to claim 13, characterized in that the string parser comprises a model trained to find search terms from speech, including commands, and is post-trained using quantization to reduce the model size.
16. Device, according to claim 13, characterized in that the time analyzer comprises a Regex module with rules for mapping words extracted from the text string tokens to temporal information corresponding to the command.
17. Device, characterized in that it comprises: a processor; and a memory device coupled to the processor and having a program stored thereon for execution by the processor to perform operations comprising: receiving a voice request to execute a command to a device; obtaining a text string corresponding to the voice request; tokenizing the text string to generate text string tokens; inserting the text string tokens into a trained text classifier to identify command classes corresponding to the voice request, the command classes including a search class command, a time class command, and other commands; extracting information from the text string tokens by means of a string parser for a search class command;Extract timing information from text string tokens using a time parser; and provide the resulting commands and extracted information to the device for execution.
18. Device according to claim 17, characterized in that the device comprises an interactive screen.
19. Device, according to claim 17, characterized in that the text classifier includes a softmax output layer with a number of neurons equal to a number of command classes and classifies an intent of the text string, and wherein the text classifier is trained with an equal number of samples for each class and is post-trained using quantization to reduce the model size.
20. Device according to claim 19, characterized in that the string parser comprises a model trained to find search terms from speech, including commands, and is post-trained using quantization to reduce the model size, and wherein the timing parser comprises a Regex module with rules for mapping words extracted from text string tokens to temporal information corresponding to the command. Petition 870250023302, dated 03 / 25 / 2025, pp. 44 / 49