Medium range speech communication systems, methods, and devices

TWI937651BActive Publication Date: 2026-09-01IOAIRE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
TW113149875
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-20
Publication Date
2026-09-01
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

Communication within large and complex facilities, such as warehouses and medical centers, is challenging due to the limitations of traditional cellular and Wi-Fi networks, which have limited range and difficulty penetrating walls, requiring numerous devices for coverage and high power consumption.

Method used

A medium-range voice communication system utilizing 802.11ah (HaLow Wi-Fi) networks with human-machine interface devices and computing devices that translate voice commands into text and vice versa, enabling robust communication through walls and reducing infrastructure and power consumption.

Benefits of technology

Provides efficient, hands-free communication with extended range and lower power consumption, allowing seamless voice communication within large facilities without the need for extensive cellular infrastructure, and enabling faster, less distracting interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001908629_001
    Figure TWG2TB001908629_001
  • Figure TWG2TB001908629_002
    Figure TWG2TB001908629_002
  • Figure TWG2TB001908629_003
    Figure TWG2TB001908629_003
Patent Text Reader

Abstract

This document provides systems, methods, and apparatus for providing mid-range communication between portable devices. One method includes: receiving a query from a portable device within a mid-range communication area of ​​a mid-range connection device; determining whether the query is in voice format; determining whether the query is in text format readable by a computing device; determining an action to be taken on the computing device and taking the determined action, the action producing a result of the query; and returning the result of the action to the portable device from which the query was received.
Need to check novelty before this filing date? Find Prior Art

Description

Medium-range voice communication system, method and device The present invention relates to a medium-range voice communication system, method and device. Facilities can include complexes such as warehouses, medical centers, universities, and the like. These facilities are often large and complex (e.g., large buildings, multiple floors, facilities with multiple buildings), and thus, communication between user devices can be difficult. Typically, communication methods are cellular networks used with mobile devices and Wi-Fi networks primarily for computer-to-computer communication. Some embodiments of the present invention provide a medium-range voice communication system. The medium-range voice communication system includes: an 802.11ah medium-range network having a plurality of human-machine interface devices and a plurality of computing devices communicating through a plurality of 802.11ah network nodes distributed within a facility; at least one of the human-machine interface devices is a first audio user interface device, the first audio user interface device receives a voice command from a user including a command for a first one of the plurality of computing devices, creates a voice command data file, and transmits the voice command data file to the first computing device via one of the plurality of 802.11ah network nodes, the first computing device having a processor and a memory, the commands stored in the memory being operable to execute commands on the first computing device. The processor executes the following: receiving the voice command data file; translating the voice command data file into a text command data file; searching a database for an appropriate response to the command; creating a text response data file; translating the text response data file into a voice response data file; and transmitting the voice response data file from the first computing device via one of the 802.11ah network nodes; and the first audio user interface device has a processor and a memory, and instructions stored in the memory can be executed on the processor to: receive the voice response data file; and play the voice response data file to the user via a speaker on the first audio user interface device. Some embodiments of the present invention provide a portable mid-range communication device. The portable mid-range communication device includes a processor and memory, the memory having instructions executable on the processor to: receive a query from a portable device within a mid-range communication area of ​​a mid-range connection device; determine whether the query is in voice format; determine whether the query is in text format readable by a computing device; determine an action to be taken on the computing device and take the determined action, the action producing a result of the query; and return the result of the action to the portable device from which the query was received. Embodiments of the present invention relate to systems, methods, and devices for providing medium-range voice communication between portable devices. For example, a method includes: receiving a query from a portable device within a medium-range communication area of ​​a medium-range connection device; determining whether the query is in voice format; determining whether the query is in text format readable by a computing device; determining an action to be taken on the computing device and taking the determined action, the action producing a result of the query; and returning the result of the action to the portable device from which the query was received. Particularly helpful embodiments of the present invention are those utilizing HaLow Wi-Fi (also known as 802.11ah). Wi-Fi HaLow is unique in that it uses low-power connectivity, making it suitable for small, portable devices. It also has a longer range than many other Internet of Things (IoT) technology options, requiring less infrastructure. It also provides a more robust connection in challenging environments due to its unique ability to penetrate walls and other barriers within a building or facility. Wi-Fi HaLow uses a sub-gigahertz (S1G) radio operating in a frequency range between 750 MHz and 1 GHz. Its coverage range is one mile from its source device (e.g., a HaLow network access point). The use of a HaLow network can provide communication coverage to a large facility without the substantial infrastructure of a cellular network (e.g., towers and transmitters), and has the ability to enable the functions discussed herein that cannot be provided by a typical Wi-Fi network, particularly in facilities where Wi-Fi signals need to propagate between buildings (e.g., through the exterior and / or interior walls of multiple buildings). Traditional Wi-Fi solutions have a range of only about 75 feet from their source and moderate wall penetration capabilities. This means that quite a few traditional Wi-Fi network devices are needed to provide wireless coverage to the same area (if possible). HaLow Wi-Fi can also operate at a power consumption significantly lower than that of traditional Wi-Fi devices. For example, a coin cell battery-powered device can operate for months or even years. Many devices have a power consumption of less than 10 microwatts. HaLow Wi-Fi uses a unique algorithm to provide longer sleep times between beacon responses (checking in with other network devices), resulting in significantly longer device battery life. Furthermore, HaLow Wi-Fi has power-saving capabilities that allow for ultra-low power consumption that is unattainable with traditional Wi-Fi devices. In the following detailed description, reference is made to the accompanying drawings which form a part hereof. The drawings show by way of illustration how one or more embodiments of the invention may be practiced. These embodiments are described in sufficient detail to enable one of ordinary skill in the art to practice one or more embodiments of the present invention. It is understood that other embodiments may be utilized and process, electrical and / or structural changes may be made without departing from the scope of the present invention. As will be appreciated, elements shown in the various embodiments herein may be added, exchanged, combined, and / or eliminated to provide several additional embodiments of the present invention. The proportions and relative scales of the elements provided in the drawings are intended to illustrate embodiments of the present invention and should not be considered limiting. As used herein, "a," "an," or "several" things may refer to one or more of such things, and "plurality" things may refer to more than one of such things. For example, "several buildings" may refer to one or more buildings, and "plurality of buildings" may refer to more than one building. Figure 1 illustrates a medium-range voice communication system at a facility in accordance with one or more embodiments of the present invention. As used herein, medium-range communication is between 350 feet and 1.5 miles. 1 , a medium-range communication system 100 is shown having a facility 101 with several buildings 108. Within the buildings 108 (with roofs removed for ease of viewing) are shelving units 110, each of which has a different type of product 112 on it. These different types of products are represented by blocks of different sizes and shapes, but these are not the only distinguishing characteristics, as will be discussed below. System 100 includes communication devices 104 and a system control computing device 102. In some embodiments, the system control computing device may also include communication device functionality, allowing it to serve a dual purpose. The communication devices shown are 802.11ah devices, allowing for long range (a range of 1 mile is shown between each device 104 and its neighboring device(s)). 802.11ah devices can also penetrate at least one wall of a building 108. In the illustrated embodiment, the device's signal can penetrate two building walls, allowing the left and right devices to communicate with the center device, which allows information to be transmitted between the left and right devices. In some embodiments where the left and right devices are within communication range of each other, they can be designed to penetrate the four walls between them and any other obstructions in the signal transmission path between the two communicating devices (such as shelves and products). Figure 1 also shows two people with portable devices 106. As discussed herein, one person may want to locate or contact another person to initiate a voice call. In some embodiments, one person may want to request an inquiry about a product 112 at facility 101 via their device 106. Device 106 may be referred to as an audio user device, as used herein. These processes will be discussed in more detail below. An example embodiment of a mid-range voice communication system is provided below. In this example, the communication system is an 802.11ah mid-range network having a plurality of human-machine interface devices (e.g., portable devices capable of audio communication via a Wi-Fi network, such as walkie-talkies and cell phones) and a plurality of computing devices (e.g., computer servers) communicating through a plurality of 802.11ah network nodes (e.g., access points) distributed throughout a facility. At least one of the human-machine interface devices is a first audio user interface device that receives a voice command from a user for a first one of the plurality of computing devices, creates a voice command data file (a voice format query), and transmits the voice command data file to the first computing device via one of the 802.11ah nodes. The first computing device has a processor and a memory, and the commands are stored in the memory. The instructions are executable on the processor to: receive a voice command data file; translate the voice command data file into a text command data file (text format query); search a database (e.g., in data storage 361) for an appropriate response to the command; upon determining an appropriate response, create a text response data file; translate the text response data file into a voice response data file; and transmit the voice response data file from the first computing device via one of the 802.11ah network nodes. The first audio user device has a processor and a memory, and instructions are stored in the memory. The instructions are executable on the processor to receive a voice response data file and play the voice response data file to the user via a speaker on the first audio user device. Through the above example embodiments, a user can ask a query by speaking into a microphone on their portable device and receive an audio response from the system through a speaker on the device. Such a system can provide substantial benefits, such as being much faster than typing a query, being less distracting to the user, and being essentially hands-free. FIG2 illustrates a computing device for medium-range voice communication according to one or more embodiments of the present invention. For example, a computing device 220 as shown in FIG2 can be used by a user as a portable device in a medium-range communication system, such as shown at 106 in FIG1 . 2 , the device 206 includes a computing device 220, an audio input 234, a wireless transmitter / receiver 236, an audio output device 238, and a display 239. The computing device 220 includes a processor 222 for executing executable instructions 226 stored in a memory 224. The executable instructions may, for example, initiate a voice call, record an audio query (via the audio input device 234, such as a microphone) to create a voice format file, send the voice format file to the system control device, receive a voice format file containing the results of the query, play the voice format file to the user (via the audio output device 238, such as a speaker), or conduct a voice call, among other functions. The memory also contains data 228 that can be used during the above functions. For example, the data can include voice format files, information about the user of this particular portable device, information about this particular portable device, and other useful information. Examples of this information include the user's name, an identifier (if an identifier has been assigned to the user), the user's portable device identifier, and the last location of the user's portable device. The computing device 220 also includes a network interface 230. For example, the network interface 230 can be used to receive updated executable instructions and / or data and can be part of an 802.11ah network or a different type of network. Input / output interface 232 connects components 234, 236, 238, and 239. Audio input can receive audio files containing queries to be processed from a portable device. Transmitter / receiver 236 can facilitate transferring audio files to and / or from the portable device. For example, the user input device can be a keyboard or mouse. The display can be a visual screen that displays information to a user viewing the display. FIG3 is a functional diagram of a computing device for medium-range voice communication according to one or more embodiments of the present invention. At the top of FIG3 , several actions 344 , 346 , 348 , and 349 to be taken by system 340 are shown. To initiate a request for an action to be taken, a user may make a query at 342 by speaking into their portable user device. Each of these actions utilizes different elements stored in memory 360 . Memory 360 can be volatile or non-volatile memory. Memory 360 can also be removable (e.g., portable) memory or non-removable (e.g., internal) memory. For example, memory 360 can be random access memory (RAM) (e.g., dynamic random access memory (DRAM) and / or phase change random access memory (PCRAM)), read-only memory (ROM) (e.g., electrically erasable programmable read-only memory (EEPROM) and / or compact disc read-only memory (CD-ROM)), flash memory, a laser disk, a digital versatile disk (DVD) or other optical disk storage, and / or a magnetic medium such as a cassette, magnetic tape, or disk, among other types of memory. Within memory 360, executable instructions at 370 may, for example, determine whether a query is in a text or voice format, initiate a voice-to-text module or a text-to-speech module, determine an action to be taken based on the content of the query, perform the action, initiate a voice call connection, conduct a voice call, return the results of the action to the device from which the query was received, and other functions. The memory also includes data storage 361 that can be used during the above functions. For example, the data can include voice format files, text format files, information about people who have portable devices at the facility, information about portable devices at the facility, information about the characteristics of products in the facility's inventory, and other useful information. Furthermore, memory 360 and processor 350 are located within a computing device within the medium-range voice communication system. In the illustrated system shown in FIG1 , this computing device may be device 102. Although depicted as being located within computing device 102, embodiments of the present invention are not limited thereto. For example, memory 360 may also be located within another computing resource (e.g., enabling computer-readable instructions to be downloaded via the Internet or another wired or wireless connection). For example, for a request to locate a person 344 at a facility, the processor 350 uses the query module 369 to receive the query file and determine whether it needs to be converted from speech to text, and if so, uses the speech-to-text module 362 executed by the processor to do so. The data storage 361 may have information 365 about the location of the person in the facility. For example, the data storage may include the person's name, an identifier (if an identifier has been assigned to the person), the user's portable device identifier, and the user's portable device's last location. To place a call, a user initiates a request query 346 by speaking into their portable device (e.g., 106 in FIG. 1 ). For example, the user may say, "I want to speak to Jeff." The system may process the query and say, "There are three Jeffs. Do you mean Jeff Stier, Jeff Odens, or Jeff Cameron?" The user may then respond by saying, "Jeff Stier." The processor 350 may then execute instructions for the call connection module 364. The call connection module 364 locates Jeff Stier's portable device and, through it, alerts Mr. Stier that someone wants to call. Once connected, executable instructions 370 for providing voice to the voice communication module 368 can be executed and a call can be made. The query module 269 can also be used to convert query results from text to voice via the text-to-speech module 366. In some embodiments, the voice response data file returned to the querying portable device (the first audio user device) may include a confirmation that the second audio user device is the correct device for establishing a voice communication session. For example, this may be achieved by including the name of a user of the second audio user device. In various embodiments, the voice response data file includes a request for the user to provide a voice or text confirmation that the user's name is correct. This can be confirmed by the user simply saying "yes". In some embodiments, the system further includes a second audio user device, and the query is to locate the second audio user device on the 802.11ah network. In this embodiment, the voice response data file may, for example, include a description of the second audio user device's location within a facility. For example, the system may return a result indicating that the second audio user device is located outside the southwest corner of Building B. This information may be derived through GPS tracking, signal strength triangulation based on the location of access points, or any other suitable positioning procedures available to the system. System 340 also allows users to request inquiries about items tracked at a facility. For example, at 348, a user may request inventory characteristics data via their user portable device. The request may be processed by query module 369 with the help of speech-to-text module 362 and item information data 363 stored in data storage 361. For example, data such as quantity inventory data may include an item type, item identifier, item name, the number of items at the facility, item model identifier, part identifier, item location, a description of what the item is, a description of what the item does, a description of the item's location within the facility, a set of physical dimensions for the item, an item weight, an item location history, and other useful information. The results of the query can then be converted from text to speech by the text-to-speech module 366 and then sent back to the portable user device, where the results can be audibly played through a speaker on the portable user device. For example, the user can speak into the portable user device and say, "How many flange parts are available?" The system can access the data described above, and the portable user device can say, "There are 4 flange parts available." The user can then follow up and say "Where are they located?" and the system can audibly respond "Two flanges are located on shelf 3 of shelving unit #1, and two flanges are located on shelf 5 of shelving unit #8. The closest location to you is unit #8. It's two rows to the right against the wall." Another example of how the system can be used is in the automotive industry. For example, an item could be the number, make, model, year, and / or type (e.g., SUVs, sedans, etc.) of vehicles in a fleet for sale, resale, maintenance, and / or new vehicle preparation. Inventory characteristics could include equipment listed on window stickers, a new or resale vehicle defect list (with items that need to be completed before the vehicle is delivered to the buyer), maintenance items that need to be addressed by a repair technician, and / or the location of the vehicle at the facility. The action to be taken is to provide this information to a user and / or update this information. In some embodiments, instead of a voice file, the system can send another type of media file (e.g., an image or video file) in response to a query. For example, a video file can provide updated information to the user. For example, a video showing a package being delivered to a loading dock or an image demonstrating that an action has occurred are two situations where an image or video would be helpful. In various embodiments, the system may allow a request for more extensive information, such as an item's entire chain of custody. This can be helpful if a buyer wants to audit an item's path. This can be helpful, for example, if a user wants to identify who damaged an item in transit. Another useful query is for updating an item's inventory. Here, a query is presented at 349 in which the user speaks into their portable device by saying, "I took a flange from shelving unit #8." The system will answer the query by updating the inventory in data storage 361 and audibly replying, "The system has updated the inventory," via the user's portable device. Inventory can also be added. For example, a query is presented at 349 where the user speaks into their portable device by saying, "I placed three flanges on shelf #2 at shelf unit #12." The system will answer the query by updating the inventory in data storage 361 and audibly replying, "The system has updated the inventory," via the user's portable device. It should be noted that in all of the above examples, each spoken query is converted to text, and each text reply is converted to speech. This allows for a fluent audible discussion with the system 340 that is so fluent that the user is unaware that the other party to the communication is non-human. 4 is a flow chart of a method for mid-range voice communication according to one or more embodiments of the present invention. The process 480 begins at 481 by receiving a query from a portable device within a mid-range communication area of ​​a mid-range connection device. At 482, the process determines whether the query is in spoken format. If so, at 483, executable instructions executed by the processor initiate speech-to-text conversion via the speech-to-text module. If the query is not in spoken format, at 484, executable instructions are executed to determine whether the query is in a text format readable by the computing device. If the format is neither spoken nor readable text, at 485, the query is transferred to another processing module (if any processing module is used) to determine whether the query can be processed in another manner. If the query is in a computing device readable text format (either directly from the query itself or from conversion via a speech-to-text converter), then at 486 the processor reads the query and executes instructions to determine an action to be taken on the computing device based on the information provided in the text query. At 487, 488, 489, and 490, the determined action is taken to produce a result for the query. At 493, the result of the action is returned to the portable device from which the query was received. In some embodiments, at 491, executable instructions are executed to determine whether the result of the action needs to be converted. If it does need to be converted, a text-to-speech module is initiated at 492, and the result is converted into a voice file or data. Although specific embodiments have been illustrated and described herein, those of ordinary skill in the art will appreciate that any arrangement calculated to achieve the same technique may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations of various embodiments of the present invention. It should be understood that the above description has been made in an illustrative manner and not a restrictive manner. After reading the above description, those skilled in the art will understand the combination of the above embodiments and other embodiments not explicitly described herein. The scope of the various embodiments of the present invention includes any other applications using the above structures and methods. Therefore, the scope of the various embodiments of the present invention should be determined with reference to the appended claims and the full scope of equivalents to which such claims are entitled. In the foregoing detailed description, various features are grouped together in the example embodiments depicted in the figures for the purpose of simplifying the invention. This method of disclosure should not be interpreted as reflecting an intention that embodiments of the invention require more features than expressly recited in each claim. Instead, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the detailed description, with each claim standing on its own as a separate embodiment. 100: Medium-range communication system 101: Facility 102: System control computing device 104: Communication device 106: Portable device 108: Building 110: Shelving unit 112: Product 206: Equipment 220: Computing device 222: Processor 224: Memory 226: Executable instructions 228: Data 230: Network interface 232: Input / output interface 234: Audio input / audio input device / component 236: Wireless transmitter / receiver / component 238: Audio Output device / component 239: Display / component 340: System 342: Action 344: Action 346: Action 348: Action 349: Action 350: Processor 360: Memory 361: Data storage 362: Voice-to-text module 363: Item information 364: Call connection module 365: Information 366: Text-to-speech module 368: Voice-to-voice communication module 369: Query module 370: Executable instructions 480: Process 481: Receive from A query of a portable device within a medium-range communication area of ​​a medium-range connection device 482: The process determines whether the query is in voice format 483: The executable instructions executed by the processor initiate voice-to-text conversion via the voice-to-text module 484: The executable instructions are executed to determine whether the query is in a text format readable by a computing device 485: The query is transferred to other processing modules (if any processing modules are used) to see if the query can be processed in another way 486: The processor reads the query and executes instructions to determine an action to be taken on the computing device based on the information provided in the text query 487: The determined action is taken to produce a result of the query 488: The determined action is taken to produce a result of the query 489: The determined action is taken to produce a result of the query 490: The determined action is taken to produce a result of the query 491: The executable instructions are executed to determine whether the result of the action needs to be converted 492: The text-to-speech module is initiated 493: The result of the action is returned to the portable device from which the query was received FIG. 1 illustrates a mid-range voice communication system at a facility according to one or more embodiments of the present invention. FIG. 2 illustrates a computing device for medium-range voice communication according to one or more embodiments of the present invention. FIG3 is a functional diagram of a computing device for medium-range voice communication according to one or more embodiments of the present invention. FIG4 illustrates a method flow for medium-range voice communication according to one or more embodiments of the present invention. 100: Medium-range communication system 101: Facilities 102: System control computing device 104: Communication device 106: Portable Device 108: Buildings 110: Shelving Unit 112: Products

Claims

1. A medium-range voice communication system, comprising: An 802.11ah midrange network comprising a plurality of 802.11ah network nodes distributed across a plurality of buildings within a facility; a plurality of computing devices within the buildings of the facility, each computing device including a separate processor and a separate memory; and a plurality of portable human interface devices within the buildings of the facility, each including a separate microphone and a separate speaker, and each of the portable human interface devices being associated with a separate user of the facility, wherein each of the portable human interface devices is configured to: receive a voice command from the separate user via the separate microphone; create a voice command data file based on the voice command; transmit the voice command data file to the separate computing device via the 802.11ah midrange network; wherein instructions stored in the separate memory can be executed on the separate processor to: receive the voice command data file; The voice command data file is translated into a text command data file; an appropriate response to the voice command is searched in a database; a text response data file is created; the text response data file is translated into a voice response data file; and the voice response data file is transmitted via the 802.11ah midrange network; and each of the portable human-machine interface devices is further configured to: receive the voice response data file; and play the voice response data file to the respective user via the respective speaker.

2. The system of claim 1, wherein the voice command locates a different user of the facility on the 802.11ah network; and wherein the command is executed to connect the portable human-machine interface device associated with the respective user for voice communication of the portable human-machine interface device associated with the different user of the facility.

3. The system as requested in item 2, wherein the voice response data file includes confirmation that the portable human-machine interface device of the different user associated with the facility is a correct device for the voice communication.

4. The system as requested in item 3, wherein the confirmation that the portable human-machine interface device associated with the different users of the facility is the correct device for the voice communication includes the name of one of the different users.

5. The system as described in request item 4, wherein the voice response data file contains a request from the different user to confirm, either by voice or text, that the name of the different user is correct.

6. A system as requested in any of items 1 to 4, wherein the voice command locates a different user of the facility on the 802.11ah mid-range network; wherein the voice response data file contains a description of the location of a portable human-machine interface device associated with the different user within the facility.

7. A portable human-machine interface mid-range communication device, comprising: A processor and memory having instructions executable on the processor to: receive a query from a portable human-machine interface device within a mid-range communication area of ​​a computing device in a plurality of buildings of a facility, wherein the portable human-machine interface device is associated with a particular user of the facility; determine whether the query is in voice format; determine whether the query is in text format readable by a computing device; determine an action to be taken on the computing device and take the determined action, the action producing a result of the query; and return the result of the action to the portable human-machine interface device from which the query was received.

8. The apparatus of request item 7, wherein when the query is determined to be in a text format readable by the computing device, a query module is initiated to determine the action to be taken.

9. The device of request item 7, wherein if the result of the query is in text format, then a text-to-speech module is initiated to convert the text-format result of the query into a speech-format result of the query.

10. The apparatus of any one of requests 7 to 9, wherein the instructions further include determining whether the result of the query needs to be transformed before returning the query to the portable human-machine interface device from which the query was received.

11. The device of any of the requests 7 to 9, wherein the determined action is a request to connect with a different user within the facility to make a voice call.

12. The device that requests any one of items 7 to 9, wherein the determined action is a request for inventory characteristic data.

Citation Information

Patent Citations

  • Method, apparatus and computer program product for providing adaptive gesture analysis

    TW201019239A

  • Technologies for conversational interfaces for system control

    US20190391541A1

  • Systems and methods for configuring and using an audio transcript correction machine learning model

    US20230360652A1

  • System and method for locational image processing

    US20230363557A1