Vehicle-mounted toll terminal with local AI computing power and voice interaction and implementation method
Patent Information
- Application Number
- CN202611004996.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-08-18
AI Technical Summary
(1)依赖手动操作,乘客需手动刷卡/扫码,老年、残障等特殊群体操作不便;查询余额、换乘信息等功能需通过触屏点击,步骤繁琐;
(1)本发明构建了本地AI算力架构,摆脱了对云端的依赖,实现了语音指令毫秒级处理;语音唤醒-识别-反馈全流程延迟≤1.5s,噪音环境下仍稳定响应,适配车载网络波动场景;
Smart Images

Figure CN122598321A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle-mounted toll collection terminal technology, and in particular to a vehicle-mounted toll collection terminal and its implementation method that has local AI computing power and voice interaction function. Background Technology
[0002] With societal development, the number of private cars is increasing, placing a significant burden on the environment. The government is strongly advocating for green travel. As one of the main modes of green transportation, public buses use card or QR code payments, but these methods have significant limitations in terms of interaction: (1) It relies on manual operation. Passengers need to manually swipe their cards / scan codes, which is inconvenient for the elderly, disabled and other special groups. Functions such as checking balance and transfer information need to be accessed through the touch screen, which is cumbersome. (2) Without local AI computing power support, although some devices support basic voice response, they rely on cloud computing power to process commands. Network fluctuations in the vehicle environment can easily lead to response delay (>3s) or failure. (3) Poor adaptability of voice interaction, unable to adapt to dialects (such as Cantonese and Sichuan dialect) and noisy environments (engine noise, passenger noise), low recognition accuracy (<80%), and lack of user voice habit learning ability; (4) Insufficient integration of functions: Voice commands and core charging functions (such as QR code triggering and abnormal transaction feedback) are not deeply integrated, which can easily lead to problems such as "voice command execution and transaction not being synchronized".
[0003] Therefore, it is necessary to design an in-vehicle toll terminal solution that integrates local AI computing power, supports highly adaptive voice interaction, and is precisely linked with the toll collection function to solve the above pain points. Summary of the Invention
[0004] In view of the above-mentioned defects or deficiencies in the existing technology, it is desirable to provide an in-vehicle toll terminal and its implementation method with local AI computing power and voice interaction function, build a local AI computing power architecture, get rid of dependence on the cloud, achieve millisecond-level processing of voice commands, and provide an excellent user experience.
[0005] The present invention provides a vehicle-mounted toll terminal with local AI computing power and voice interaction function, including a hardware module for communication connection and an AI voice interaction module; The hardware module includes: The main control and storage unit is used for data processing and storing local language models and user interaction logs; The local AI computing unit is electrically connected to the main control and storage unit and is used for local speech recognition, instruction understanding and user model learning. A multimodal microphone array, electrically connected to the main control and storage unit, is used for sound pickup; A voice output unit, electrically connected to the main control and storage unit, is used for voice feedback during the interaction process; and The transaction function unit is electrically connected to the main control and storage unit and is used to work in conjunction with the AI voice interaction module to synchronize voice commands and transaction operations. The AI voice interaction module includes: The voice wake-up unit is electrically connected to the main control and storage unit and pre-stores wake-up words; A local AI voice engine, electrically connected to the main control and storage unit, is used for speech recognition and parsing and to generate natural speech feedback locally. The AI computing power scheduling unit, electrically connected to the main control and storage unit, is used to dynamically allocate the computing power of the local AI computing power unit; and The functional linkage unit is electrically connected to the main control and storage unit and is used to establish a mapping relationship between voice commands and transaction functions.
[0006] Furthermore, the main control and storage unit includes a processor, memory, and flash memory.
[0007] Furthermore, the local AI computing unit integrates a dedicated NPU chip, supporting INT8 / FP16 precision operations.
[0008] Furthermore, the multimodal microphone array is a 2-microphone linear array that supports 180° sound pickup, with built-in echo cancellation and noise suppression circuits, and a pickup distance of 0.3-1.5m, suitable for vehicle noise environments.
[0009] Furthermore, the voice output unit includes a full-range speaker and a voice synthesis chip, with a volume adjustable up to 85dB.
[0010] Furthermore, the transaction function unit integrates an NFC card reader module, a QR code recognition module, and a PSAM security module.
[0011] Furthermore, the local AI voice engine includes: The speech recognition module loads a lightweight dialect model, supports 7 major Chinese dialects, has a local recognition accuracy of ≥95%, a noisy environment accuracy of ≥90%, and a recognition latency of ≤500ms. The instruction understanding module uses a rule-based and deep learning hybrid model to parse the user's instruction intent and generate executable instructions for the device. The speech synthesis module generates natural speech feedback locally with a synthesis latency of ≤400ms.
[0012] Furthermore, during voice interaction, the AI computing power scheduling unit allocates 50-60% of the computing power of the local AI computing power unit to the local AI voice engine to ensure response speed; when there is no voice command, the AI computing power scheduling unit reduces the computing power allocated to the local AI voice engine to 20-30% for learning user voice habits and optimizing subsequent recognition accuracy.
[0013] Furthermore, the AI voice interaction module also includes a security verification unit, which triggers secondary verification for high-risk commands. The command is executed only after the verification is successful, preventing accidental operation or malicious commands.
[0014] Furthermore, this invention also provides a method for implementing an in-vehicle toll terminal with local AI computing power and voice interaction functions as described above, comprising the following steps: 1) Device initialization; When the vehicle-mounted toll terminal is powered on, the local AI computing unit loads the local language model, the multimodal microphone array is activated, and the AI computing power scheduling unit enters standby mode, allocating 10-20% of the computing power to listen for wake-up words. 2) Voice wake-up; When a passenger utters a wake-up word, the voice wake-up unit detects the wake-up word, the voice output unit provides feedback voice information, and at the same time, the AI computing power scheduling unit allocates 50-60% of the computing power of the local AI computing power unit to the local AI voice engine. 3) Command input and recognition; Passengers speak voice commands, a multimodal microphone array collects the speech and reduces noise, and a local AI voice engine recognizes the commands and interprets their intent; 4) Function execution and feedback; When the function linkage unit is triggered, the voice output unit provides voice feedback, and the screen displays the result synchronously. 5) Computing power recovery and learning; If no voice commands are given within a set time, the AI computing power scheduling unit will reduce the computing power allocated to the local AI voice engine to 20-30%. At the same time, it will learn the passenger's pronunciation characteristics, automatically generate a user voice characteristic report every week, and incrementally update the local language model in the background to improve the accuracy of subsequent voice recognition without affecting the transaction function.
[0015] Compared with the prior art, the beneficial effects of the present invention are: (1) This invention constructs a local AI computing power architecture, gets rid of the dependence on the cloud, and realizes millisecond-level processing of voice commands; the delay of the entire process of voice wake-up-recognition-feedback is ≤1.5s, and it still responds stably in noisy environments, adapting to vehicle network fluctuation scenarios; (2) This invention improves the adaptability of voice interaction, supports dialect recognition and noise suppression, adapts to the complex environment of the vehicle, covers the elderly and dialect users, and has a recognition accuracy far exceeding that of existing devices, thus improving the ease of use for special groups. (3) This invention achieves precise linkage between voice interaction and core payment functions (card swiping / scanning, balance inquiry, transaction feedback), taking into account both ease of operation and transaction security, reducing manual operation steps, and significantly improving interaction efficiency; (4) This invention has secondary verification of high-risk instructions, which ensures safe use; at the same time, it optimizes user habits through AI learning, and the long-term user experience is continuously improved. (5) The present invention adopts a low power consumption design and dynamic computing power scheduling to avoid the local AI computing power unit from running at full load. Compared with fixed computing power allocation, it effectively reduces power consumption.
[0016] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0017] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a structural block diagram of an on-board toll collection terminal. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0019] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] Please refer to Figure 1 The present invention provides an in-vehicle toll terminal with local AI computing power and voice interaction function, including a hardware module for communication connection and an AI voice interaction module; The hardware modules include: The main control and storage unit includes a quad-core Arm Cortex-A55 processor (1.8GHz) + 4GB RAM (DDR4) + 64GB eMMC flash memory, used for data processing and storing local language models and user interaction logs; The local AI computing unit integrates a dedicated NPU chip (such as the RK3588 NPU with a computing power of ≥2TOPS), supports INT8 / FP16 precision operations, and is electrically connected to the main control and storage units for local speech recognition, instruction understanding and user model learning. The multimodal microphone array uses a 2-microphone linear array that supports 180° sound pickup, and has built-in echo cancellation (AEC) and noise suppression (NS) circuits. The pickup distance is 0.3-1.5m, which is suitable for vehicle noise environment (≤85dB). It is electrically connected to the main control and storage unit for sound pickup. The voice output unit includes a 2W full-range speaker (supporting a frequency response of 300-3400Hz) and a voice synthesis chip. The volume is adjustable to 85dB to cover in-vehicle noise. It is electrically connected to the main control and storage unit for voice feedback during interaction. The transaction function unit integrates an NFC card reader module, a QR code recognition module, and a PSAM security module. It is electrically connected to the main control and storage unit and is used to work with the AI voice interaction module to synchronize voice commands and transaction operations. The AI voice interaction module includes: The voice wake-up unit is electrically connected to the main control and storage unit and has pre-stored wake-up words (such as "bus assistant"). The local AI voice engine, including an ASR (Acoustic Speech Recognition) module, a NLU (Natural Language Understanding) module, and a TTS (Text-to-Speech) module, is electrically connected to the main control and storage units. It is used for speech recognition and parsing, and for locally generating natural speech feedback. The ASR module loads a lightweight dialect model, supporting seven major Chinese dialects (Mandarin, Wu, Yue, Min, Hakka, Gan, and Xiang), achieving a local recognition accuracy of ≥95%, ≥90% in noisy environments, and a recognition latency of ≤500ms. The NLU module uses a rule-based + deep learning hybrid model to parse user command intent (such as "check balance," "scan to pay," "transfer route") and generates executable commands. The TTS module locally generates natural speech feedback (such as "balance 2.5 yuan," "scan successful, please board"), with a synthesis latency of ≤400ms. The AI computing power scheduling unit is electrically connected to the main control and storage units and is used to dynamically allocate the computing power of the local AI computing power unit. During voice interaction, the AI computing power scheduling unit allocates 50-60% of the computing power of the local AI computing power unit to the local AI voice engine to ensure response speed. When there are no voice commands, the AI computing power scheduling unit reduces the computing power allocated to the local AI voice engine to 20-30% for learning user voice habits (such as commonly used user commands and pronunciation features) to optimize subsequent recognition accuracy. The functional linkage unit, electrically connected to the main control and storage units, is used to establish a mapping relationship between voice commands and transaction functions. For example, the voice command "scan to pay" triggers the QR code recognition module to start, and the voice output unit simultaneously prompts "Please show your payment code." Another example is the voice command "check balance," which triggers the NFC card reader module to read the NFC card balance data, and the voice output unit displays the card balance information on the screen of the vehicle-mounted payment terminal. Furthermore, for abnormal voice commands (such as "issue an invoice"), the voice output unit prompts "This function is not currently supported" and records the request in a log. The security verification unit triggers secondary verification (such as voice verification code "Please repeat the number 357") for high-risk commands (such as "Change payment password"). The command is executed only after the verification is successful to prevent accidental operation or malicious commands.
[0021] In this embodiment, the local AI speech engine is developed based on C++. The speech recognition module adopts MFCC feature extraction + CNN-LSTM model, the instruction understanding module adopts intent vocabulary (containing 50+ commonly used instructions) + fuzzy matching algorithm, and the speech synthesis module adopts waveform splicing synthesis technology. The AI computing power scheduling unit is based on the Linux kernel process priority mechanism. The voice interaction process is set to the highest priority (PRI 90), and when there is no interaction, it is reduced to PRI 40 to release computing power resources. The AI voice interaction module automatically generates a user voice feature report once a week (such as the top 5 most frequently used commands and pronunciation deviation features) and updates the local language model. The model update does not affect normal trading functions (incremental update in the background).
[0022] Preferably, the AI voice interaction module is configured with an exception handling mechanism, such as: 1) Wake-up failure: The wake-up word is not recognized after 3 consecutive attempts. The voice output unit prompts "Please come closer to the microphone and say 'Bus Assistant' again", while the screen displays the wake-up word prompt. 2) Command recognition error: When the accuracy of the voice recognition module is less than 85%, the voice output unit will prompt "I didn't hear you clearly, please say it again" and automatically adjust the microphone gain (increase by 5dB). 3) Insufficient computing power: When multiple tasks are concurrent (such as voice interaction + big data upload) cause computing power to be strained, the AI computing power scheduling unit will suspend the user learning function and prioritize the voice interaction and transaction functions. Learning will resume after the load decreases.
[0023] Furthermore, embodiments of the present invention also provide a method for implementing an in-vehicle toll terminal with local AI computing power and voice interaction functions as described above, comprising the following steps: 1) Device initialization; When the vehicle-mounted toll terminal is powered on, the local AI computing unit loads the local language model, the multimodal microphone array is activated, and the AI computing power scheduling unit enters standby mode, allocating 10% of the computing power to listen for wake-up words. 2) Voice wake-up; When a passenger says the wake-up word "bus assistant", the voice wake-up unit detects the wake-up word and the voice output unit responds with the voice message "Hello, how can I help you?" At the same time, the AI computing power scheduling unit allocates 60% of the computing power of the local AI computing power unit to the local AI voice engine. 3) Command input and recognition; When a passenger says the voice command "Check balance", the multimodal microphone array collects the voice and reduces noise, the local AI voice engine recognizes the command and interprets the command intent as "Read NFC card balance". 4) Function execution and feedback; The function linkage unit triggers the NFC card reader module to start. After the passenger swipes the card, the balance data (such as 2.5 yuan) is read, and the voice output unit responds with the voice message "Your card balance is 2.5 yuan", which is displayed on the screen simultaneously. Afterwards, if the passenger continues to say the voice command "scan to pay" within the set time (30 seconds), the voice recognition module will analyze and trigger the QR code recognition module. The voice output unit will prompt "Please show the payment code". After the code is successfully scanned, the voice output unit will reply "Payment of 1.8 yuan, successful". 5) Computing power recovery and learning; If no voice command is given within the set time (30 seconds), the AI computing power scheduling unit will reduce the computing power allocated to the local AI voice engine to 30%. At the same time, it will learn the passenger's pronunciation characteristics, automatically generate a user voice characteristic report every week, and incrementally update the local language model in the background to improve the accuracy of subsequent voice recognition without affecting the transaction function.
[0024] In this embodiment, this application can be applied to the following application scenarios: 1) Daily passenger interaction: Elderly passengers can check their balance and scan codes in their local dialect without manual operation; commuters can use voice to ask "where is the next stop?" and "transfer route", improving travel efficiency; 2) Driver Equipment Management: Drivers can check today's revenue and restart equipment via voice commands without leaving the vehicle to operate the terminal, ensuring driving safety; 3) Special scenario adaptation: In noisy road sections (such as business districts and intersections), the microphone noise reduction function ensures accurate command recognition; when the network is interrupted, the local AI still works normally, with no risk of functional failure.
[0025] This application greatly improves the convenience and adaptability of vehicle-mounted toll collection terminals, providing an excellent user experience and possessing significant value for widespread application.
[0026] In the description of this specification, the terms "connection," "installation," and "fixing," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0027] In the description of this specification, the terms "one embodiment," "some embodiments," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0028] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A vehicle-mounted toll terminal with local AI computing power and voice interaction function, characterized in that, This includes a hardware module for communication connectivity and an AI voice interaction module; The hardware module includes: The main control and storage unit is used for data processing and storing local language models and user interaction logs; The local AI computing unit is electrically connected to the main control and storage unit and is used for local speech recognition, instruction understanding and user model learning. A multimodal microphone array, electrically connected to the main control and storage unit, is used for sound pickup; A voice output unit, electrically connected to the main control and storage unit, is used for voice feedback during the interaction process; and The transaction function unit is electrically connected to the main control and storage unit and is used to work in conjunction with the AI voice interaction module to synchronize voice commands and transaction operations. The AI voice interaction module includes: The voice wake-up unit is electrically connected to the main control and storage unit and pre-stores wake-up words; A local AI voice engine, electrically connected to the main control and storage unit, is used for speech recognition and parsing and to generate natural speech feedback locally. The AI computing power scheduling unit, electrically connected to the main control and storage unit, is used to dynamically allocate the computing power of the local AI computing power unit; and The functional linkage unit is electrically connected to the main control and storage unit and is used to establish a mapping relationship between voice commands and transaction functions.
2. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, The main control and storage unit includes a processor, memory, and flash memory.
3. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, The local AI computing unit integrates a dedicated NPU chip, supporting INT8 / FP16 precision computation.
4. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, The multimodal microphone array is a 2-microphone linear array that supports 180° sound pickup, with built-in echo cancellation and noise suppression circuits, and a pickup distance of 0.3-1.5m, suitable for vehicle noise environments.
5. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, The voice output unit includes a full-range speaker and a voice synthesis chip, with a volume adjustable up to 85dB.
6. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, The transaction function unit integrates an NFC card reader module, a QR code recognition module, and a PSAM security module.
7. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, The local AI voice engine includes: The speech recognition module loads a lightweight dialect model, supports 7 major Chinese dialects, has a local recognition accuracy of ≥95%, a noisy environment accuracy of ≥90%, and a recognition latency of ≤500ms. The instruction understanding module uses a rule-based and deep learning hybrid model to parse the user's instruction intent and generate executable instructions for the device. The speech synthesis module generates natural speech feedback locally with a synthesis latency of ≤400ms.
8. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, During voice interaction, the AI computing power scheduling unit allocates 50-60% of the computing power of the local AI computing power unit to the local AI voice engine to ensure response speed; when there is no voice command, the AI computing power scheduling unit reduces the computing power allocated to the local AI voice engine to 20-30% for user voice habit learning and to optimize subsequent recognition accuracy.
9. The vehicle-mounted toll terminal with local AI computing power and voice interaction function as described in claim 1, characterized in that, The AI voice interaction module also includes a security verification unit, which triggers secondary verification for high-risk commands. The command is executed only after the verification is successful, preventing accidental operation or malicious commands.
10. A method for implementing an in-vehicle toll terminal with local AI computing power and voice interaction function as described in any one of claims 1-9, characterized in that, Includes the following steps: 1) Device initialization; When the vehicle-mounted toll terminal is powered on, the local AI computing unit loads the local language model, the multimodal microphone array is activated, and the AI computing power scheduling unit enters standby mode, allocating 10-20% of the computing power to listen for wake-up words. 2) Voice wake-up; When a passenger utters a wake-up word, the voice wake-up unit detects the wake-up word, the voice output unit provides feedback voice information, and at the same time, the AI computing power scheduling unit allocates 50-60% of the computing power of the local AI computing power unit to the local AI voice engine. 3) Command input and recognition; Passengers speak voice commands, a multimodal microphone array collects the speech and reduces noise, and a local AI voice engine recognizes the commands and interprets their intent; 4) Function execution and feedback; When the function linkage unit is triggered, the voice output unit provides voice feedback, and the screen displays the result synchronously. 5) Computing power recovery and learning; If no voice commands are given within a set time, the AI computing power scheduling unit will reduce the computing power allocated to the local AI voice engine to 20-30%. At the same time, it will learn the passenger's pronunciation characteristics, automatically generate a user voice characteristic report every week, and incrementally update the local language model in the background to improve the accuracy of subsequent voice recognition without affecting the transaction function.