Intelligent asset management method and system based on multi-modal interaction and intent understanding
By leveraging multimodal interaction and intent understanding technologies, combined with speech recognition and sensor verification, structured business records are generated, resolving the information gap problem in existing asset management equipment and enabling automated, reliable asset management records and compliant data collection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 周乐熙
- Filing Date
- 2026-03-02
- Publication Date
- 2026-05-29
AI Technical Summary
Existing asset management equipment cannot automatically and accurately map physical access operations into structured records in the digital space, resulting in information gaps and management blind spots, failing to meet the refined management needs of enterprises or families.
By using multimodal interaction and intent understanding technologies, user voice commands are collected and intent is recognized. Consistency verification is performed by combining sensor array data, generating structured business records and guiding users to provide relevant information. This achieves a paradigm shift from "protecting assets from unauthorized access" to "ensuring that every legitimate operation is reliably recorded."
It enables seamless automatic recording of compliance requirements, improves the authenticity and reliability of recorded data, provides data foundation support for advanced applications, and ensures the availability of core functions through a degradation path when voice interaction fails, thus ensuring data integrity.
Smart Images

Figure CN122116513A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of Internet of Things (IoT) and artificial intelligence (AI) interaction technology, specifically to an intelligent asset management method and system based on multimodal interaction and intent understanding. Background Technology
[0002] Currently, the technological development of asset management equipment, especially safes and vaults, mainly focuses on physical security and access control. Traditional mechanical safes only provide key- or password-based locking functions and cannot record any operational information. While subsequent electronic safes have introduced fingerprint and IC card authentication mechanisms and possess simple log recording capabilities, they are limited to recording "when and who" opened the door, creating a so-called "empty log." With the development of IoT technology, some high-end devices have begun to have remote alarm or status monitoring functions, but their core logic remains within the scope of "access control systems," focusing only on the opening permissions and status of the door, without sensing the details of asset movement after the door is opened. This results in asset management still relying on manual post-event recording or trust mechanisms, leading to information gaps and management blind spots.
[0003] Existing technologies attempt to compensate for missing information by adding features such as image capture or manual note input, but these solutions have fundamental flaws. First, requiring users to make additional records after an operation violates the "least effort principle," resulting in extremely low compliance rates in practical use and failing to generate effective data. Second, even if users input notes, the data is in the form of unstructured images or text, making automated analysis and business correlation difficult, creating "data silos." Finally, existing devices lack the ability to perceive the intent of operations, failing to distinguish between the distinct business actions of "retrieving the official seal" and "retrieving petty cash," rendering their log data inadequate for the refined management needs of enterprises or households. Ultimately, the technical thinking of existing technologies is limited to "access control" and has not risen to the level of "business operation management."
[0004] Therefore, how to automatically, accurately, and seamlessly map asset access operations in the physical world into structured records rich in business semantics in the digital space, and embed compliance control into the operational process, thereby achieving a paradigm shift from "protecting assets from unauthorized access" to "ensuring that every legitimate operation is credibly recorded," has become a pressing technical problem to be solved in this field. Summary of the Invention
[0005] This invention embeds multimodal interaction and intent understanding into the physical access operation process, realizing a "seamless compliance" asset management paradigm that automatically maps a physical asset access operation to a structured business record, effectively solving the problems in the background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent asset management method based on multimodal interaction and intent understanding, applied to intelligent asset management equipment, comprising: In response to successful user authentication, the system collects the user's voice commands. The voice commands are subjected to intent recognition and slot filling to extract structured operation information, which includes at least the operation intent, asset identifier, and operation quantity. Based on the structured operation information, a structured business record corresponding to this access operation is generated; Based on the structured business records, an authorization decision is made, and the access door of the device is opened after the decision is approved.
[0007] Furthermore, based on the user's historical operation context or user configuration information, dynamic inquiry voice is generated and played to guide the user to provide voice commands related to the current access operation.
[0008] Furthermore, after extracting the structured operation information, the process also includes: Obtain physical state change data corresponding to the asset identifier, the physical state change data being collected by a sensor array integrated within the device; Perform a consistency check between the physical state change data and the number of operations; If the verification fails, a clarification message is generated, and the user's confirmation command is collected again.
[0009] Furthermore, the sensor array includes a weighing sensor disposed within the storage compartment; the physical state change data is a weight change value Δm; the operation quantity is the number or weight of assets; and the consistency verification includes comparing the weight change value Δm with the weight value represented by the operation quantity.
[0010] Furthermore, if no valid voice command is collected within a preset time, or if intent recognition fails, a delayed recording mode is entered. This delayed recording mode includes: Control the opening of the access bay door; Record the timestamp, user ID, and physical state change data collected by the sensor for this operation, and generate a structured record with a pending task flag; A reminder message with supplementary instructions is sent to the user terminal to update the structured record upon receiving subsequent supplementary voice commands from the user.
[0011] An intelligent asset management system based on multimodal interaction and intent understanding includes: The identity authentication module is used to authenticate the user's identity; The voice acquisition module is used to collect the user's voice commands after identity authentication is successful; An edge computing controller is connected to both the identity authentication module and the voice acquisition module, and the edge computing controller includes: An intent understanding unit is used to perform intent recognition and slot filling on the voice commands to extract structured operation information, which includes at least the operation intent, asset identifier, and operation quantity. The record generation unit is used to generate a structured business record corresponding to the current access operation based on the structured operation information. The permission decision unit is used to execute permission decisions based on the structured business records; An electronic lock control module is used to controllably open the access compartment door of the device after the authorization decision is passed.
[0012] Furthermore, the edge computing master controller also includes a dialogue management unit, which generates dynamic query text based on the user's historical operation context or user configuration information after identity authentication is successful; the system also includes a speech synthesis module, which synthesizes the dynamic query text into guiding speech and plays it.
[0013] Furthermore, it also includes a sensor array installed within the storage bay for collecting data on changes in the physical state of the assets; the edge computing controller also includes: The data fusion verification unit is used to verify the consistency between the number of operations extracted by the intent understanding unit and the physical state change data collected by the sensor array, and to trigger the speech synthesis module to play a clarification prompt based on the verification result.
[0014] Furthermore, the sensor array includes at least a weighing sensor for sensing changes in mass within the storage compartment; the number of operations extracted by the intent understanding unit is a weight value or a quantity value, and the data fusion verification unit is used to compare the weight change value Δm collected by the weighing sensor with the weight value corresponding to the number of operations.
[0015] Furthermore, the edge computing master controller also includes a degradation processing unit, which is used to control the electronic lock control module to open the access door and generate a to-do record containing a timestamp, user identifier and sensor data when the voice acquisition module fails to acquire valid voice within a preset time or the intent understanding unit fails to recognize it, and to trigger the sending of a reminder message with supplementary instructions to the user terminal.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. After user authentication is passed but before the access action is executed, this invention guides the user to naturally express their operation intention through dynamic voice inquiry, and automatically performs intention recognition and slot filling, transforming the traditional passive recording or post-event filling into active collection in the process; this mechanism seamlessly embeds compliance requirements into the core operation path, and automatically generates structured business records containing operation intention, asset identification, quantity and purpose without additional user operation, fundamentally solving the technical problems of log hollowness and low compliance data collection rate in the existing technology.
[0017] 2. This invention performs real-time cross-verification between the operational information extracted from speech semantic understanding and the physical state change data collected by the sensor array. For example, it compares and verifies the quantity stated by the user in speech with the weight change value detected by the weighing sensor. This multimodal fusion mechanism effectively identifies and corrects user misstatements or system misidentifications, greatly improves the authenticity and reliability of the recorded data, achieves anti-fraud effects that cannot be achieved by single-modal technology, and enhances the robustness of the system.
[0018] 3. Since each operation record is standardized data rich in business semantics, this invention provides a data foundation for subsequent advanced applications such as automatic reconciliation, balance calculation, operation behavior analysis, and abnormal risk warning. Simultaneously, the system is designed with a delayed record degradation path, ensuring the availability of core access functions even when voice interaction fails or the user is unable to respond. Furthermore, a to-do reminder mechanism ensures the ultimate integrity of the data, realizing a business model innovation that extends from single device sales to data services. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the intelligent asset management method provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1: This invention provides an intelligent asset management system based on multimodal interaction and intent understanding, including: an identity authentication module 101 for authenticating user identities. In this example, the identity authentication module 101 employs a high-security fingerprint sensor, specifically an FPC series fingerprint sensor, which features liveness detection to effectively prevent fake fingerprint attacks. It is understood that in other examples, the identity authentication module 101 may also employ one or more combinations of a face recognition camera, iris recognition sensor, or near-field communication card reader to meet the needs of different security levels and application scenarios.
[0022] The voice acquisition module 102 is used to collect the user's voice commands after successful identity authentication. Considering that smart asset management devices are usually placed in indoor environments and there may be mechanical noise when opening the box, the voice acquisition module 102 in this embodiment adopts a microphone array composed of multiple microphones, specifically a dual-microphone array, which supports beamforming and echo cancellation functions. It can clearly pick up the user's voice within a range of 0.5 meters to 3 meters and effectively suppress environmental noise and mechanical vibration noise when the box door is opened.
[0023] The edge computing controller 103 is connected to the identity authentication module 101 and the voice acquisition module 102, respectively. The edge computing controller 103 employs a high-performance embedded system-on-a-chip (SoC). In this embodiment, the Rockchip RK3568 chip is selected, which integrates a quad-core ARM Cortex-A55 processor and a neural network processing unit with 0.8 TOPS computing power, capable of meeting the computational requirements of real-time speech recognition and lightweight deep learning model inference. The edge computing controller 103 includes an intent understanding unit 1031, used for intent recognition and slot filling of voice commands to extract structured operation information. The structured operation information includes at least the operation intent, asset identifier, and operation quantity. In this embodiment, the intent understanding unit 1031 runs a lightweight natural language understanding model. This model adopts a joint architecture of bidirectional long short-term memory network and conditional random field, pre-trained with labeled corpora, and is able to identify predefined intent categories and fill corresponding slots. For example, when a user says "Take 500 yuan to pay for express delivery", the intent understanding unit 1031 outputs: intent = take_out, item name = cash, quantity = 500, unit = yuan, purpose = express delivery fee.
[0024] The record generation unit 1032 is used to generate a structured business record corresponding to the current access operation based on structured operation information. The record is encapsulated in JSON format and includes user ID, timestamp, operation intent, asset identifier, operation quantity, purpose description, and sensor data fields, which facilitates subsequent storage, transmission, and analysis.
[0025] The permission decision-making unit 1033 is used to execute permission decisions based on structured business records. The permission decision-making unit 1033 has a built-in rule engine that makes judgments based on preset access control lists and business rules. For example, for the "retrieve official seal" operation, the rule engine checks whether the current user is in the "official seal authorization list" and whether the current time is within the permitted operation time window.
[0026] The electronic lock control module 104 is used to controllably open the access door of the device after the authorization decision is passed. In this embodiment, the electronic lock control module 104 adopts an electromagnetically driven lock with a response time of less than 200 milliseconds and has an unlocking status feedback function, which can transmit the door status back to the edge computing master controller 103 in real time.
[0027] As a further improvement, the system also includes a speech synthesis module 105, used to synthesize text into speech and play it. The speech synthesis module 105 can use an offline speech synthesis chip, has multiple preset timbres, supports mixed Chinese and English reading, and can convert system prompts, queries, confirmation messages, etc. into natural and fluent speech output.
[0028] Sensor array 106, installed within the storage compartment, is used to collect data on changes in the physical state of assets. In this embodiment, sensor array 106 includes at least a weighing sensor 1061, integrated into the load-bearing plate at the bottom of the compartment. The weighing sensor 1061 has a range of 0-10 kg and an accuracy of ±1 g, enabling real-time monitoring of changes in the total weight of items within the compartment and outputting the weight change value Δm. It is understood that, depending on the type of asset being managed, sensor array 106 may also include one or more combinations of infrared photocell arrays, RFID readers, or miniature cameras.
[0029] The edge computing controller 103 also includes a dialogue management unit 1034, which generates dynamic query text based on the user's historical operation context or user configuration information after successful authentication. The dialogue management unit 1034 maintains a dialogue state tracker, recording information such as the current session round, filled slots, and slots to be queried, and generates the next query content according to a predefined dialogue strategy. For example, when it is detected that user A's last operation was "retrieved the official seal" and it has not yet been returned, the dialogue management unit 1034 generates the query text "Do you want to return the official seal?".
[0030] The data fusion verification unit 1035 is used to verify the consistency between the number of operations extracted by the intent understanding unit 1031 and the physical state change data collected by the sensor array 106, and trigger the speech synthesis module 105 to play a clarification prompt based on the verification result. In this embodiment, the specific verification logic of the data fusion verification unit 1035 is as follows: when the number of operations extracted by the intent understanding unit 1031 involves weight or items that can be converted into weight, the difference Δm between the weighing sensor 1061 before and after the current operation is read; Δm is compared with the theoretical weight value corresponding to the number of operations; if the error exceeds a preset threshold, the verification is determined to have failed, and a clarification prompt message is generated.
[0031] The degradation processing unit 1036 is used to control the electronic lock control module 104 to open the access compartment door when the voice acquisition module 102 fails to acquire valid voice within a preset time or the intent understanding unit 1031 fails to recognize it, and to generate a to-do record containing a timestamp, user identifier and sensor data, and trigger a reminder message to be sent to the user terminal with supplementary instructions.
[0032] The system also includes a communication module 107 for data synchronization between the edge computing controller 103 and the cloud platform or user mobile terminal. The communication module 107 supports both Wi-Fi and Bluetooth modes. When Wi-Fi is available, it automatically connects to the cloud for encrypted data upload. When the network is unavailable, the data is temporarily stored in the local storage chip and will be uploaded again when the network is restored. Example
[0033] Please refer to the following: Figure 1 This embodiment provides an intelligent asset management method based on multimodal interaction and intent understanding, applied to the intelligent asset management device described in Embodiment 1. The method includes the following steps: Step S201: In response to successful user authentication, collect the user's voice commands.
[0034] Specifically, users authenticate themselves using a fingerprint sensor. After the edge computing controller 103 verifies that the fingerprint information matches the pre-stored fingerprint template, authentication is successful, and the system enters a "pre-operation state." At this time, the voice acquisition module 102 starts up, enters voice wake-up monitoring mode, and waits for user voice input. To ensure a natural interaction, the system can play a short prompt tone to indicate that the user can start speaking.
[0035] As a preferred embodiment, step S200 is included before step S201: generating and playing dynamic inquiry voice based on the user's historical operation context or user configuration information to guide the user to provide voice commands related to the current access operation.
[0036] Specifically, the dialogue management unit 1034 retrieves the user's recent operation records and personalized configurations from local storage or the cloud. For example, if the system detects that user B retrieved the "Project Department Seal" two hours ago and has not yet returned it, it generates the query text "Do you want to return the Project Department Seal?", which is played by the speech synthesis module 105. If the user is using the service for the first time or there is no relevant context, it plays the general query "Please explain the purpose of this operation." This guidance mechanism can significantly improve the standardization and information completeness of user voice commands, and improve the accuracy of subsequent intent recognition.
[0037] Step S202: Perform intent recognition and slot filling on the voice commands to extract structured operation information, which includes at least the operation intent, asset identifier, and operation quantity.
[0038] Specifically, the analog speech signal acquired by the speech acquisition module 102 is converted from analog to digital and then sent to the intent understanding unit 1031. The intent understanding unit 1031 first performs automatic speech recognition, converting the speech signal into text. This embodiment uses the Vosk lightweight offline speech recognition engine, which supports Chinese and English recognition. The vocabulary can be customized according to the application scenario, such as pre-setting commonly used asset management terms like "cash," "official seal," "contract," and "gold bar."
[0039] The identified text is input to the Natural Language Understanding (NLE) module. The NLE module uses a miniaturized pre-trained BERT-based model for intent classification and combines it with a bidirectional LSTM-CRF model for slot filling. Example of model output: For the input text "I want to take two gold bars as gifts for clients," the intent classification result is "take_out," and the slot filling result is: {"item_name": "gold bar", "amount": "2", "unit": "bar", "purpose": "gift for clients"}. Here, "gold bar" is identified through dictionary matching, "2" is extracted as a number using regular expressions, and "gift for clients" is categorized into the "business gift" purpose category through keyword classification.
[0040] Step S203: Obtain physical state change data corresponding to the asset identifier. The physical state change data is collected by a sensor array integrated into the device.
[0041] Specifically, while the user is interacting via voice, the sensor array 106 continuously collects data. Taking the weighing sensor 1061 as an example, the system records the initial weight value W0 at the moment of successful identity authentication, and reads the current weight value W1 after step S202, calculating the weight change value Δm = W1 - W0. A positive value of Δm indicates that an item has been stored, and a negative value indicates that an item has been retrieved.
[0042] Step S204: Perform a consistency check between the physical state change data and the number of operations.
[0043] The data fusion verification unit 1035 compares the number of operations extracted in step S202 with the weight change value collected in step S203. For example, if the user's voice command is to take out two gold bars, and the system pre-stores the standard weight of a single gold bar as 50g, then the theoretical weight reduction should be 100g. If the actual detected Δm is within the range of -98g to -102g, the verification passes; if Δm is -50g or -150g, the verification fails, indicating that there may be misrepresentation by the user, incorrect item specifications, or taking more or less than expected. For items such as cash whose weight per bill cannot be predicted, the system can set a weight change threshold. If the absolute value of Δm exceeds the preset range, such as banknotes weighing approximately 1g each (the weight of a 500 yuan banknote varies depending on its thickness), a ±5g error can be set, triggering a verification failure.
[0044] Step S205: If the verification fails, a clarification message is generated, and the user's confirmation command is collected again.
[0045] If the verification in step S204 fails, the dialogue management unit 1034 generates a clarification prompt text, such as "The detected weight change does not match your description. How much did you actually remove?", which is played by the speech synthesis module 105. The voice acquisition module 102 re-acquires the user's voice, the intent understanding unit 1031 extracts the information again, and performs a second verification with the sensor data. If the verification fails three times consecutively, the system can proceed to the manual review process, marking the record as "pending manual confirmation" and notifying the administrator.
[0046] Step S206: Based on the structured operation information, generate a structured business record corresponding to this access operation.
[0047] The record generation unit 1032 integrates the structured operation information extracted in step S202 with the sensor data, user ID, timestamp, and other information collected in step S203 to generate a complete JSON format record. An example record is shown below: {"record_id":"20230520153045001", "timestamp":"2023-05-2015:30:45", "user_id":"zhangsan_001", "auth_method":"fingerprint", “intent”:“take_out”, “item_name”:“cash”, "amount":500, "unit": "yuan", "purpose": "express delivery fee", "sensor_data": { "weight_delta": -5.2, "unit": "g" }, "status": "confirmed" } Step S207: Perform permission adjudication based on the structured business record.
[0048] The permission adjudication unit 1033 calls the rule engine for adjudication based on the record generated in step S206. The rule engine supports custom rules. For example: Rule 1: Any "retrieval" operation requires checking whether the user has the operation permission for the item; Rule 2: A secondary approval is required for a single cash withdrawal exceeding 1000 yuan; Rule 3: Operations during non-working hours require administrator confirmation. If the current operation meets all the rule conditions, the adjudication result is "passed"; otherwise, it is "rejected" or "pending approval". The adjudication result is returned to the edge computing master 103.
[0049] Step S208: Control the opening of the access hatch of the device after the adjudication is passed.
[0050] If the adjudication in step S207 is passed, the edge computing master 103 sends an unlocking instruction to the electronic lock control module 104, and the electronic lock control module 104 drives the electromagnetic lock to open the hatch. At the same time, the voice synthesis module 105 plays a prompt voice: "Unlocking successful, please pick up the item". If the adjudication fails, a rejection prompt is played, such as "You have no right to take out the official seal, and the administrator has been notified", and the record of this operation is marked as "denied access" and uploaded to the cloud.
[0051] It should be noted that the core innovation of this invention is to transform the unlocking action from "unlock upon authentication" to a complete business process of "authentication - interaction - record - adjudication - unlock". In the traditional solution, unlocking is an inevitable result of authentication; while in this invention, unlocking is the final action after passing the business compliance review. This transformation upgrades the device from a passive "gatekeeper" to an active "business administrator". Embodiment
[0052] Based on the above embodiment, this embodiment further describes the specific implementation of the delayed recording mode.
[0053] After step S201, if the voice acquisition module 102 does not detect valid voice input within a preset time, or if the confidence level of the intent understanding unit 1031 after recognizing the user's voice is lower than the threshold, the system determines that the interaction has failed and enters the delayed recording mode.
[0054] The specific process of the delayed recording mode is as follows: First, the degradation processing unit 1036 controls the electronic lock control module 104 to directly open the access compartment door, ensuring that the user's emergency access needs are not affected. This design reflects the high availability principle of the system and avoids the extreme situation where users cannot use the device due to voice interaction failure.
[0055] Secondly, the sensor array 106 automatically records the physical state changes during the operation. For example, the weighing sensor 1061 records the weight change Δm before and after the operation; if infrared sensors are installed in the compartment, they record which item compartments were touched. Simultaneously, the system records the timestamp and user identifier of the operation.
[0056] Then, the record generation unit 1032 generates a structured record marked with "to-do". This record does not contain business semantic information (such as purpose or specific item name), but only contains a timestamp, user identifier, sensor data, and a unique to-do ID. This record is stored locally and marked as pending synchronization.
[0057] Finally, the downgrade processing unit 1036 sends a push notification to the user's bound mobile terminal via the communication module 107, reminding the user to supplement the operation instructions. After clicking the notification, the user can supplement the specific details of this operation via voice or text input on the mobile app, such as "retrieve the property certificate for mortgage processing". After receiving the supplementary information, the system updates the original pending record to a complete structured business record and associates it with the sensor data.
[0058] The technical advantage of the delayed recording mode is that, while ensuring the availability of core access functions, it maximizes the integrity of business data through a post-event reminder mechanism, thus achieving a balance between user experience and data compliance. Example
[0059] This embodiment, combined with a specific application scenario, further illustrates the implementation process and technical effects of the present invention.
[0060] Scenario 1: Corporate Petty Cash Management User Li Si needs to withdraw petty cash for company purchases. Li Si authenticates his identity with his fingerprint on the smart safe. After successful authentication, the system plays a dynamic inquiry: "Hello Li Si, please explain the matter." Li Si replies: "I need to withdraw 500 yuan to pay for express delivery." The intent understanding unit 1031 recognizes the intent as "withdrawal," the item name as "cash," the quantity as "500," and the purpose as "express delivery fee." The weighing sensor detects a weight reduction of approximately 5 grams (matching the weight of 500 yuan banknotes), and the data fusion verification passes. The permission adjudication unit 1033 checks Li Si's petty cash withdrawal permission and finds that his daily limit is 1000 yuan. After this withdrawal, the daily cumulative limit is 500 yuan, which is within the limit, and the adjudication passes. The system plays a confirmation voice: "Recorded: 500 yuan cash withdrawn, purpose: express delivery fee. The petty cash balance will be updated to 1500 yuan. Unlocking." After unlocking, the structured record is uploaded to the company's financial system, the petty cash account is automatically reduced by 500 yuan, and the financial staff can view the cash balance in real time in the management backend.
[0061] Scenario 2: Emergency access to documents at home The father needed to retrieve the property certificate late at night, but the family was already asleep and didn't want to speak. After the father's fingerprint authentication was successful, the system played a voice prompt, but the father didn't respond. After waiting 3 seconds, the system determined the interaction had failed, entered delayed recording mode, and directly opened the locker door. After the father took the property certificate, the system recorded the time of the operation, the user, and the signal triggered by the infrared sensor in the document compartment inside the locker (detecting that the property certificate's location had been touched). The next morning, the father received an app push notification on his phone: "Please add the retrieval details from last night at 22:35." The father clicked on it and voice-completed, "Retrieving the property certificate for mortgage processing," and the system updated the record, completing the data loop.
[0062] The two scenarios described above fully demonstrate the adaptability of this invention in different situations: achieving automated recording of "seamless compliance" in normal scenarios; ensuring availability through a degradation path in special scenarios, and ensuring data integrity through post-event reminders. Example
[0063] This embodiment further explains the hardware selection and specific parameters of each module of the system to facilitate implementation by those skilled in the art.
[0064] Identity authentication module 101: In addition to the fingerprint sensor, a face recognition module can be selected, using an RGB+infrared dual-lens camera, supporting liveness detection, with a face recognition accuracy of ≥99.5% and a recognition time of <1s. For high-security scenarios, a finger vein recognition module can be selected, with a recognition time of <0.5s and a false recognition rate of <0.0001%.
[0065] Voice acquisition module 102: The microphone array adopts a 4-microphone circular array, supporting 360-degree sound source localization and beamforming, with a pickup radius of up to 5 meters. Employing echo cancellation and noise reduction algorithms, it maintains over 90% speech recognition accuracy even in 60dB ambient noise conditions.
[0066] Edge computing controller 103: In addition to RK3568, embedded platforms such as Allwinner A133 or Qualcomm QCS8250 can also be selected. For scenarios requiring higher AI computing power, an external AI acceleration module, such as Rockchip RK1808 compute stick, can be connected, providing up to 3.0 TOPS of INT8 computing power.
[0067] Speech synthesis module 105: can use the Yuyin Tianxia SYN6658 speech synthesis chip, supports mixed reading of Chinese and English, provides 16 voice options, and the synthesis speed is adjustable from 80 to 400 words per minute.
[0068] Weighing sensor 1061: Employs a resistance strain gauge pressure sensor, arranged at four corners, and acquires signals via a Wheatstone bridge and a 24-bit ADC (such as HX711), with a sampling frequency of 80Hz and a resolution of up to 0.1g. The system requires zero-point and full-scale calibration during installation, and periodic temperature compensation calibration to ensure long-term stability.
[0069] Communication Module 107: In addition to Wi-Fi and Bluetooth, a 4G / 5G communication module can be added for enterprise-level applications to ensure real-time data uploads even in environments with unstable networks. The communication protocol uses MQTT over TLS to ensure the encryption and reliability of data transmission.
[0070] It is worth noting that the hardware models and parameters disclosed in this embodiment are merely illustrative examples. Those skilled in the art can make reasonable selections and adjustments based on actual application scenarios and cost requirements. As long as the functions and effects described in this invention can be achieved, they all fall within the protection scope of this invention. For home application scenarios, the lower-cost ESP32-S3 chip can be selected as the edge computing master controller, which has a built-in neural network accelerator sufficient to run lightweight speech recognition models. For enterprise-level high-concurrency scenarios, the higher-performance RK3588 chip can be used, supporting multi-channel concurrent processing.
[0071] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A smart asset management method based on multimodal interaction and intent understanding, applied to smart asset management equipment, characterized in that: include: In response to successful user authentication, the system collects the user's voice commands. The voice commands are subjected to intent recognition and slot filling to extract structured operation information, which includes at least the operation intent, asset identifier, and operation quantity. Based on the structured operation information, a structured business record corresponding to this access operation is generated; Based on the structured business records, an authorization decision is made, and the access door of the device is opened after the decision is approved.
2. The method according to claim 1, characterized in that, Before collecting the user's voice commands, the method also includes: Based on the user's historical operation context or user configuration information, generate and play dynamic inquiry voice to guide the user to provide voice commands related to the current access operation.
3. The method according to claim 1 or 2, characterized in that, After extracting the structured operation information, the process also includes: Obtain physical state change data corresponding to the asset identifier, the physical state change data being collected by a sensor array integrated within the device; Perform a consistency check between the physical state change data and the number of operations; If the verification fails, a clarification message is generated, and the user's confirmation command is collected again.
4. The method according to claim 3, characterized in that, The sensor array includes a weighing sensor installed in the storage compartment; the physical state change data is a weight change value Δm; the operation quantity is the number or weight of assets; and the consistency verification includes comparing the weight change value Δm with the weight value represented by the operation quantity.
5. The method according to claim 1, characterized in that, If no valid voice command is collected within a preset time, or if intent recognition fails, the system enters a delayed recording mode, which includes: Control the opening of the access bay door; Record the timestamp, user ID, and physical state change data collected by the sensor for this operation, and generate a structured record with a pending task flag; A reminder message with supplementary instructions is sent to the user terminal to update the structured record upon receiving subsequent supplementary voice commands from the user.
6. An intelligent asset management system based on multimodal interaction and intent understanding, characterized in that, include: The identity authentication module is used to authenticate the user's identity; The voice acquisition module is used to collect the user's voice commands after identity authentication is successful; An edge computing controller is connected to both the identity authentication module and the voice acquisition module, and the edge computing controller includes: An intent understanding unit is used to perform intent recognition and slot filling on the voice commands to extract structured operation information, which includes at least the operation intent, asset identifier, and operation quantity. The record generation unit is used to generate a structured business record corresponding to the current access operation based on the structured operation information. The permission decision unit is used to execute permission decisions based on the structured business records; An electronic lock control module is used to controllably open the access compartment door of the device after the authorization decision is passed.
7. The system according to claim 6, characterized in that, The edge computing master controller also includes a dialogue management unit, which generates dynamic query text based on the user's historical operation context or user configuration information after identity authentication is successful; the system also includes a speech synthesis module, which synthesizes the dynamic query text into guiding speech and plays it.
8. The system according to claim 6 or 7, characterized in that, It also includes a sensor array installed within the storage bay for collecting data on changes in the physical state of the assets; the edge computing controller also includes: The data fusion verification unit is used to verify the consistency between the number of operations extracted by the intent understanding unit and the physical state change data collected by the sensor array, and to trigger the speech synthesis module to play a clarification prompt based on the verification result.
9. The system according to claim 8, characterized in that, The sensor array includes at least a weighing sensor for sensing changes in mass within the storage compartment; the number of operations extracted by the intent understanding unit is a weight value or a quantity value; and the data fusion verification unit is used to compare the weight change value Δm collected by the weighing sensor with the weight value corresponding to the number of operations.
10. The system according to claim 6, characterized in that, The edge computing master controller also includes a degradation processing unit, which is used to control the electronic lock control module to open the access compartment door and generate a to-do record containing a timestamp, user identifier and sensor data when the voice acquisition module fails to acquire valid voice within a preset time or the intent understanding unit fails to recognize it, and to trigger the sending of a reminder message with supplementary instructions to the user terminal.