Multi-modal office assistant system

The multimodal office assistant system combines voice and visual data for wake-up and command recognition, solving the problem of low efficiency caused by the single interaction mode of traditional office systems. It realizes efficient, secure and personalized office assistant functions, improving user experience and system management capabilities.

CN121053985APending Publication Date: 2025-12-02CISDI INFORMATION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511314119.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Traditional office systems have a single interaction mode, which leads to a decline in office efficiency. Users need to perform tedious manual input operations using a keyboard and mouse.

Method used

The system employs a multimodal office assistant system, which uses a voice wake-up module for wake-up probability prediction and a command execution module for command recognition. By combining voice and visual data for multimodal feature fusion, the system can wake up and execute commands, including functions such as identity verification, schedule management, meeting organization, call processing, and email automation.

Benefits of technology

It improves office efficiency, enhances the accuracy of wake-up probability prediction and command recognition, meets personalized needs, improves user experience and system security, and supports automated and intelligent management of various office scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053985A_ABST
    Figure CN121053985A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent office, in particular to a multi-mode office assistant system, which comprises a voice wake-up module for performing wake-up probability prediction based on first multi-mode data to obtain a wake-up probability prediction value; if the wakeup probability predicted value is larger than a wakeup threshold value, multi-mode office assistant system wakeup is carried out, and the wakeup threshold value is in positive correlation with the current environment noise intensity; the command execution module performs command identification based on the second multi-modal data to determine a target command and execute the target command, and the acquisition time of the first multi-modal data is earlier than the acquisition time of the second multi-modal data; compared with a single user manual input mode, the system can effectively improve the office efficiency, and can better improve the accuracy of wake-up probability prediction and the accuracy of the obtained target command.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent office technology, and in particular to a multimodal office assistant system. Background Technology

[0002] In today's enterprise operations and management, office systems, such as office automation (OA) systems, enterprise resource planning (ERP) systems, and various collaborative office platforms, have become indispensable core tools. These systems aim to optimize workflows, improve organizational efficiency, and reduce operating costs through information technology.

[0003] Currently, traditional or mainstream office systems typically rely on a single, passive method of manual user input for information interaction. Specifically, users primarily use a keyboard and mouse to manually type text and numbers into pre-defined form fields or select from drop-down menus to complete data entry, queries, and command issuance. However, this method is cumbersome and severely impacts office efficiency. Summary of the Invention

[0004] This application provides a multimodal office assistant system to solve the technical problem of decreased office efficiency caused by the single interaction mode of existing office systems.

[0005] This application provides a multimodal office assistant system, the system comprising:

[0006] The voice wake-up module receives first multimodal data; based on the first multimodal data, it performs wake-up probability prediction to obtain a wake-up probability prediction value; if the wake-up probability prediction value is greater than the wake-up threshold, it wakes up the multimodal office assistant system, wherein the wake-up threshold is positively correlated with the current ambient noise intensity.

[0007] The command execution module receives second multimodal data, performs command recognition based on the second multimodal data to determine the target command, and executes the target command. Both the first and second multimodal data include voice data and visual data. The visual data is video data or multi-frame image data containing the image of the voice speaker. The acquisition time of the first multimodal data is earlier than the acquisition time of the second multimodal data.

[0008] In one embodiment of this application, the voice wake-up module is specifically used to extract features from the voice data in the first multimodal data to obtain a first voice feature;

[0009] Feature extraction is performed on the visual data in the first multimodal data to obtain the first lip movement feature and the first limb movement feature of the speaker;

[0010] Based on the first speech feature, the first lip movement feature, and the first limb movement feature, multimodal feature information fusion is performed to obtain the target feature;

[0011] Based on the target characteristics, a wake-up probability prediction is performed to obtain the predicted wake-up probability value.

[0012] In one embodiment of this application, the system further includes: an identity recognition module, which receives target multimodal data before the command execution module receives the second multimodal data; and performs voiceprint recognition based on the voice data in the target multimodal data to determine the identity of the voice speaker;

[0013] Based on the visual data in the target multimodal data, facial recognition is performed on the speaker to obtain a facial recognition result; based on the fingerprint data in the target multimodal data, fingerprint recognition is performed on the speaker to obtain a fingerprint recognition result; based on the facial recognition result and the fingerprint recognition result, the speaker is authenticated.

[0014] If the verification is successful, the target operating environment information corresponding to the current voice speaker is obtained based on the preset operating environment mapping relationship. The operating environment mapping relationship is the mapping relationship between the user and the operating environment information. The operating environment information includes the working interface layout, language setting information, and preference parameters. Based on the target operating environment information, the preset display interface is adjusted.

[0015] In one embodiment of this application, the command execution module includes:

[0016] The schedule management unit, when the target command is a create schedule command, generates schedule entries based on the schedule information in the create schedule command, and synchronizes the schedule entries to the current user's schedule.

[0017] A calendar synchronization unit, which, when the target command is a schedule synchronization command, synchronizes the user's schedule from the original device to one or more target devices;

[0018] The event reminder module sends schedule reminders to the user based on the schedule entries in the schedule and a preset first reminder method; if there is a time conflict in the schedule entries, a schedule conflict reminder is sent to the user according to a preset second reminder method.

[0019] In one embodiment of this application, the command execution module includes:

[0020] The meeting organization unit, when the target command is a meeting organization command, generates initial schedule entries based on the meeting information in the meeting organization command. The meeting information includes: meeting name, meeting start time, meeting end time, and participant information, including participant name and priority for each participant. Based on the participant names, it obtains the schedules of multiple participants to determine each participant's free time. It defines the time period between the meeting start time and the meeting end time as a target time period. If any participant's free time includes the target time period, then that participant is identified as a target participant. The ratio between the number of target participants and the total number of participants is defined as a target percentage.

[0021] If the target percentage is less than a preset percentage threshold, and / or the target participants do not include participants whose priority in the participant information is the target priority, the meeting start time and meeting end time will be adjusted at least once until the target percentage is greater than or equal to the percentage threshold, and the target participants include participants whose priority in the participant information is the target priority. Based on the adjusted meeting start time and meeting end time, the initial schedule entry will be updated to obtain intermediate schedule entries. The office location information of the target participants and the available meeting room information within the target time period will be obtained. Based on the office location information and the meeting room information, the target meeting room will be determined. The information of the target meeting room will be updated to the intermediate schedule entry to obtain the final schedule entry. The final schedule entry will be synchronized to the schedules of multiple participants in the meeting information. If there are pre-configured remote participants in the meeting information, a meeting invitation email and a meeting link will be generated based on the final schedule entry. The meeting link will be embedded in the meeting invitation email, and a meeting request email with the meeting link embedded will be sent to the preset email address of the remote participants.

[0022] In one embodiment of this application, the command execution module includes:

[0023] The call unit answers the current call when the target command is an answer call command; when the target command is an automatic reply command, it saves the automatic reply information in the automatic reply command, the automatic reply information including automatic reply text or automatic reply voice, and sends the automatic reply information to the caller when a call request is received.

[0024] When the target command is a message command, if a call request is received, the message function is activated, prompting the caller to start leaving a message. By performing keyword recognition and tone analysis on the message content, the priority of the message content is determined, and based on the priority, a message processing reminder is sent to the user.

[0025] In one embodiment of this application, the command execution module includes:

[0026] The email automation processing unit, when the target command is an email reading command, displays the content of the email to be read in the email reading command on a preset display interface;

[0027] If the target command is an email reply command, the reply content is obtained based on the email reply command, and the reply content is sent back to the sender of the current email;

[0028] When the target command is an email template retrieval command, the target template in the email template retrieval command is retrieved, and an email to be sent is generated based on the target template and the email content input by voice. The recipient's email address is determined according to the recipient information input by voice, and the email to be sent is sent to the recipient's email address.

[0029] In one embodiment of this application, the command execution module includes:

[0030] The leave application unit, when the target command is a leave application submission command, calls a preset leave form template, fills the leave time and leave reason in the leave application submission command into the leave form template to obtain the target leave form and completes the submission of the target leave form;

[0031] The leave approval unit determines whether the approval result of the leave application in the currently displayed interface is approved or rejected when the target command is a leave application approval instruction.

[0032] In one embodiment of this application, the command execution module includes:

[0033] The robot scheduling module, when the target command is a robot scheduling command, acquires the status and position information of multiple robots based on the communication channel between the multimodal office assistant system and multiple preset robots; identifies robots with idle status information as selectable robots; and acquires the starting position and destination position information of the object to be transported in the robot scheduling command.

[0034] Based on the starting position information and the position information of the selectable robots, a target robot is determined from the plurality of selectable robots, wherein the distance between the target robot and the object to be transported is less than the distance between the other selectable robots and the object to be transported; a control command is sent to the target robot to instruct the target robot to complete the transport task of the object to be transported.

[0035] In one embodiment of this application, the system further includes: a positioning service module, which receives employee positioning information sent by a preset positioning device, the positioning device being set in the office area, and the employee positioning information including the location information of multiple employees;

[0036] The command execution module includes a location query unit. When the target command is a location query command and the voice sender has completed the location query permission verification, the unit sends a query instruction to the location service module based on the employee information to be queried in the location query command, and receives the location information of the employee to be queried from the location service module; and displays the location information of the employee to be queried on a preset display interface or sends it to the terminal device preset by the voice sender.

[0037] The beneficial effects of this application are as follows: The multimodal office assistant system proposed in this application includes: a voice wake-up module, which receives first multimodal data; performs wake-up probability prediction based on the first multimodal data to obtain a predicted wake-up probability value; if the predicted wake-up probability value is greater than a wake-up threshold, the multimodal office assistant system is woken up, and the wake-up threshold is positively correlated with the current ambient noise intensity; and a command execution module, which receives second multimodal data, performs command recognition based on the second multimodal data to determine the target command, and executes the target command. Both the first and second multimodal data include voice data and visual data. The visual data is video data or multi-frame image data containing the image of the speaker. The acquisition time of the first multimodal data is earlier than the acquisition time of the second multimodal data. This system can receive multimodal data, which can effectively improve office efficiency compared to a single manual input method. Furthermore, by performing wake-up probability prediction based on the first multimodal data, the accuracy of the wake-up probability prediction can be improved significantly, and by performing command recognition based on the second multimodal data, the accuracy of the obtained target command can be effectively improved. Attached Figure Description

[0038] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0039] In the attached diagram:

[0040] Figure 1 This is a schematic diagram of the structure of a multimodal office assistant system provided in an embodiment of this application. Detailed Implementation

[0041] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0042] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0043] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.

[0044] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of a multimodal office assistant system provided in an embodiment of this application, as shown below. Figure 1 As shown, the system includes:

[0045] The voice wake-up module 110 receives first multimodal data; based on the first multimodal data, it performs wake-up probability prediction to obtain a wake-up probability prediction value; if the wake-up probability prediction value is greater than the wake-up threshold, it wakes up the multimodal office assistant system, wherein the wake-up threshold is positively correlated with the current ambient noise intensity.

[0046] The command execution module 120 receives second multimodal data, performs command recognition based on the second multimodal data to determine the target command, and executes the target command. Both the first multimodal data and the second multimodal data include voice data and visual data. The visual data is video data or multi-frame image data containing the image of the voice speaker. The acquisition time of the first multimodal data is earlier than the acquisition time of the second multimodal data.

[0047] In some examples of this embodiment, the voice wake-up module 110 can specifically input the first multimodal data into a preset first neural network model to perform wake-up probability prediction and obtain a wake-up probability prediction value. The command execution module 120 can specifically input the second multimodal data into a preset second neural network model to perform target command recognition and obtain a target command. The aforementioned first neural network model and second neural network model can both be convolutional neural network models or long short-term memory network models, etc.

[0048] Understandably, the command execution module 120 starts running after the multimodal office assistant system is woken up; before the multimodal office assistant system is woken up, the command execution module 120 is in a dormant state.

[0049] It should be noted that if the acquisition of the second multimodal data fails, such as only acquiring the voice data of the speaker without acquiring the visual data of the speaker (e.g., due to a malfunction in the visual data acquisition device or lack of visual data acquisition or acquisition permissions), then the aforementioned wake-up probability prediction and command recognition will be performed based on the voice data of the speaker. It is understood that the multimodal office assistant system in this embodiment not only supports the acquisition and processing of multimodal data but also supports the acquisition and processing of single data.

[0050] It should also be noted that the aforementioned voice and visual data were obtained with the appropriate data access permissions and under legal and compliant conditions.

[0051] Understandably, this embodiment, by dynamically adjusting the wake-up threshold based on ambient noise intensity, can help improve wake-up accuracy and reduce the probability of false wake-ups. For example, if the current ambient noise intensity is low (e.g., less than a preset noise intensity threshold), the wake-up threshold is reduced (e.g., reduced by 0.1) based on the preset initial wake-up threshold value (which can be set according to actual needs, such as 0.8), to improve the recognition accuracy of the first multimodal data, i.e., the recognition accuracy and sensitivity of the wake-up command; if the current ambient noise intensity is high (e.g., greater than or equal to the noise intensity threshold), the wake-up threshold is increased (e.g., increased by 0.1) based on the preset initial wake-up threshold value, to reduce the probability of false wake-ups.

[0052] In some examples of this embodiment, the wake-up threshold can also be adjusted based on the user's (voice sender's) usage habits. For example: collect multiple historical wake-up time points of the user (historical wake-up times of the multimodal office assistant system); based on the multiple historical wake-up time points, obtain a target time set, the target time set including at least one target time period, the target time period including multiple historical wake-up time points, the number of historical wake-up time points in the target time period being greater than a preset number threshold (e.g., 50); if the current time is within the target time period, then the current wake-up threshold is decreased; if the current time is outside the target time period, then the current wake-up threshold is increased, the decrease and increase magnitudes can be set according to actual conditions (e.g., 0.1). In this way, the accuracy of waking up the multimodal office assistant system can be improved.

[0053] Understandably, the multimodal office assistant system in this embodiment can receive multimodal data, which can effectively improve office efficiency compared to a single manual input method. Furthermore, by adopting the above method, the accuracy of wake-up probability prediction can be improved, as well as the accuracy of the obtained target commands, which facilitates the subsequent precise command execution.

[0054] In some embodiments, the voice wake-up module 110 is specifically used to extract features from the voice data in the first multimodal data to obtain a first voice feature;

[0055] Feature extraction is performed on the visual data in the first multimodal data to obtain the first lip movement feature and the first limb movement feature of the speaker;

[0056] Based on the first speech feature, the first lip movement feature, and the first limb movement feature, multimodal feature information fusion is performed to obtain the target feature;

[0057] Based on the target characteristics, a wake-up probability prediction is performed to obtain the predicted wake-up probability value.

[0058] In some examples of this embodiment, "multimodal feature information fusion" can be achieved by superimposing multiple features to obtain a new feature vector, i.e., the target feature. Through this method, a highly accurate wake-up probability prediction value can be obtained.

[0059] In some embodiments, the system further includes: an identity recognition module, which receives target multimodal data before the command execution module 120 receives the second multimodal data; and performs voiceprint recognition based on the voice data in the target multimodal data to determine the identity of the voice speaker;

[0060] Based on the visual data in the target multimodal data, facial recognition is performed on the speaker to obtain a facial recognition result; based on the fingerprint data in the target multimodal data, fingerprint recognition is performed on the speaker to obtain a fingerprint recognition result; based on the facial recognition result and the fingerprint recognition result, the speaker is authenticated.

[0061] If the verification is successful, the target operating environment information corresponding to the current voice speaker is obtained based on the preset operating environment mapping relationship. The operating environment mapping relationship is the mapping relationship between the user and the operating environment information. The operating environment information includes the working interface layout, language setting information, and preference parameters. Based on the target operating environment information, the preset display interface is adjusted.

[0062] In some examples of this embodiment, if verification fails, it is re-verified. It is understood that by organically combining multimodal data such as voice data, visual data, and fingerprint data, the security of the system can be effectively improved. Furthermore, by setting a target operating environment corresponding to the user (i.e., the current voice speaker) after successful verification, the personalized needs of different users can be better met, enhancing the user experience.

[0063] In some embodiments, the command execution module 120 includes:

[0064] The schedule management unit, when the target command is a create schedule command, generates schedule entries based on the schedule information in the create schedule command, and synchronizes the schedule entries to the current user's schedule.

[0065] A calendar synchronization unit, which, when the target command is a schedule synchronization command, synchronizes the user's schedule from the original device to one or more target devices;

[0066] The event reminder module sends schedule reminders to the user based on the schedule entries in the schedule and a preset first reminder method; if there is a time conflict in the schedule entries, a schedule conflict reminder is sent to the user according to a preset second reminder method.

[0067] In some examples of this embodiment, the schedule information includes the schedule name (such as the meeting name), the start time of the schedule, and the end time (such as the start time and end time of the meeting).

[0068] In some examples of this embodiment, the event reminder module can provide reminders in various ways, including voice broadcasts, SMS notifications, and desktop pop-ups, to ensure that users do not miss any important events.

[0069] In some examples of this embodiment, the system (multimodal office assistant system) runs user-defined reminder settings, that is, users can choose to receive reminders at different time points before the event (schedule) occurs, or set priority reminders according to the importance of the event.

[0070] In some examples of this embodiment, if the schedule management unit has email access permissions, it can automatically scan and identify the schedule information in the received emails in the mailbox, generate the corresponding schedule entry based on the schedule information, and synchronize the schedule entry to the current user's schedule.

[0071] In some embodiments, the command execution module 120 includes:

[0072] The meeting organization unit, when the target command is a meeting organization command, generates initial schedule entries based on the meeting information in the meeting organization command. The meeting information includes: meeting name, meeting start time, meeting end time, and participant information, including participant name and priority for each participant. Based on the participant names, it obtains the schedules of multiple participants to determine each participant's free time. It defines the time period between the meeting start time and the meeting end time as a target time period. If any participant's free time includes the target time period, then that participant is identified as a target participant. The ratio between the number of target participants and the total number of participants is defined as a target percentage.

[0073] If the target percentage is less than a preset percentage threshold, and / or the target participants do not include participants whose priority in the participant information is the target priority, the meeting start time and meeting end time will be adjusted at least once until the target percentage is greater than or equal to the percentage threshold, and the target participants include participants whose priority in the participant information is the target priority. Based on the adjusted meeting start time and meeting end time, the initial schedule entry will be updated to obtain intermediate schedule entries. The office location information of the target participants and the available meeting room information within the target time period will be obtained. Based on the office location information and the meeting room information, the target meeting room will be determined. The information of the target meeting room will be updated to the intermediate schedule entry to obtain the final schedule entry. The final schedule entry will be synchronized to the schedules of multiple participants in the meeting information. If there are pre-configured remote participants in the meeting information, a meeting invitation email and a meeting link will be generated based on the final schedule entry. The meeting link will be embedded in the meeting invitation email, and a meeting request email with the meeting link embedded will be sent to the preset email address of the remote participants.

[0074] In some examples of this embodiment, vacant meeting rooms within a target time period are identified as vacant meeting rooms. For any vacant meeting room, the target distance between it and the office location of any target participant is obtained. The average of multiple target distances is calculated to obtain the mean distance, which corresponds one-to-one with a vacant meeting room. The vacant meeting room with the smallest mean distance is identified as the target meeting room. This method improves the rationality of meeting room arrangement and reduces unnecessary movement for target participants.

[0075] Understandably, this embodiment automatically creates calendar entries and adjusts meeting times to ensure that as many people as possible can participate. Furthermore, this embodiment also considers remote participants by automatically generating meeting invitation emails.

[0076] In some examples of this embodiment, after receiving the final schedule, the meeting organizing unit can send it to all participants via email, SMS, or instant messaging. Simultaneously, the system supports personalized notification methods based on user preferences, such as setting SMS reminders for key personnel while others receive notifications via email. The system also features dynamic update and reminder functions. For example, if there are changes to the meeting time or location, the system can automatically update the information and send a correction notification to avoid information oversight.

[0077] In some embodiments, for tasks requiring long-term planning, this system supports generating time management reports such as weekly and monthly reports to help users review and optimize time usage efficiency.

[0078] In some embodiments, the command execution module 120 includes:

[0079] The call unit answers the current call when the target command is an answer call command; when the target command is an automatic reply command, it saves the automatic reply information in the automatic reply command, the automatic reply information including automatic reply text or automatic reply voice, and sends the automatic reply information to the caller when a call request is received.

[0080] When the target command is a message command, if a call request is received, the message function is activated, prompting the caller to start leaving a message. By performing keyword recognition and tone analysis on the message content, the priority of the message content is determined, and based on the priority, a message processing reminder is sent to the user.

[0081] Understandably, the call unit in this embodiment can well meet the user's call needs.

[0082] In some examples of this embodiment, different priorities of message content can correspond to different message processing reminder methods, which can be set according to the actual situation. For example, a prompt sound can be emitted for higher priority messages to provide an immediate reminder and remind the user to handle them as soon as possible; for lower priority messages, a reminder can be sent after a preset time period, etc.

[0083] In some examples of this embodiment, the system also supports converting message content into text information and storing the text information in a preset storage device so that users can view and process it at a convenient time.

[0084] In some embodiments, the command execution module 120 includes:

[0085] The email automation processing unit, when the target command is an email reading command (such as "read the latest email"), displays the content of the email to be read in the email reading command on a preset display interface.

[0086] When the target command is an email reply command (such as "Reply to Zhang San with the content that the meeting arrangements have been confirmed"), the reply content is obtained based on the email reply command, and the reply content is sent back to the sender of the current email;

[0087] When the target command is an email template retrieval command, the target template in the email template retrieval command is retrieved, and an email to be sent is generated based on the target template and the email content input by voice. The recipient's email address is determined according to the recipient information input by voice, and the email to be sent is sent to the recipient's email address.

[0088] In some examples of this embodiment, the system supports automatic email classification and prioritization. Using Natural Language Processing (NLP) and machine learning algorithms, it analyzes email content and historical data to automatically classify emails into categories such as urgent, general, and spam. Users can also define custom rules, such as marking emails from specific senders as high priority for automatic processing. Users can schedule emails to be sent at specific times using target commands, such as scheduling meeting invitations or project progress reports at specific times.

[0089] In some embodiments, the command execution module 120 includes:

[0090] The leave application unit, when the target command is a leave application submission command, calls a preset leave form template, fills the leave time and leave reason in the leave application submission command into the leave form template to obtain the target leave form and completes the submission of the target leave form;

[0091] The leave approval unit determines whether the approval result of the leave application in the currently displayed interface is approved or rejected when the target command is a leave application approval instruction.

[0092] In some embodiments, the command execution module 120 is also suitable for common office processes such as submitting expense reports or project reports, reducing manual input time and the probability of input errors. Approving personnel can complete the approval of leave applications by inputting multimodal data, such as voice data and visual data, or by completing approval decisions through simple gesture operations. This system can track the progress of each process in real time and, when necessary, provide voice prompts to relevant personnel for the next step. In OA (Office Automation) process management, this system can also utilize data analysis functions to help enterprises optimize process design. By analyzing large amounts of process data, it can identify which links are prone to bottlenecks, thereby providing improvement suggestions for enterprises.

[0093] In some embodiments, the command execution module 120 includes:

[0094] The robot scheduling module, when the target command is a robot scheduling command, acquires the status and position information of multiple robots based on the communication channel between the multimodal office assistant system and multiple preset robots; identifies robots with idle status information as selectable robots; and acquires the starting position and destination position information of the object to be transported in the robot scheduling command.

[0095] Based on the starting position information and the position information of the selectable robots, a target robot is determined from the plurality of selectable robots, wherein the distance between the target robot and the object to be transported is less than the distance between the other selectable robots and the object to be transported; a control command is sent to the target robot to instruct the target robot to complete the transport task of the object to be transported.

[0096] For example, when a department urgently needs to replenish office supplies, employees can notify the system via voice commands and / or other interactive methods. The system will then dispatch the nearest logistics robot to pick up and deliver the items. Understandably, the system can rationally schedule tasks based on factors such as task priority, execution time, and the robot's current location.

[0097] In some embodiments, the system further includes: a location service module, which receives employee location information sent by a preset location device, the location device being set in the office area, and the employee location information including the location information of multiple employees;

[0098] The command execution module 120 includes a location query unit. When the target command is a location query command and the voice sender has completed the location query permission verification, the unit sends a query instruction to the location service module based on the employee information to be queried in the location query command, and receives the location information of the employee to be queried from the location service module; and displays the location information of the employee to be queried on a preset display interface or sends it to the terminal device preset by the voice sender.

[0099] In some examples of this embodiment, the positioning device may be a locator with a camera, etc. It is understood that by setting up the aforementioned positioning service module and positioning query module, employee attendance management can be better supported, recording employee arrival and departure times, avoiding the use of traditional clock-in methods, and improving "clocking in" efficiency. In some embodiments, by analyzing employee location information, this system can also obtain the dynamic trends of employee movement, which helps to provide data support for the design and layout of office spaces, further improving the comfort of the office environment.

[0100] Understandably, this system supports multimodal interaction, which can effectively improve user work efficiency and enhance the user experience. It helps drive the development of office systems from automation to intelligence and humanization, and supports enterprise digital transformation. Simultaneously, it promotes the innovation and application of related technologies and drives the development of artificial intelligence in the office field.

[0101] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. A multimodal office assistant system, characterized in that, include: The voice wake-up module receives first multimodal data; based on the first multimodal data, it performs wake-up probability prediction to obtain a wake-up probability prediction value; if the wake-up probability prediction value is greater than the wake-up threshold, it wakes up the multimodal office assistant system, wherein the wake-up threshold is positively correlated with the current ambient noise intensity. The command execution module receives second multimodal data, performs command recognition based on the second multimodal data to determine the target command, and executes the target command. Both the first and second multimodal data include voice data and visual data. The visual data is video data or multi-frame image data containing the image of the voice speaker. The acquisition time of the first multimodal data is earlier than the acquisition time of the second multimodal data.

2. The multimodal office assistant system according to claim 1, characterized in that, The voice wake-up module is specifically used to extract features from the voice data in the first multimodal data to obtain the first voice feature; Feature extraction is performed on the visual data in the first multimodal data to obtain the first lip movement feature and the first limb movement feature of the speaker; Based on the first speech feature, the first lip movement feature, and the first limb movement feature, multimodal feature information fusion is performed to obtain the target feature; Based on the target characteristics, a wake-up probability prediction is performed to obtain the predicted wake-up probability value.

3. The multimodal office assistant system according to claim 1, characterized in that, The system further includes: an identity recognition module, which receives target multimodal data before the command execution module receives the second multimodal data; and performs voiceprint recognition based on the voice data in the target multimodal data to determine the identity of the voice speaker; Based on the visual data in the target multimodal data, facial recognition is performed on the speaker to obtain a facial recognition result; based on the fingerprint data in the target multimodal data, fingerprint recognition is performed on the speaker to obtain a fingerprint recognition result; based on the facial recognition result and the fingerprint recognition result, the speaker is authenticated. If the verification is successful, the target operating environment information corresponding to the current voice speaker is obtained based on the preset operating environment mapping relationship. The operating environment mapping relationship is the mapping relationship between the user and the operating environment information. The operating environment information includes the working interface layout, language setting information, and preference parameters. Based on the target operating environment information, the preset display interface is adjusted.

4. The multimodal office assistant system according to claim 1, characterized in that, The command execution module includes: The schedule management unit, when the target command is a create schedule command, generates schedule entries based on the schedule information in the create schedule command, and synchronizes the schedule entries to the current user's schedule. A calendar synchronization unit, which, when the target command is a schedule synchronization command, synchronizes the user's schedule from the original device to one or more target devices; The event reminder module sends schedule reminders to the user based on the schedule entries in the schedule and a preset first reminder method; if there is a time conflict in the schedule entries, a schedule conflict reminder is sent to the user according to a preset second reminder method.

5. The multimodal office assistant system according to claim 1, characterized in that, The command execution module includes: The meeting organization unit, when the target command is a meeting organization command, generates initial schedule entries based on the meeting information in the meeting organization command. The meeting information includes: meeting name, meeting start time, meeting end time, and participant information, including participant name and priority for each participant. Based on the participant names, it obtains the schedules of multiple participants to determine each participant's free time. It defines the time period between the meeting start time and the meeting end time as a target time period. If any participant's free time includes the target time period, then that participant is identified as a target participant. The ratio between the number of target participants and the total number of participants is defined as a target percentage. If the target percentage is less than a preset percentage threshold, and / or the target participants do not include participants whose priority in the participant information is the target priority, the meeting start time and meeting end time will be adjusted at least once until the target percentage is greater than or equal to the percentage threshold, and the target participants include participants whose priority in the participant information is the target priority. Based on the adjusted meeting start time and meeting end time, the initial schedule entry will be updated to obtain intermediate schedule entries. The office location information of the target participants and the available meeting room information within the target time period will be obtained. Based on the office location information and the meeting room information, the target meeting room will be determined. The information of the target meeting room will be updated to the intermediate schedule entry to obtain the final schedule entry. The final schedule entry will be synchronized to the schedules of multiple participants in the meeting information. If there are pre-configured remote participants in the meeting information, a meeting invitation email and a meeting link will be generated based on the final schedule entry. The meeting link will be embedded in the meeting invitation email, and a meeting request email with the meeting link embedded will be sent to the preset email address of the remote participants.

6. The multimodal office assistant system according to claim 1, characterized in that, The command execution module includes: The call unit answers the current call when the target command is an answer call command; when the target command is an automatic reply command, it saves the automatic reply information in the automatic reply command, the automatic reply information including automatic reply text or automatic reply voice, and sends the automatic reply information to the caller when a call request is received. When the target command is a message command, if a call request is received, the message function is activated, prompting the caller to start leaving a message. By performing keyword recognition and tone analysis on the message content, the priority of the message content is determined, and based on the priority, a message processing reminder is sent to the user.

7. The multimodal office assistant system according to claim 1, characterized in that, The command execution module includes: The email automation processing unit, when the target command is an email reading command, displays the content of the email to be read in the email reading command on a preset display interface; If the target command is an email reply command, the reply content is obtained based on the email reply command, and the reply content is sent back to the sender of the current email; When the target command is an email template retrieval command, the target template in the email template retrieval command is retrieved, and an email to be sent is generated based on the target template and the email content input by voice. The recipient's email address is determined according to the recipient information input by voice, and the email to be sent is sent to the recipient's email address.

8. The multimodal office assistant system according to claim 1, characterized in that, The command execution module includes: The leave application unit, when the target command is a leave application submission command, calls a preset leave form template, fills the leave time and leave reason in the leave application submission command into the leave form template to obtain the target leave form and completes the submission of the target leave form; The leave approval unit determines whether the approval result of the leave application in the currently displayed interface is approved or rejected when the target command is a leave application approval instruction.

9. The multimodal office assistant system according to claim 1, characterized in that, The command execution module includes: The robot scheduling module, when the target command is a robot scheduling command, acquires the status and position information of multiple robots based on the communication channel between the multimodal office assistant system and multiple preset robots; identifies robots with idle status information as selectable robots; and acquires the starting position and destination position information of the object to be transported in the robot scheduling command. Based on the starting position information and the position information of the selectable robots, a target robot is determined from the plurality of selectable robots, wherein the distance between the target robot and the object to be transported is less than the distance between the other selectable robots and the object to be transported; a control command is sent to the target robot to instruct the target robot to complete the transport task of the object to be transported.

10. The multimodal office assistant system according to claim 1, characterized in that, The system also includes a location service module, which receives employee location information sent by a preset location device, the location device being set up in the office area, and the employee location information including the location information of multiple employees; The command execution module includes a location query unit. When the target command is a location query command and the voice sender has completed the location query permission verification, the unit sends a query instruction to the location service module based on the employee information to be queried in the location query command, and receives the location information of the employee to be queried from the location service module; and displays the location information of the employee to be queried on a preset display interface or sends it to the terminal device preset by the voice sender.

Citation Information

Patent Citations

  • Voiceprint recognition method and device, storage medium and loudspeaker box

    CN108766446A

  • Intelligent equipment awakening method and device, and intelligent equipment registration method and device

    CN111179941A

  • Voiceprint awakening method, device and equipment and storage medium

    CN111223490A

  • Instruction input method and system for vehicle, storage medium and vehicle

    CN113157080A

  • Display device and voice wake-up method

    CN118366434A