Task processing method and device, electronic equipment and computer readable storage medium

By recognizing user intent and converting audio and video tasks into appropriate execution components, the problem of limited audio channel resources in augmented reality devices is solved, improving user experience and system stability.

CN120705639BActive Publication Date: 2025-12-23FALCON INNOVATIONS TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511142798.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-12-23
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing smart wearable devices such as augmented reality glasses suffer from limited audio channel resources when performing multiple audio tasks, leading to task interruptions and a degraded user experience.

Method used

By recognizing user intent, audio and video tasks are converted to each other, with audio tasks executed by the audio output module and video tasks executed by the virtual screen, thus avoiding information overload.

Benefits of technology

This improves system stability and user experience, allowing users to obtain task information from both visual and auditory perspectives, thus avoiding information confusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705639B_ABST
    Figure CN120705639B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a task processing method and device, electronic equipment and computer readable storage medium, and relate to the technical field of computers; the method is applied to an extended reality device, the extended reality device comprises a virtual screen and an audio output module; the method comprises: when a first task of a first category and a second task exist simultaneously, identifying a user intention; according to the user intention, determining a task to be converted and an original category task from the first task and the second task; converting the task to be converted into a task of a second category; executing the original category task and executing the task of the second category; wherein the task of the first category is executed by one of the virtual screen and the audio output module, and the task of the second category is executed by the other of the virtual screen and the audio output module. In this way, the audio task and the video task can be converted in the present application, so that the audio channel / video channel information overload can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of computer, in particular to a task processing method and device, electronic equipment and computer readable storage medium. BACKGROUND

[0002] With the wide application of smart wearable devices such as augmented reality (AR) glasses, users can use voice to perform tasks such as navigation, reminders, and answering phone calls. However, these devices generally have the problem of single audio channel resource. When there are multiple audio tasks, other audio tasks are often interrupted to retain only one audio task, and the interruption of audio tasks affects the user experience. For example, when a user is listening to music, if a phone call is suddenly received, in order to avoid information overload of the audio channel, the music playback is usually interrupted, and the interruption of music playback affects the user experience. SUMMARY

[0003] Embodiments of the present application provide a task processing method and device, electronic equipment and computer readable storage medium. Based on the technical solution of the embodiments of the present application, audio tasks and video tasks can be converted to each other, thereby avoiding information overload of the audio channel / video channel.

[0004] In a first aspect, embodiments of the present application provide a task processing method applied to an extended reality device, the extended reality device including a virtual screen and an audio output module; the method comprising:

[0005] In response to triggering a first task of a first category, determining whether a second task of the first category is currently being executed;

[0006] When the second task of the first category is currently being executed, identifying a user intention;

[0007] According to the user intention, determining a to-be-converted task and an original category task from the first task and the second task;

[0008] Converting the to-be-converted task into a task of a second category;

[0009] Executing the original category task and the task of the second category.

[0010] Wherein, the task of the first category is executed by one of the virtual screen and the audio output module, and the task of the second category is executed by the other of the virtual screen and the audio output module.

[0011] In one of the embodiments, the identifying a user intention comprises:

[0012] obtaining first information of a first task of the first category and second information of a second task of the first category;

[0013] obtaining user behavior and event context;

[0014] obtaining a user intent recognition model;

[0015] analyzing the first information, the second information, the user behavior and the event context based on the user intent recognition model, and determining the user intent.

[0016] In one of the embodiments, the first category of tasks is an audio task; and the recognizing the user intent while the second task of the first category is currently being executed comprises:

[0017] detecting whether the audio output module has a left-right channel independent playing capability while the second task of the first category is currently being executed;

[0018] recognizing the user intent when the audio output module does not have the left-right channel independent playing capability.

[0019] In one of the embodiments, the method further comprises:

[0020] playing audio corresponding to any one of the first and second tasks of the first category via the left channel when the audio output module has the left-right channel independent playing capability;

[0021] playing audio corresponding to the other of the first and second tasks of the first category via the right channel.

[0022] In one of the embodiments, when the task to be converted is a voice call task, the converting the task to be converted into a task of a second category comprises:

[0023] converting audio of the voice call into text;

[0024] generating a call summary according to the text;

[0025] the executing the task of the second category comprises:

[0026] displaying the text and / or the call summary via the virtual screen.

[0027] In one of the embodiments, when the task to be converted is a voice navigation task, the converting the task to be converted into a task of a second category comprises:

[0028] obtaining a navigation route map corresponding to the voice navigation task;

[0029] acquire a current position in real time;

[0030] generate navigation instruction text according to the current position and the navigation map;

[0031] The execution of the second category of tasks includes:

[0032] display the navigation map and the navigation instruction text in real time by using the virtual screen.

[0033] In one of the embodiments, when the task to be converted is a navigation display task, the conversion of the task to be converted into a second category of tasks includes:

[0034] acquire a navigation map corresponding to the navigation display task;

[0035] acquire a current position in real time;

[0036] generate navigation instruction speech according to the current position and the navigation map;

[0037] The execution of the second category of tasks includes:

[0038] play the navigation instruction speech by using the audio output module.

[0039] In a second aspect, the embodiments of the present application provide a task processing apparatus applied to an extended reality device, the extended reality device including a virtual screen and an audio output module; the apparatus includes:

[0040] a task triggering module configured to determine whether a second task of a first category is being executed currently in response to triggering a first task of the first category;

[0041] an intention recognition module configured to recognize a user intention when the second task of the first category is being executed currently;

[0042] a task determination module configured to determine a task to be converted and an original category task from the first task and the second task according to the user intention;

[0043] a task conversion module configured to convert the task to be converted into a second category of tasks;

[0044] a task execution module configured to execute the original category task and the second category of tasks.

[0045] The first category of tasks is executed by one of the virtual screen and the audio output module, and the second category of tasks is executed by the other of the virtual screen and the audio output module.

[0046] In one of the embodiments, the intent recognition module is specifically configured to perform:

[0047] obtain first information of the first task of the first category and second information of the second task of the first category;

[0048] obtain user behavior and event context;

[0049] obtain a user intent recognition model;

[0050] analyze the first information, the second information, the user behavior and the event context based on the user intent recognition model, and determine the user intent.

[0051] In one of the embodiments, the task of the first category is an audio task; and the intent recognition module is specifically configured to perform:

[0052] detect whether the audio output module has left-right channel independent playing capability when the second task of the first category is currently being performed;

[0053] recognize the user intent when the audio output module does not have the left-right channel independent playing capability.

[0054] In one of the embodiments, the apparatus further comprises:

[0055] a left channel playing module configured to play audio corresponding to any one of the first task and the second task of the first category by using a left channel when the audio output module has the left-right channel independent playing capability;

[0056] a right channel playing module configured to play audio corresponding to the other of the first task and the second task of the first category by using a right channel.

[0057] In one of the embodiments, when the task to be converted is a voice call task, the task conversion module is specifically configured to perform:

[0058] convert audio of the voice call into text;

[0059] generate a call summary according to the text;

[0060] the task execution module is specifically configured to perform:

[0061] display the text and / or the call summary by using the virtual screen.

[0062] In one of the embodiments, when the task to be converted is a voice navigation task, the task conversion module is specifically configured to perform:

[0063] acquire a navigation route map corresponding to the voice navigation task;

[0064] acquire a current position in real time;

[0065] generate a navigation instruction text according to the current position and the navigation route map;

[0066] The task execution module is specifically configured to perform:

[0067] display the navigation route map and the navigation instruction text in real time by using the virtual screen.

[0068] In one of the embodiments, when the task to be converted is a navigation display task, the task conversion module is specifically configured to perform:

[0069] acquire a navigation route map corresponding to the navigation display task;

[0070] acquire a current position in real time;

[0071] generate a navigation instruction voice according to the current position and the navigation route map;

[0072] The task execution module is specifically configured to perform:

[0073] play the navigation instruction voice by using the audio output module.

[0074] In a third aspect, the embodiments of the present application further provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is executed by the processor to implement the steps in the task processing method.

[0075] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps in the task processing method.

[0076] In a fifth aspect, the embodiments of the present application further provide a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the method provided in the various optional implementation manners of the embodiments of the present application.

[0077] The embodiments of the present application have the following beneficial effects:

[0078] In response to triggering the first task of the first category, it can be determined whether the second task of the first category is currently being executed, when the second task of the first category is currently being executed, it can be determined that the first task and the second task of the first category exist simultaneously; because the tasks of the first category are executed by the same component (the virtual screen or the audio output module), if the first task and the second task of the first category are directly executed based on the component, information overload of the component will occur, and the user cannot distinguish the information of the two tasks; therefore, the user intention can be identified, according to the user intention, the task to be converted and the original category task are determined from the first task and the second task, and the task to be converted is converted into a task of the second category; in this way, the original category task (the task of the first category) and the task of the second category can be executed by the virtual screen and the audio output module respectively, the information overload of the single component is avoided, the functions of the two components are fully utilized, the processing efficiency is improved, and the system stability is enhanced; and the user can obtain the information of the two tasks from the two dimensions of vision and hearing, the problem that the user cannot distinguish the information of the two tasks is avoided, and the use experience of the user is improved. BRIEF DESCRIPTION OF DRAWINGS

[0079] In order to more clearly illustrate the technical solutions in the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0080] Figure 1 is a step schematic diagram of a task processing method provided by an embodiment of the present application;

[0081] Figure 2 is a structure schematic diagram of a task processing system provided by an embodiment of the present application;

[0082] Figure 3 is a flow schematic diagram of a task processing method provided by an embodiment of the present application;

[0083] Figure 4 is a schematic diagram of a virtual screen display interface provided by an embodiment of the present application;

[0084] Figure 5 is a structure schematic diagram of a task processing device provided by an embodiment of the present application;

[0085] Figure 6 is a structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0086] The technical solutions in the present application will be described clearly and completely in the present application by combining the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0087] In one embodiment, as shown in Figure 1 A task processing method is provided, although a logical sequence is shown in the step diagram, in some cases, the steps shown or described can be performed in an order different from that shown in the figure. Specifically, the task processing method can be applied to an extended reality (XR) device. Extended reality includes virtual reality (VR), augmented reality, and mixed reality (MR). The extended reality device can have an independent operating system, support users to install software, game applications, etc., can be controlled by voice or action to complete functions such as adding schedules, map navigation, interacting with friends, taking photos and videos, video calling with friends, and task processing, and can realize wireless network access through a mobile communication network.

[0088] The extended reality device can include a mobile handheld terminal such as a mobile phone, a tablet, etc., include a heads-up display such as a car head-up display, and include a near-eye display terminal worn on the head of a human body, such as a display in the form of glasses, a display in the form of a helmet, etc. In the embodiments of the present application, the hardware structure of the extended reality device is described by taking the optical see-through glasses display, i.e., the extended reality glasses, as an example.

[0089] The extended reality glasses include a frame body, i.e., a temple and a frame, to support the extended reality glasses worn on the head of a human body; include an optical display module mainly including a micro display screen, an optical lens, and an optical waveguide sheet to display virtual content to a user; include an audio module mainly including a microphone and a speaker to collect and play sound; include a sensor mainly including a camera, a gyroscope, a barometer, and an infrared light emitter-receiver to collect information data related to the human body, the glasses body, and the external environment; include an integrated processor mainly including a microcontroller unit (MCU) or a central processing unit (CPU) to process data and calculate; include a circuit board, which is flexible or rigid, to connect other electronic components to form an electronic circuit system, which is generally placed in the inner cavity of the frame and the temple; and can further include a battery power supply.

[0090] In an embodiment, the extended reality device can include a virtual screen and an audio output module. The virtual screen can display virtual elements, which can be various visual elements such as text, graphics, and icons. The audio output module can output any audio. The virtual screen of the extended reality device is not a physically tangible physical screen in the traditional sense, but is a display area that can display various virtual information, images, videos, and interactive interfaces, etc. that are generated by technical means and perceived by the user through light, images, and other elements. The virtual screen of the extended reality device can have a large enough field of view, which determines the size of the range covered by the virtual screen in the user's field of view. A larger field of view can allow the user to experience a wider virtual space. The audio output module can be a headset, a speaker, a micro speaker, or an array speaker, etc.

[0091] The following are described in detail respectively. It should be noted that the order of the following embodiments is not limited as the priority order of the embodiments.

[0092] According to Figure 1 The task processing method shown in the figure includes at least steps S110 to S150, which are described in detail as follows:

[0093] In step S110, in response to triggering a first task of a first category, it is determined whether a second task of the first category is currently being executed.

[0094] In step S120, when the second task of the first category is currently being executed, a user intention is identified.

[0095] In step S130, according to the user intention, a task to be converted and an original category task are determined from the first task and the second task.

[0096] In step S140, the task to be converted is converted into a task of a second category.

[0097] In step S150, the original category task is executed, and the task of the second category is executed.

[0098] Among them, the task of the first category is executed by one of the virtual screen and the audio output module, and the task of the second category is executed by the other of the virtual screen and the audio output module. The task of the first category is one of an audio task and a video task, and the task of the second category is the other of the audio task and the video task. The audio task is executed by the audio output module, and the video task is executed by the virtual screen.

[0099] The audio task can include but is not limited to one or more of voice call, voice navigation, voice reminder, and playing music. The video task can include but is not limited to one or more of navigation display, text display, and video display.

[0100] When the first task of the first category is detected, it can be detected whether a second task of the same category (the first category) is currently being executed. When the second task of the first category is currently being executed, it can be determined that there are multiple tasks of the first category at the same time, and the channel (the video channel or the audio channel) corresponding to the first category can have information overload. If the first task and the second task are executed through the component (the virtual screen or the audio output module) corresponding to the first category at the same time, the user can be confused by the information of the first task and the second task.

[0101] Therefore, when the second task of the first category is currently being executed, the user intention can be recognized, and the user intention represents which one of the first task or the second task the user tends to convert into a task of the second category. The user intention can be determined based on a user intention recognition model. The user intention recognition model can be an intention recognition large model. The user intention recognition model can predict the user intention based on the first information of the first task, the second information of the second task, user behavior, and event context.

[0102] According to the user intention, one of the first task and the second task can be determined as a to-be-converted task, and the other can be determined as a task of the original category. The to-be-converted task refers to a task that needs to be converted into another category (the second category), and the task of the original category refers to a task that does not need to be converted.

[0103] After the to-be-converted task is determined, the information of the to-be-converted task can be converted into information that can be output by the component (the virtual screen or the audio output module) corresponding to the second category, and then the to-be-converted task is executed through the component corresponding to the second category; and the task of the original category is directly executed through the component corresponding to the first category.

[0104] For example, the task of the first category is an audio task, and the task of the second category is a video task. The to-be-converted task is a voice navigation task of the first category, which can be converted into a navigation display task of the second category, and the navigation map and the navigation instruction text can be displayed through the virtual screen.

[0105] For another example, the task of the first category is a video task, and the task of the second category is an audio task. The to-be-converted task is a navigation display task of the first category, which can be converted into a voice navigation task of the second category, and the navigation instruction voice can be played through the audio output module.

[0106] The technical solution of the embodiment of the present application can determine whether the second task of the first category is being executed at present in response to the first task of the first category being triggered, and can determine that the first task and the second task of the first category exist at the same time at present when the second task of the first category is being executed at present. Because the tasks of the first category are executed by the same component (the virtual screen or the audio output module), if the first task and the second task of the first category are directly executed based on the component, information overload of the component will occur, and the user cannot distinguish the information of the two tasks. Therefore, the user intention can be recognized, the task to be converted and the original category task can be determined from the first task and the second task according to the user intention, and the task to be converted is converted into a task of the second category. In this way, the original category task (the task of the first category) and the task of the second category can be executed by the virtual screen and the audio output module respectively, the information overload of the single component is avoided, the functions of the two components are fully utilized, the processing efficiency is improved, and the system stability is enhanced. The user can obtain the information of the two tasks from two dimensions of vision and hearing, the problem that the user cannot distinguish the information of the two tasks is avoided, and the use experience of the user is improved.

[0107] On the basis of the above technical solution, as an embodiment, the extended reality device can include a task processing system, and the execution of the task processing method is controlled by the task processing system. Figure 2 FIG. 1 is a structural schematic diagram of a task processing system provided by an embodiment of the present application. Referring to FIG. 1, Figure 2 The task processing system can include a voice navigation task module, a call detection module, a sound channel state recognition module, an intention recognition model, a task scheduling control module, an artificial intelligence (AI) assistant module, a dual sound channel control module, and an AR interface display module. The voice navigation task module can perform a voice navigation task. The call detection module can detect whether there is a call. The sound channel state recognition module can detect whether the sound channel has the ability to independently play left and right sound channels, and can also detect whether the sound channel is currently executing an audio task. The intention recognition model can recognize user intention. The task scheduling control module can schedule tasks, including converting tasks. The AI assistant module can automatically answer a call, and can convert the audio of a voice call into text and automatically generate a call summary. The dual sound channel control module can control the output of each of the two sound channels. The AR interface display module can control the virtual screen to display an AR interface.

[0108] On the basis of the above technical solutions, as an embodiment, the identifying the user intent can include: obtaining first information of a first task of the first category and second information of a second task of the first category; obtaining user behavior and event context; obtaining a user intent recognition model; and analyzing the first information, the second information, the user behavior, and the event context based on the user intent recognition model to determine the user intent.

[0109] The information of the task (the first information and / or the second information) can include various information of the task. In an embodiment, the types of information of different tasks that need to be obtained for identifying the user intent can be pre-configured, and when obtaining the information of the task for identifying the user intent, the information of the task can be obtained according to the configuration.

[0110] In an embodiment, the information of the voice call task can include, but is not limited to, one or more of whether a call number is for takeout or express delivery, whether the call number is for an important contact, whether the call number has ever called, and historical processing (answering or hanging up) of the call number that has ever called.

[0111] In an embodiment, the information of the voice navigation task can include, but is not limited to, one or more of a starting position, an ending position, a current position, a planned navigation route, and real-time navigation instructions.

[0112] The user behavior can include, but is not limited to, one or more of gesture behavior, eye movement behavior, voice behavior (such as speaking different voice instructions), and head behavior (such as head swinging, nodding, shaking, etc.).

[0113] The event context can include, but is not limited to, one or more of background information, environmental conditions, cause and effect, and associated factors surrounding the event. The event context constitutes the overall situation of understanding the event. The event context helps to fully and accurately grasp the meaning, logic, and impact of the event, etc. Without the event context, the event can appear isolated, ambiguous, and even misinterpreted.

[0114] The event context can include, but is not limited to, one or more of time information, space information, cause and effect, and associated subject and relationship. The time information can include a specific time (such as a season, a time) at which the event occurs, and an era characteristic (such as a technical level) of the time. The space information can include a location (such as a region, a specific place) at which the event occurs, and a physical environment, a cultural atmosphere, a geographical characteristic, and the like of the space. For example, when navigating, the space information can include information about whether the current position of the user is a fork in the road, whether there is an obstacle, and the like. The cause and effect can include a reason for the occurrence of the event and a subsequent influence. For example, if the audio task is converted into a video task, the subsequent influence can include the allocation of the AR interface, the determination of the display position, and the like. The associated subject and relationship can include a person associated with the event, a relationship of the person with the wearer of the extended reality device, and the like. For example, the associated subject and relationship in the voice call event can include a user at the other end of the voice call, an identity of the user at the other end, and the like.

[0115] The first information, the second information, the user behavior, and the event context can be input into a user intent recognition model, which can determine, based on the input information, which task the user currently intends to convert in the category, and thus determine the task to be converted and the original category task from the first task and the second task.

[0116] The user intent recognition model can be obtained based on a large number of labeled training samples through supervised training, or obtained based on a large number of samples through deep learning. The training of the user intent recognition model can refer to related technologies.

[0117] By adopting the technical solutions of the embodiments of the present application, the user intent can be quickly recognized based on the user intent recognition model, the user intent can be predicted based on the user behavior, the predicted user intent can be more consistent with the user's idea, the user intent can be predicted based on the event context, the entire event can be more clearly recognized, the user intent can be predicted based on the first information and the second information, and the information of the first task and the second task can be fully considered, so that the predicted user intent can be more accurate.

[0118] On the basis of the above technical solutions, as an embodiment, the first type of task is an audio task. The identifying the user intention while the second task of the first type is being executed currently comprises: detecting whether the audio output module has the ability of independent playing of left and right channels while the second task of the first type is being executed currently; and identifying the user intention when the audio output module does not have the ability of independent playing of left and right channels. When the audio output module has the ability of independent playing of left and right channels, the left channel is used to play audio corresponding to any one of the first and second tasks of the first type, and the right channel is used to play audio corresponding to the other of the first and second tasks of the first type.

[0119] The first type of task is an audio task, and when the second task of the first type is being executed currently, it can be determined that two audio tasks exist simultaneously. Whether the audio output module has the ability of independent playing of left and right channels can be determined by a channel state identifying module. The ability of independent playing of left and right channels means that the left and right channels of the audio output module can respectively receive, process and output different audio signals, and realize independent sound emission without mutual interference.

[0120] If the audio output module does not have the ability of independent playing of left and right channels, and two audio tasks are executed simultaneously by using the audio output module, the user cannot distinguish the content of the two audio tasks, and therefore, the user intention needs to be identified, and one of the audio tasks is converted into a video task, so as to reduce the pressure of audio processing and audio output, avoid overload of audio channel information, and enable the user to obtain information of the two tasks from two dimensions of hearing and vision.

[0121] If the audio output module has the ability of independent playing of left and right channels, the left and right channels can be used to respectively play audio of the two audio tasks. For example, the left channel can be used to play audio of the first task, and the right channel can be used to play audio of the second task, or the left channel can be used to play audio of the second task, and the right channel can be used to play audio of the first task.

[0122] In one of the embodiments, if the audio output module has the ability of independent playing of left and right channels, it can be determined whether a video task is being executed currently. If a video task is being executed currently, the left and right channels can be used to respectively play audio of the two audio tasks. If a video task is not being executed currently, the user intention can still be identified to convert one of the audio tasks into a video task.

[0123] The two ears of the user respectively receive the left channel and the right channel sound, instead of the same ear receiving two audio task mixed sound, so that the user can distinguish the information of the two audio tasks when playing the respective audio of the two audio tasks by using the left and right different channels. In this way, more resources can be reserved for the video task.

[0124] Figure 3 is a flowchart of a task processing method provided by an embodiment of the present application; refer to Figure 3 The first type of task is an audio task. When the user uses the AR glasses for voice navigation, voice channel state recognition can be performed to determine whether there is currently an audio task running. If an incoming call is detected, a user intention recognition model in the user intention recognition module can be called to determine the user intention. It can be determined whether dual sound is supported (whether it has the ability to play independently by left and right channels). When dual sound is supported, channel shunting control can be performed to control two channels to perform different audio tasks, for example, the left channel can be used for voice navigation and the right channel can be used for voice call. When dual sound is not supported, an AI assistant can be used to determine the processing strategy, and determine whether the caller is a takeout, express delivery or important contact person. The AI assistant can automatically answer the phone and generate a voice summary, and can also push related notifications (such as voice summaries) to the AR interface for display. The voice navigation task can also be converted into a navigation display task, and a multi-modal AR interface (i.e. a virtual screen display interface) can be used for multi-modal display. Figure 4 is a schematic diagram of a virtual screen display interface provided by an embodiment of the present application; as shown in Figure 4 The virtual screen display interface can display a navigation picture in the navigation map area, and can also display an AI-generated voice call summary in a bubble, and can also display a gesture control prompt. The gesture control prompt can include continue navigation, hang up, and redial, etc.

[0125] On the basis of the above technical solutions, as an embodiment, when the task to be converted is a voice call task, the converting the task to be converted into a second type of task can include: converting the audio of the voice call into text; generating a call summary according to the text. The executing the second type of task can include: displaying the text and / or the call summary by using the virtual screen.

[0126] When multiple audio tasks exist at the same time, and the voice call task among them is determined to be a task to be converted, the voice call task needs to be converted into a video task.

[0127] The AI assistant can automatically answer the voice call to obtain the audio of the voice call, which can include the audio of the caller and the audio of the AI assistant. The audio of the voice call can be converted into text, which can include the identity of the speaker for each sentence. The text can be subjected to semantic analysis to generate a call summary.

[0128] When the AI assistant automatically answers the voice call, the audio of the voice call can be obtained at a volume of 0 or a lower volume, and the information of the audio can be directly obtained to generate text and a call summary based on the information of the audio, thereby avoiding the audio of the voice call affecting another audio task.

[0129] The virtual screen can display the text converted from the audio of the voice call and / or the call summary. Whether to display the text or the call summary, or both, can be pre-set by the user.

[0130] The technical solutions of the embodiments of the present application can convert a voice call task into text and generate a call summary, and then display the text and / or the call summary on a virtual screen. In this way, the user can obtain information of the voice call task by watching the content displayed on the virtual screen, and the audio output module can output another audio task, thereby avoiding the voice call task affecting another audio task.

[0131] On the basis of the above technical solutions, as an embodiment, when the task to be converted is a voice navigation task, the conversion of the task to be converted into a second type of task can include: obtaining a navigation route map corresponding to the voice navigation task; obtaining a current position in real time; and generating navigation instruction text according to the current position and the navigation route map. The execution of the second type of task can include: displaying the navigation route map on the virtual screen and displaying the navigation instruction text in real time.

[0132] When multiple audio tasks exist at the same time and a voice navigation task is determined to be a task to be converted, the voice navigation task needs to be converted into a video task.

[0133] The navigation route map corresponding to the voice navigation task can be obtained. The navigation route map can include a map, a starting position, an end position, and a navigation route from the starting position to the end position. The current position of the user can be obtained in real time, and the current position of the user on the navigation route can be determined according to the current position of the user and the navigation route map, and then it can be determined how the user should act at the current position, such as whether to go straight, turn left, or turn right. According to how the user should act, the navigation instruction can be determined, and the navigation instruction can be displayed in text to obtain navigation instruction text.

[0134] The navigation route map can be displayed by using the virtual screen, and the navigation instruction text corresponding to the current position can be displayed in real time.

[0135] By using the technical solution of the embodiment of the application, when multiple audio tasks exist at the same time, the voice navigation task can be converted into a video task, and then the navigation route map and the navigation instruction text are displayed by using the virtual screen. In this way, the user can obtain navigation information by watching the content displayed on the virtual screen, and the audio output module outputs another audio task, so as to avoid the voice navigation task affecting another audio task.

[0136] On the basis of the above technical solution, as an embodiment, when the task to be converted is a navigation display task, the conversion of the task to be converted into a task of the second category can include: obtaining a navigation route map corresponding to the navigation display task; obtaining a current position in real time; generating a navigation instruction voice according to the current position and the navigation route map. The execution of the task of the second category can include: playing the navigation instruction voice by using the audio output module.

[0137] When multiple video tasks exist at the same time, and the navigation display task among them is determined as a task to be converted, the navigation display task needs to be converted into an audio task.

[0138] The navigation route map corresponding to the navigation display task can be obtained. The navigation route map can include a map, a starting position, an ending position, and a navigation route from the starting position to the ending position. The current position of the user can be obtained in real time. According to the current position of the user and the navigation route map, the current position of the user in the navigation route can be determined, and then it can be determined how the user should act at the current position, such as whether to go straight, turn left, and turn right, etc. According to how the user should act, the navigation instruction can be determined, the navigation instruction voice can be obtained by using the voice to represent the navigation instruction. The navigation instruction voice can be played by using the audio output module.

[0139] When the user hears the navigation instruction voice, the user can determine how to act at the current position according to the navigation instruction voice, such as whether to go straight, turn left, and turn right, etc.

[0140] By using the technical solution of the embodiment of the application, when multiple video tasks exist at the same time, the navigation display task can be converted into an audio task, and then the navigation instruction voice is output by using the audio output module. In this way, the user can determine how to act at the current position according to the navigation instruction voice, and the virtual screen displays other video tasks, so as to avoid the display area of the virtual screen being unable to display multiple video tasks, and the user being unable to clearly see the information of the video tasks.

[0141] To facilitate better implementation of the task processing method of the present application, the present application also provides a task processing device based on the above task processing method. The meanings of the terms are the same as in the above task processing method, and the specific implementation details can be referred to the description in the method embodiment.

[0142] Please refer to Figure 5 , Figure 5 is a structural schematic diagram of a task processing device provided by an embodiment of the present application, wherein the task processing device is applied to an extended reality device, and the extended reality device includes a virtual screen and an audio output module. The task processing device includes:

[0143] The task triggering module 501 is configured to determine whether a second task of the first category is being executed currently in response to triggering a first task of the first category.

[0144] The intent recognition module 502 is configured to recognize a user intent when the second task of the first category is being executed currently.

[0145] The task determination module 503 is configured to determine a task to be converted and an original category task from the first task and the second task according to the user intent.

[0146] The task conversion module 504 is configured to convert the task to be converted into a task of a second category.

[0147] The task execution module 505 is configured to execute the original category task and the task of the second category.

[0148] The task of the first category is executed by one of the virtual screen and the audio output module, and the task of the second category is executed by the other of the virtual screen and the audio output module.

[0149] In one embodiment, the intent recognition module 502 is specifically configured to perform:

[0150] Obtain first information of the first task of the first category and second information of the second task of the first category.

[0151] Obtain user behavior and event context.

[0152] Obtain a user intent recognition model.

[0153] Analyze the first information, the second information, the user behavior and the event context based on the user intent recognition model to determine the user intent.

[0154] In one embodiment, the task of the first category is an audio task, and the intent recognition module 502 is specifically configured to perform:

[0155] detecting whether the audio output module has left-right channel independent playing capability when a second task of the first category is being executed;

[0156] recognizing the user intention when the audio output module does not have the left-right channel independent playing capability.

[0157] In one of the embodiments, the apparatus further comprises:

[0158] a left channel playing module configured to play audio corresponding to any one of the first task and the second task of the first category by using a left channel when the audio output module has the left-right channel independent playing capability;

[0159] a right channel playing module configured to play audio corresponding to the other one of the first task and the second task of the first category by using a right channel.

[0160] In one of the embodiments, when the task to be converted is a voice call task, the task conversion module 504 is specifically configured to perform:

[0161] convert audio of the voice call into text;

[0162] generate a call summary according to the text;

[0163] The task execution module 505 is specifically configured to perform:

[0164] display the text and / or the call summary by using the virtual screen.

[0165] In one of the embodiments, when the task to be converted is a voice navigation task, the task conversion module 504 is specifically configured to perform:

[0166] obtain a navigation route map corresponding to the voice navigation task;

[0167] obtain a current position in real time;

[0168] generate a navigation instruction text according to the current position and the navigation route map;

[0169] The task execution module 505 is specifically configured to perform:

[0170] display the navigation route map by using the virtual screen, and display the navigation instruction text in real time.

[0171] In one of the embodiments, when the task to be converted is a navigation display task, the task conversion module 504 is specifically configured to perform: obtain a navigation route map corresponding to the navigation display task;

[0172] acquire a current position in real time;

[0173] generate a navigation instruction voice according to the current position and the navigation roadmap;

[0174] The task execution module 505 is specifically configured to execute:

[0175] The audio output module is configured to play the navigation instruction voice.

[0176] According to the technical solutions of the embodiments of the present application, in response to triggering the first task of the first category, it can be determined whether the second task of the first category is being executed at present. When the second task of the first category is being executed at present, it can be determined that the first task and the second task of the first category exist simultaneously at present. Because the tasks of the first category are executed by the same component (the virtual screen or the audio output module), if the first task and the second task of the first category are directly executed based on the component, information overload of the component will occur, which causes the user to be unable to distinguish the information of the two tasks. Therefore, the user's intention can be identified, and according to the user's intention, the task to be converted and the original category task are determined from the first task and the second task, and the task to be converted is converted into a task of the second category. In this way, the original category task (the task of the first category) and the task of the second category can be executed by the virtual screen and the audio output module respectively, the information overload of the single component is avoided, the functions of the two components are fully utilized, the processing efficiency is improved, and the system stability is enhanced. Moreover, the user can obtain the information of the two tasks from two dimensions of vision and hearing respectively, the problem that the user is unable to distinguish the information of the two tasks is avoided, and thus the user's use experience is improved.

[0177] The specific limitations of the task processing apparatus can be referred to the limitations of the task processing method in the foregoing, which will not be repeated here. Each module in the task processing apparatus described above can be realized by software, hardware and combinations thereof in whole or in part. Each module described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in the form of software, so as to call and execute the operations corresponding to each module by the processor.

[0178] In addition, the present application also provides an electronic device, which can be the extended reality device described above, and the extended reality device can be smart glasses, such as Figure 6 As shown in the figure, it shows the structural schematic diagram of the electronic device related to the present application, specifically:

[0179] The electronic device can include a processor 601 with one or more processing cores and a memory 602 with one or more computer readable storage media, etc. Those skilled in the art can understand that, Figure 6The electronic device structure shown in the figures does not constitute a limitation on the electronic device, and can include more or fewer components than shown, or combine certain components, or arrange the components differently. Among them:

[0180] The processor 601 is the control center of the electronic device, connects various parts of the entire electronic device through various interfaces and lines, and performs various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 602 and calling data stored in the memory 602, thereby overall monitoring the electronic device. Optionally, the processor 601 can include one or more processing cores; preferably, the processor 601 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 601.

[0181] The memory 602 can be used to store software programs and modules, and the processor 601 executes various functions and data processing by running the software programs and modules stored in the memory 602. The memory 602 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 602 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 602 can also include a memory controller to provide access for the processor 601 to the memory 602.

[0182] In one embodiment, the electronic device further includes a power supply 603 for supplying power to each component, and preferably the power supply 603 can be logically connected to the processor 601 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 603 can also include one or more than one direct current or alternating current power supply, a recharging system, a power supply device debugging circuit, a power supply converter or inverter, a power supply state indicator, etc. any component.

[0183] In one embodiment, the electronic device can further include an input unit 604, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0184] When the electronic device is specifically smart glasses, in addition to the above structure, at least a glasses body frame, an optical display component, an electronic circuit component, a sensor, etc. are further included, and the sensors built in the glasses include a heart rate monitor, a blood glucose detector, a microphone, a camera, and / or an eye tracker.

[0185] Although not shown, the electronic device can further include a display unit, etc., which will not be described here. Specifically, in the present embodiment, the processor 601 in the electronic device will load the executable file corresponding to the process of one or more application programs into the memory 602 according to the following instructions, and run the application program stored in the memory 602 by the processor 601, thereby implementing the steps in any of the task processing methods provided in the present application.

[0186] Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. Specifically, the electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0187] In one embodiment, an electronic device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in any embodiment of the present application.

[0188] In one embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program is executed by a processor to implement the method described in any embodiment of the present application.

[0189] In some embodiments, a computer program product is also provided, including a computer program or instructions, which are executed by a processor to implement the method described in any embodiment of the present application.

[0190] The specific implementation of each of the above operations can be referred to the previous embodiments, which will not be described here.

[0191] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by related hardware controlled by the instructions, which can be stored in a computer readable storage medium and loaded and executed by a processor.

[0192] To this end, the present application provides a computer readable storage medium, storing a computer program, which can be loaded by a processor to execute the steps in any of the task processing methods provided in the present application.

[0193] The specific implementation of the above operations can refer to the foregoing embodiments, which will not be repeated here.

[0194] The computer readable storage medium can include a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0195] Due to the instructions stored in the computer readable storage medium, the steps of any one of the task processing methods provided in the present application can be executed, and thus the beneficial effects of any one of the task processing methods provided in the present application can be achieved. Details are described in the foregoing embodiments, which will not be repeated here.

[0196] Finally, it should be noted that, in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or terminal device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or terminal device including the element.

[0197] The above describes in detail one task processing method, device, electronic device and computer readable storage medium provided in the present application. The principle and implementation manner of the present application are described by applying specific examples in this document. The above embodiment description is only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range will be changed; in conclusion, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A task processing method characterized by, The method is applied to an extended reality device, the extended reality device comprising a virtual screen and an audio output module; the method comprises: in response to triggering a first task of a first category, determining whether a second task of the first category is currently being executed; when the second task of the first category is currently being executed, identifying a user intention, comprising: obtaining first information of the first task of the first category and second information of the second task of the first category; obtaining user behavior and event context; obtaining a user intention identification model; analyzing the first information, the second information, the user behavior and the event context based on the user intention identification model to determine the user intention; according to the user intention, determining a to-be-converted task and an original category task from the first task and the second task; converting the to-be-converted task into a task of a second category; executing the original category task and the task of the second category; wherein the task of the first category is executed by one of the virtual screen and the audio output module, and the task of the second category is executed by the other of the virtual screen and the audio output module.

2. The method of claim 1, wherein, The task of the first category is an audio task; when the second task of the first category is currently being executed, identifying the user intention comprises: when the second task of the first category is currently being executed, detecting whether the audio output module has left and right channel independent playback capability; when the audio output module does not have the left and right channel independent playback capability, identifying the user intention.

3. The method of claim 2, wherein, The method further comprises: when the audio output module has the left and right channel independent playback capability, playing audio corresponding to any one of the first task and the second task of the first category using a left channel; playing audio corresponding to the other of the first task and the second task of the first category using a right channel.

4. The method of claim 1, wherein, when the to-be-converted task is a voice call task, converting the to-be-converted task into a task of a second category comprises: converting audio of the voice call into text; generating a call summary according to the text; the execution of the task of the second category comprises: displaying the text and / or the call summary using the virtual screen.

5. The method of claim 1, wherein, when the to-be-converted task is a voice navigation task, converting the to-be-converted task into a task of a second category comprises: obtaining a navigation route map corresponding to the voice navigation task; obtaining a current position in real time; generating a navigation instruction text according to the current position and the navigation route map; the execution of the task of the second category comprises: displaying the navigation route map and the navigation instruction text in real time using the virtual screen.

6. The method of claim 1, wherein, when the to-be-converted task is a navigation display task, converting the to-be-converted task into a task of a second category comprises: obtaining a navigation route map corresponding to the navigation display task; obtaining a current position in real time; generating a navigation instruction speech according to the current position and the navigation route map; the execution of the task of the second category comprises: playing the navigation instruction speech using the audio output module.

7. A task processing apparatus characterized by comprising: The application is applied to an extended reality device, the extended reality device comprising a virtual screen and an audio output module; the device comprises: a task triggering module, configured to determine whether a second task of a first category is being executed currently in response to triggering a first task of the first category; an intention recognition module, configured to recognize a user intention when the second task of the first category is being executed currently, comprising: obtaining first information of the first task of the first category and second information of the second task of the first category; obtaining user behavior and event context; obtaining a user intention recognition model; analyzing the first information, the second information, the user behavior and the event context based on the user intention recognition model to determine the user intention; a task determination module, configured to determine a task to be converted and an original category task from the first task and the second task according to the user intention; a task conversion module, configured to convert the task to be converted into a task of a second category; a task execution module, configured to execute the original category task and execute the task of the second category; wherein the task of the first category is executed by one of the virtual screen and the audio output module, and the task of the second category is executed by the other of the virtual screen and the audio output module.

8. An electronic device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps in the task processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps in the task processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video switching method and device, equipment and storage medium

    CN116600170A

  • Speech-enabled augmented reality

    US20230055477A1