Method and apparatus for optimizing information display of smart glasses
By collecting multimodal interaction data and using visual-assisted sound source localization technology, smart glasses achieve personalized display and immersive experience in cross-voice conferencing, solving the problem of insufficient display optimization in existing technologies and improving the user's immersive experience and information acquisition efficiency.
Patent Information
- Application Number
- CN202511099581.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing smart glasses have shortcomings in display optimization during cross-voice conferencing, affecting the immersive experience and practical usability, especially in providing personalized processing and immersive experience when dealing with complex sound environments and visual information.
By collecting multimodal interaction data, including voice, eye movement, and touch data, the backtracking function is triggered, displaying a timeline floating window and history labels. The content displayed in the floating window is optimized according to the data type. Combined with visual-assisted sound source localization and speech separation technology, personalized translation and image rendering are provided.
It enables accurate differentiation of speaker voices in cross-voice conferences, providing personalized display and immersive experience, improving user information acquisition efficiency and immersive experience.
Smart Images

Figure CN120595988B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent wearable devices, and particularly relates to an information display optimization method and device for intelligent glasses. BACKGROUND
[0002] In a cross-voice conference scenario, intelligent glasses are regarded as a potential solution to provide real-time translation subtitles, PPT content synchronous display and other functions in combination with XR (extended reality) technology. However, the existing intelligent glasses still have many deficiencies in display optimization, thereby affecting the immersive experience and actual usability of the cross-voice conference. SUMMARY
[0003] The present application provides an information display optimization method and device for intelligent glasses, to solve the display optimization defects of the intelligent glasses in the prior art, realize display optimization of the intelligent glasses, and improve the immersive experience of users participating in the cross-voice conference.
[0004] The present application provides an information display optimization method for intelligent glasses, comprising the following steps:
[0005] In the case that a user uses the intelligent glasses to participate in a conference, multi-modal interaction data of the user is collected; the multi-modal interaction data is interaction data between the user and the conference environment collected through multiple interaction modes;
[0006] If the backtracking function of the intelligent glasses is triggered according to the multi-modal interaction data, a timeline floating window of the intelligent glasses is called to display a plurality of historical record tags;
[0007] In response to a selection operation of the user on the timeline floating window, the content to be played back corresponding to the historical record tag selected by the user is obtained;
[0008] According to the data type of the content to be played back, the content to be played back is displayed in the form of a floating window.
[0009] According to the information display optimization method for intelligent glasses provided by the present application, the backtracking function of the intelligent glasses is triggered according to the multi-modal interaction data, which comprises:
[0010] In the case that the multi-modal interaction data is voice data, text content and emotional tone information in the voice data are identified; if the text content includes a trigger word of the backtracking function, or the emotional tone information indicates that the user is in a state of confusion, the backtracking function is triggered;
[0011] In a case where the multi-modal interaction data is eye movement data, according to the eye movement data, a line-of-sight direction and line-of-sight scanning information of the user are acquired; if the line-of-sight direction deviates from a main display area for more than a first preset time length, and / or the line-of-sight scanning information is disordered scanning information, the backtracking function is triggered.
[0012] In a case where the multi-modal interaction data is touch data, according to a temple sliding operation of the smart glasses, the backtracking function is triggered.
[0013] According to the information display optimization method of the smart glasses provided by the application, the data type of the content to be played back is determined, and the content to be played back is displayed in a floating window form, including:
[0014] According to display information of a main display area of the smart glasses, display parameters of the content to be played back and a display position of a corresponding floating window of the content to be played back are determined;
[0015] According to the data type of the content to be played back, a display mode of the content to be played back is determined;
[0016] In the display position, according to the display mode and the display parameters, the content to be played back in the floating window is rendered, and the rendered content to be played back is displayed.
[0017] According to the information display optimization method of the smart glasses provided by the application, the content to be played back includes audio translation content of a speaker in a conference, and the audio translation content is acquired according to the following manner:
[0018] According to visual data acquired by the smart glasses in real time, a three-dimensional space position of the speaker is determined, and a lip region video stream of the speaker is extracted;
[0019] According to the three-dimensional space position, a beamforming parameter of a microphone array in the smart glasses is controlled, an audio pickup beam pointing to the speaker is generated, and an audio signal of the speaker is collected;
[0020] According to the lip region video stream, a lip movement feature of the speaker is extracted, and a time sequence synchronization score of the lip movement feature and the audio signal is determined;
[0021] In a case where the time sequence synchronization score is greater than a preset score, the audio signal and the lip movement feature are input into a pre-trained speech separation model, and a target speech signal after background noise separation output by the speech separation model is obtained;
[0022] The target speech signal is translated to obtain the audio translation content.
[0023] According to the information display optimization method of the intelligent glasses provided by the application, the target voice signal is translated to obtain the audio translation content, which comprises the following steps:
[0024] An acoustic feature parameter of the target voice signal is extracted, and a text emotion feature is obtained by performing emotion keyword detection on the translation text of the target voice signal.
[0025] According to the acoustic feature parameter and the text emotion feature, a tone analysis result is obtained by performing tone analysis on the target voice signal.
[0026] According to the tone analysis result, a visual annotation is added to the translation text of the target voice signal to obtain the audio translation content.
[0027] According to the information display optimization method of the intelligent glasses provided by the application, in the case that the time sequence synchronization score is greater than the preset score, the audio signal and the lip movement feature are input into a pre-trained speech separation model to obtain a target voice signal after separating background noise output by the speech separation model.
[0028] In the case that the continuous speech duration of the target voice signal is greater than a second preset duration, or the number of words of the transcription text of the target voice signal is greater than a preset number of words, a conclusive sentence is extracted from the target voice signal to generate a summary card containing at least one key point.
[0029] The summary card is displayed in a floating form on the display interface of the intelligent glasses.
[0030] According to the information display optimization method of the intelligent glasses provided by the application, the to-be-playback content further comprises translation content corresponding to the media content displayed by the speaker, and the translation content corresponding to the media content is obtained in the following manner:
[0031] According to the image of the media content obtained by the intelligent glasses, a text region and a chart region in the media content are identified.
[0032] Text content is extracted from the text region, and chart semantic features are obtained by performing image semantic understanding on the chart region.
[0033] The target voice signal, the text content and the chart semantic features are input into a multi-modal encoder for joint encoding to generate context-aware translation content.
[0034] The application also provides an information display optimization device of intelligent glasses, which comprises the following modules:
[0035] The data collection module is configured to collect multi-modal interaction data of the user in a case where the user participates in a conference using the smart glasses.
[0036] The calling module is configured to call a timeline floating window of the smart glasses to display a plurality of historical record tags if the backtracking function of the smart glasses is triggered according to the multi-modal interaction data.
[0037] The to-be-played content acquisition module is configured to acquire to-be-played content corresponding to the historical record tag selected by the user in response to a selection operation of the user on the timeline floating window.
[0038] The optimized display module is configured to display the to-be-played content in the form of a floating window according to the data type of the to-be-played content.
[0039] The application further provides a smart glasses, including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the information display optimization method of the smart glasses according to any one of the above.
[0040] The application further provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on the processor to implement the information display optimization method of the smart glasses according to any one of the above.
[0041] The application further provides a computer program product, which includes a computer program, and the computer program is executable on the processor to implement the information display optimization method of the smart glasses according to any one of the above.
[0042] The information display optimization method and device of the smart glasses provided by the application, by collecting multi-modal interaction data of the user in a case where the user participates in a conference using the smart glasses, the multi-modal interaction data is the interaction data between the user and the conference environment collected through multiple interaction modes, if the backtracking function of the smart glasses is triggered according to the multi-modal interaction data, the timeline floating window of the smart glasses is called to display a plurality of historical record tags, in response to the selection operation of the user on the timeline floating window, the to-be-played content corresponding to the historical record tag selected by the user is acquired, and the to-be-played content is displayed in the form of a floating window according to the data type of the to-be-played content. The application triggers the backtracking function through the multi-modal interaction data, and displays the historical record tags, so that the user can quickly backtrack the key information in the conference, and the display optimization of the to-be-played content is realized by displaying the to-be-played content through the data type and the floating window form, thereby improving the immersive experience of the user participating in the cross-voice conference. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0044] Figure 1 It is the application environment diagram of the information display optimization method of the smart glasses provided by the present application.
[0045] Figure 2 It is the flowchart of the information display optimization method of the smart glasses provided by the present application.
[0046] Figure 3 It is the structural schematic diagram of the information display optimization device of the smart glasses provided by the present application.
[0047] Figure 4 It is the structural schematic diagram of the smart glasses provided by the present application. DETAILED DESCRIPTION
[0048] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort belong to the protection scope of the present application.
[0049] In the related art, although the translation device (such as a translation machine) can provide certain translation function, it often needs to be held by the user, which is inconvenient to use and lacks immersive experience. At the same time, for the complex sound environment in the conference site, it is difficult to accurately distinguish the speaker's voice from other noise, and it is also impossible to perform personalized processing according to different speakers. In addition, when processing the PPT or other visual information in the speech, the translation device cannot well combine the image rendering and translation function, and cannot meet the multiple needs of the user in the cross-voice conference. Therefore, the present application proposes an information display optimization method of smart glasses which is more intelligent, more convenient and has more immersive experience. The method can accurately distinguish the speaker's voice, automatically recognize and translate the voice, and perform personalized display according to different speakers, while combining image rendering to provide immersive viewing experience, solving the problem that the existing translation device cannot well combine image rendering and translation function, and cannot meet the multiple display optimization needs of the user in the cross-voice conference.
[0050] The information display optimization method and device of the smart glasses of the present application will be described below in combination with Figures 1-4 the drawings.
[0051] The information display optimization method of the smart glasses provided by the application can be applied to the application environment as shown in Figure 1 The terminal 20 communicates with the server 10 through a network, and the data storage system can store data required to be processed by the server 10. The data storage system can be integrated on the server 10, or placed on a cloud or other network server. The terminal 20 can collect multi-modal interaction data of a user, trigger a backtracking function according to the multi-modal interaction data, display a plurality of historical record tags through a time axis floating window, and display optimization of the to-be-playback content corresponding to the historical record tag selected by the user, thereby improving the immersive experience of the user participating in the cross-voice conference. Alternatively, the terminal 20 can send the collected multi-modal interaction data to the server 10, determine whether to trigger the backtracking function through the server 10, return the triggering information to the terminal 20 by the server 10, and realize the display optimization of the to-be-playback content according to the triggering information of the backtracking function by the terminal 20. The terminal 20 can be a smart wearable device, such as smart glasses.
[0052] It should be noted that the execution subject of the information display optimization method of the smart glasses provided by the application is a smart wearable device, such as smart glasses.
[0053] Figure 2 The information display optimization method of the smart glasses provided by the application is a flowchart as shown in Figure 2 The method comprises the following steps:
[0054] Step 101, in the case that a user uses smart glasses to participate in a conference, multi-modal interaction data of the user is collected.
[0055] The multi-modal interaction data is the interaction data between the user and the conference environment collected through a plurality of interaction modes. Specifically, the multi-modal interaction data is the data reflecting the dynamic interaction process between the user and the conference environment, which is synchronously collected through a sensor array built in the smart glasses.
[0056] The multi-modal interaction data can include voice data, eye movement data, touch data, gesture data, and head posture data of the user, etc. It should be understood that the voice data can include voice instructions such as “open a file”, “search information”, etc., and emotional information in the voice of the user such as happiness, sadness, anger, etc. The eye movement data can include gaze coordinates, gaze duration, saccade patterns, blink frequencies, etc. The touch data can include mirror leg sliding operation data, positions of the user's fingers touching the screen, force of the user touching the screen, gestures such as sliding and zooming of the user on the screen through the fingers, etc. The gesture data can include gesture actions such as waving and pointing, etc. The head posture data can include data such as rotation and inclination of the head, etc.
[0057] When users participate in meetings using smart glasses, multimodal interaction data is collected in real time through multimodal sensors. For example, voice data is collected through microphone arrays or voice recognition sensors; eye movement data is collected through cameras or eye-tracking sensors; touch data is collected through piezoelectric film sensors or pressure sensors; gesture data is collected through cameras; and head posture data is collected through inertial measurement units (IMUs).
[0058] Step 102: If the backtracking function of the smart glasses is triggered based on the multimodal interaction data, the timeline floating window of the smart glasses is retrieved to display multiple historical record tags.
[0059] It should be understood that the rewind function is a feature that allows users to return to a previous state or step, for example, allowing users to quickly retrieve and re-experience missed meeting content.
[0060] Based on preset trigger rules, pattern recognition is performed on multimodal interaction data. When a valid trigger signal is detected, the smart glasses' backtracking function is activated. Then, meeting history records are retrieved from local cache or cloud server to generate a structured timeline floating window. It should be understood that the generation of the timeline floating window includes: segmenting the meeting history records according to semantic themes to form multiple time period history labels; generating summary text containing core data points and argumentation logic for each history record label; arranging the labels in chronological order according to the meeting progress in a three-dimensional display space, and marking their importance levels. For example, the timeline floating window displays history record labels such as "10:15-10:18 Technical Principles" and "10:20-10:22 Market Analysis".
[0061] For example, based on multimodal interaction data, the system determines that the user is sliding the right temple of the smart glasses, triggering the backtracking function; based on multimodal interaction data, the system obtains the user's voice input such as "repeat the previous paragraph" and triggers the backtracking function; based on multimodal interaction data, the system determines that the user is in a specific eye movement pattern (such as quickly looking to the left twice), triggering the backtracking function.
[0062] Optionally, if the user inputs a natural language command such as "repeat the previous paragraph" via voice, the history of the last 30 seconds can be automatically displayed, and the temple of the glasses will vibrate to indicate that the history has been retrieved.
[0063] Step 103: In response to the user's selection operation on the timeline floating window, obtain the content to be played back corresponding to the historical record tag selected by the user.
[0064] The user can select a history record tag in the timeline floating window through touch operation, eye movement operation, voice operation, etc. For example, the user can slide the temple or click the glasses frame to select the target tag; the user can fix his gaze on a certain tag for more than 1 second (confirmed by eye movement tracking); the user can say an instruction such as "jump to the third tag".
[0065] In response to the selection operation of the user on the timeline floating window, the smart glasses can quickly locate the corresponding conference segment from the storage according to the time stamp (such as 10:15-10:18) associated with the history record tag selected by the user, synchronously extract all associated content of the time period, for example, the original recording of the speaker, the PPT page / whiteboard screenshot, the user's notes or annotations at that time, automatically intercept the core part of the paragraph (such as skipping the silent interval), and simultaneously associate and display the relevant materials. For example, when the "market growth data" tag is selected, the corresponding chart file is automatically attached.
[0066] In step 104, the to-be-played content is displayed in the form of a floating window according to the data type of the to-be-played content.
[0067] It should be understood that the floating window is a user interface element that can display information on the main screen or main application without the need for the user to switch to another application or screen.
[0068] The data type of the to-be-played content can be identified by file extension, content header information, or a machine learning model, etc. The data type can include voice-dominant type (such as a speech segment), chart / data type (such as a column chart in a PPT), formula / code type (such as mathematical derivation or program segment), mixed type (such as a speech containing voice and chart), etc. According to the data type, a preset display template is matched, and the layout (position, size, level), visual style (font, color, animation), and interactive function (zoom, navigation, annotation) of the floating window are optimized, and the to-be-played content is displayed in the form of a floating window.
[0069] For example, in the user interface design, the history window is located on the left side of the field of view, does not block the central view, uses a timeline design, and displays the last 5-10 minutes of content. Important content is highlighted on the timeline to facilitate quick positioning. The current viewing position is clearly marked with an indicator to help users understand the current position. The backtracking content is displayed in a scrolling text manner, with key sentences and important content highlighted in bold or highlighted to enhance readability. Users can choose to display the original text + translation or only display the translation to adjust the amount of information according to their needs. Smart glasses support browsing by paragraph or sentence unit, providing flexible content navigation. In terms of interaction feedback, a slight vibration feedback is provided when the glasses are slid, an audio confirmation is provided when the voice command recognition is successful, a simple animation transition effect is displayed when backtracking, and clear current position and original time clues are provided to enhance user's operation perception and time perception.
[0070] For example, replay a segment of speech about "quarterly revenue growth of 25%": the left side displays the translated text "Q3 revenue↑25%", with "↑25%" highlighted in red; the right side displays a sound wave, with the emphasized part amplified (corresponding to the speaker raising his voice); the bottom floating window displays the PPT thumbnail at that time (semi-transparent).
[0071] For example, replay a "market share comparison column chart": the position of the floating window is superimposed on the original chart (Z-axis moves forward 10 cm), only the core column is kept (legend, coordinate axis labels are hidden), and "A company: 38% (+5%)" is directly marked on the top of the column. The "trend in the past three years" shortcut button is provided at the bottom. When the user gazes at a certain column for 2 seconds, the detailed calculation method is displayed.
[0072] For example, replay the segment where the speaker explains "user growth curve": the left window displays the speech text sentence by sentence (synchronously highlights the current sentence), and the right window dynamically draws the corresponding growth curve (extends with the speech progress). Click on the word "peak" in the text to automatically zoom the curve to the peak point.
[0073] The information display optimization method of the smart glasses provided by the embodiment of the application comprises the following steps: in the case that a user uses smart glasses to participate in a conference, collecting multi-modal interaction data of the user; the multi-modal interaction data is interaction data between the user and the conference environment collected through multiple interaction modes; if the recall function of the smart glasses is triggered according to the multi-modal interaction data, calling a timeline floating window of the smart glasses and displaying multiple historical record tags; in response to a selection operation of the user on the timeline floating window, obtaining to-be-replayed content corresponding to the historical record tag selected by the user; and displaying the to-be-replayed content in the form of a floating window according to the data type of the to-be-replayed content. The embodiment of the application triggers the recall function through the multi-modal interaction data and displays the historical record tags, so that the user can quickly recall key information in the conference, and the display optimization of the to-be-replayed content is realized through the data type and the floating window form of the to-be-replayed content, thereby improving the immersive experience of the user participating in the cross-voice conference.
[0074] According to the above embodiment, the recall function of the smart glasses is triggered according to the multi-modal interaction data, which comprises:
[0075] In step 110, in the case that the multi-modal interaction data is voice data, text content and emotional tone information in the voice data are identified; if the text content includes a trigger word of the recall function, or the emotional tone information indicates that the user is in a confused state, the recall function is triggered;
[0076] In step 120, in the case that the multi-modal interaction data is eye movement data, the gaze direction and gaze scanning information of the user are obtained according to the eye movement data; if the gaze direction deviates from the main display area for more than a first preset time length, and / or the gaze scanning information is disordered scanning information, the recall function is triggered;
[0077] In step 130, in the case that the multi-modal interaction data is touch data, the recall function is triggered according to a temple sliding operation of the smart glasses.
[0078] The smart glasses identify the text of voice transcription in real time, and detect whether the text contains a preset trigger word, such as “repeat”, “not clear”, “say again”, etc. For example, a word library containing all possible keywords or phrases that trigger the recall function can include “explain”, “repeat”, “say again”, etc. The smart glasses match the text identified in real time with the preset trigger word library, check whether the keywords are contained, and if the keywords are contained in the text, the user is identified as possibly needing to recall.
[0079] Optionally, in addition to detecting keywords, the smart glasses can also understand the user's intent in combination with the context, because some phrases can have different meanings in different contexts. For example, "Can you repeat?" may indicate that the user needs to repeat, while "Repetition is necessary" may just state an opinion and does not indicate that the user needs to repeat. The smart glasses need to be able to distinguish between the two cases.
[0080] By identifying specific keywords in the user's voice input and understanding the context, it can be determined whether the user needs to repeat or backtrack previous information, which can improve the intelligence and user experience of voice interaction, ensuring that the smart glasses can accurately respond to the user's needs.
[0081] By determining whether the user is confused through emotional tone information (such as pitch, speed, and pauses), a specific combination of emotional tone information can be used as a signal of user confusion, such as a sudden rise in pitch, a decrease in speed, and frequent fillers (such as "um...") in any combination. For example, the user mutters to himself, "What does this data mean?" (pitch up + pause), and the smart glasses detect the questioning tone + keyword "meaning" to automatically trigger the backtracking function.
[0082] If the smart glasses match the trigger words from the trigger word library and determine that the user is in a state of confusion through emotional tone information, the backtracking function is automatically triggered.
[0083] It should be understood that the central 30-degree viewing angle range of the screen can be defined as the main display area, such as the location of the speaker or the core content of the PPT. If the user's gaze continuously leaves this area for more than 3 seconds (according to requirements), it can be considered that the user may not be focused on the current presentation content and needs to backtrack or help. For example, the user's gaze point coordinates are calculated by tracking the user's pupil position using an infrared camera to determine whether the user is looking at the main display area.
[0084] It should be understood that the user's gaze will usually move in an orderly zigzag pattern, which is a normal reading behavior. If the user's gaze jumps quickly and irregularly, such as repeatedly jumping horizontally between different areas of the PPT, it may indicate that the user is looking for information or feeling confused and needs to backtrack or help. For example, by calculating the entropy value of the user's gaze path (an indicator of randomness), if the entropy value exceeds a certain threshold, it is determined to be disordered scanning.
[0085] For example, when the user is watching the PPT, he repeatedly scans the chart and title, but does not focus on any content for more than 2 seconds. The smart glasses recognize this behavior as "difficulty in understanding" and automatically pop up a floating button asking the user whether he needs to replay the chart explanation. This design can help the user obtain the information he needs without interrupting the meeting process.
[0086] It should be understood that the temple sliding operation can be predefined, for example, the user quickly slides down the temple twice, triggering the playback of the last 30 seconds of content; the user slowly slides up the temple, triggering the sentence-by-sentence rewind function. For example, the sliding direction and speed on the temple can be detected using a capacitive sensor, while an accelerometer is used to assist in de-bouncing to ensure the accuracy of the detection. After detecting the temple sliding operation, the backtracking function is triggered.
[0087] Alternatively, when the user performs the sliding temple and frowning (or other facial expression recognition) operations simultaneously, the touch operation is prioritized, because the touch operation usually indicates that the user has a more explicit intention, i.e., the active intention is stronger. If the user is detected to have a confused tone (through voice recognition technology), but the user is still looking at the PPT, the system will delay triggering the backtracking function, because the user may be thinking or trying to understand the current content, rather than immediately needing help. For example, the user unconsciously slides down the temple twice while listening to a complex term, and frowns to indicate possible confusion, the smart glasses immediately play back the explanation paragraph of the term, and display a definition floating window below the term, because the touch operation is prioritized. The user shows a confused tone while listening to a speech, but the gaze is still focused on the PPT, the smart glasses do not immediately trigger the backtracking function, but wait for the user's possible further operation or confirmation, in order to avoid interrupting the user's thinking process that may be in progress.
[0088] The embodiments of the present application make the smart glasses more intelligent and flexible in responding to the user's needs through the application of multi-modal interaction data. By analyzing the user's voice, eye movement and touch behavior, it can automatically determine whether the user needs to backtrack to the previous information, thereby providing more personalized and timely help, based on which not only the user experience is improved, but also the information acquisition and review is made more natural and convenient. On the other hand, the user can quickly backtrack the historical content, solve the problems of information overload and understanding fragmentation of traditional devices, and upgrade the smart glasses from a simple translation tool to an intelligent terminal for assisting cross-language cognition, significantly improving the information understanding efficiency and interaction convenience in meetings.
[0089] According to the above-mentioned embodiments, the displaying the to-be-playback content in the form of a floating window according to the data type of the to-be-playback content comprises:
[0090] In step 210, the display parameters of the to-be-playback content and the display position of the to-be-playback content corresponding floating window are determined according to the display information of the main display area of the smart glasses.
[0091] In step 220, the display mode of the to-be-playback content is determined according to the data type of the to-be-playback content.
[0092] At step 230, in the display position, the to-be-played content in the floating window is rendered according to the display mode and the display parameter, and the rendered to-be-played content is displayed.
[0093] It should be understood that the display information of the display area can include: information currently displayed by the smart glasses, such as text, image, video, etc.; layout mode of the display content, such as full screen, split screen, floating window, etc.; current state of the display content, such as playing, pausing, scrolling, etc.; user interaction history with the display content, such as clicking, sliding, zooming, etc.; ambient light conditions in which the user is located, such as brightness, color temperature, etc.; current line-of-sight direction of the user, etc.
[0094] It should be understood that the display parameter of the to-be-played content can include: size of the to-be-played content, such as resolution, pixel size, etc.; format of the to-be-played content, such as JPEG, PNG, MP4, MP3, etc.; display style of the to-be-played content, such as font, color, background, etc.
[0095] It should be understood that the display position of the floating window can include: specific coordinate position of the floating window in the display area, such as X, Y coordinates; relative position of the floating window relative to the display area, such as upper left corner, lower right corner, etc.
[0096] The current display content of the main display area of the smart glasses is analyzed to determine the display parameter (such as size, transparency, color, etc.) of the to-be-played content and the display position (such as upper right corner of the screen, lower left corner, etc.) of the floating window. According to the data type (such as text, image, video, etc.) of the to-be-played content, a suitable display mode is matched, for example, text is displayed in a scrolling text box, image is displayed in a thumbnail, and video is displayed in a small player. In the determined display position, the to-be-played content in the floating window is rendered according to the display mode and the display parameter, and the rendered content is displayed to the user.
[0097] For example, assume that a user is attending a meeting including a slide presentation and wants to review a missed slide through the flashback function of the smart glasses. The user says "review the previous slide" through a voice command or triggers the flashback function through a specific gesture. The smart glasses retrieve the content to be played back corresponding to the user-selected historical record tag from the timeline floating window, such as an image of a slide. The smart glasses analyze the main display area and find that there is enough space in the lower right corner and no other floating window is currently displayed. The smart glasses decide to place the floating window in the lower right corner with a size of 20% of the screen and a transparency of 50% so that the user can see the content of the floating window while also being able to see the main display area. Since the content to be played back is a slide image, it is displayed in the form of a thumbnail, and the user can view the details through a gesture. The slide image is rendered in the floating window in the lower right corner, and the size and transparency are adjusted. After rendering is complete, the floating window is displayed, and the user can see the content of the missed slide. In addition, the user can view the details of the slide through a gesture operation or close the floating window through a voice command.
[0098] The embodiments of the present application can determine display parameters and positions individually based on the display information of the main display area and the data type of the content to be played back, ensuring that the user obtains the most comfortable viewing experience. By displaying the content to be played back in the form of a floating window without interfering with the main display area, the user can view the previously missed information while maintaining attention on the current content. Based on this, the user can quickly access the history through simple interaction without manually searching or operating, thereby improving the efficiency of obtaining information.
[0099] According to the above embodiments, the content to be played back includes audio translation content of a speaker in a meeting, and the audio translation content is obtained in the following manner:
[0100] Step 310, determining a three-dimensional space position of the speaker according to visual data obtained by the smart glasses in real time, and extracting a lip region video stream of the speaker;
[0101] Step 320, controlling beamforming parameters of a microphone array in the smart glasses to generate a sound pickup beam pointing to the speaker according to the three-dimensional space position, and collecting an audio signal of the speaker;
[0102] Step 330, extracting a lip movement feature of the speaker according to the lip region video stream, and determining a time sequence synchronization score of the lip movement feature and the audio signal;
[0103] Step 340, in a case where the time sequence synchronization score is greater than a preset score, inputting the audio signal and the lip movement feature into a pre-trained speech separation model to obtain a target speech signal separated from background noise output by the speech separation model;
[0104] Step 350, translating the target speech signal to obtain the audio translation content.
[0105] During the process of the user wearing the smart glasses to attend the meeting, when the speaker starts to speak, the camera of the smart glasses will recognize the face and track, and the microphone array will collect the sound. The system in the smart glasses will fuse the visual information and the audio information, and through the visual guided beam forming technology, the sound pickup direction is aimed at the speaker. At the same time, the deep learning model will separate the speaker's voice from the background noise, such as the conversation of others and the sound of the air conditioner, in real time.
[0106] The visual aided sound source localization and separation The high-definition camera built-in the smart glasses captures the video stream of the speech scene in real time. Then, the visual aided sound source localization and separation can be realized through the following steps:
[0107] (1) Determine the three-dimensional spatial position of the speaker by using face detection and tracking algorithm. For example, by using computer vision algorithms such as Multi-task Cascaded Convolutional Networks (MTCNN), Face Recognition Network (FaceNet), etc., the faces in the video stream are detected in real time, and the speaker is continuously tracked to provide accurate position information of the speaker.
[0108] (2) According to the three-dimensional spatial position of the speaker, control the beamforming parameters of the microphone array to generate directional sound pickup beam. For example, by combining the visual information (speaker position, head posture) with the multi-microphone array data of the smart glasses, through the visual guided beam forming technology, the sound pickup direction of the microphone array is accurately aimed at the speaker, so as to maximize the suppression of noise and non-target voice from other directions.
[0109] (3) Perform lip reading analysis on the detected speaker's face area. By analyzing the synchronization between lip movement and audio signal, it can further confirm whether the current sounder is the target speaker and assist in distinguishing background human voice. Specifically, according to the video stream of the lip area, the lip movement features of the speaker are extracted, and the time sequence synchronization score of the lip movement features and the audio signal is determined.
[0110] (4) In the case where the timing synchronization score is greater than the preset score, input the audio signal processed by beamforming and the lip movement feature into a speech separation model. The speech separation model adaptively adjusts the model parameters according to the speaker's voiceprint feature, and outputs the separated target speech signal. The voiceprint feature is continuously updated through an online learning module, and the module dynamically optimizes the voiceprint model according to newly collected speech data. For example, the audio features (such as Mel-Frequency Cepstral Coefficients (MFCC), Filter Bank Features (Fbank), etc.) corresponding to the visually enhanced audio signal are fused with the visual features (such as face key points, lip movement features) in the deep learning model. For example, an attention mechanism can be used to dynamically adjust the weights of different modal features, so that the model pays more attention to the key information related to the speaker.
[0111] It should be understood that the deep learning driven noise and non-target speech separation will adopt an end-to-end (End-to-End) speech separation model based on deep learning to achieve more efficient and accurate noise and non-target speech separation.
[0112] After obtaining the target speech signal, a context-aware translation algorithm and a delay optimization algorithm are used to translate the target speech signal to obtain the audio translation content.
[0113] It should be understood that in the first use of the user or in a specific scenario, the user can be guided to register the voiceprint of the commonly used speaker. The registration process can be non-intrusive, for example, before the speech starts, the system automatically collects a small amount of speech samples of the speaker to model the voiceprint. The system will analyze the input speech signal in real time and compare it with the registered voiceprint library to identify the current speaker, which will help the system accurately lock the target speaker in a multi-speaker scenario. Once the current speaker is identified, the system will adaptively adjust the speech processing parameters according to the voiceprint features of the speaker, including:
[0114] Personalized noise reduction parameters: adjust the parameters of the noise reduction algorithm according to the speaker's tone, speech speed, etc. to achieve the best noise reduction effect;
[0115] Personalized speech enhancement: personalized speech enhancement is performed according to the speaker's speech characteristics to make the sound clearer and louder.
[0116] Personalized speech recognition model: load or fine-tune the speech recognition model optimized for the speaker to improve the accuracy of speech recognition.
[0117] In addition, the smart glasses system has continuous learning capability. During the use of the user, the system will continuously collect new voice data, and use the data to update and optimize the voiceprint model and the speech processing model online. For example, when the system finds that the recognition accuracy of a certain speaker decreases, it will automatically trigger model fine-tuning to adapt to the changes in the speaker's voice or new environmental noise patterns.
[0118] The embodiments of the present application focus on capturing the speaker's voice by forming directional sound collection with the microphone array, effectively filtering the surrounding interference; the real-time noise reduction algorithm can improve the signal-to-noise ratio by not less than 15dB according to the environmental noise separation technology based on deep learning, greatly improving the speech recognition accuracy; the context-aware translation algorithm considers the conference scene-specific terminology and context to provide professional translation results; the delay optimization algorithm controls the translation delay within 300ms through sentence segmentation and parallel processing, ensuring the smoothness of user experience.
[0119] According to the above-mentioned embodiments, the target voice signal is translated to obtain the audio translation content, which comprises:
[0120] Step 3501, extracting the acoustic feature parameters of the target voice signal, and detecting the emotional keywords of the translation text of the target voice signal to obtain the text emotional features;
[0121] Step 3502, according to the acoustic feature parameters and the text emotional features, performing tone analysis on the target voice signal to obtain a tone analysis result;
[0122] Step 3503, according to the tone analysis result, adding visual annotations to the translation text of the target voice signal to obtain the audio translation content.
[0123] It should be understood that the smart glasses system establishes a multi-level classification system, wherein the basic tone categories include tone types such as statement, emphasis, question, surprise, hesitation, etc.; the emotional intensity level divides the emotional intensity into 5 levels, providing detailed emotional expression; the multi-modal fusion technology combines voice features and text semantics for comprehensive judgment, improving recognition accuracy; the context-aware mechanism considers the context before and after the text to correct the tone judgment, avoiding errors caused by isolated judgment.
[0124] In the specific implementation of tone marking, the system provides multiple visual expression methods. For example, icon marking displays a small icon (such as an exclamation mark or a question mark) corresponding to the tone next to the text; text effects apply different text effects (such as wavy lines or bold) according to the tone type; color coding uses different shades or saturations for different emotional intensities; dynamic effects make the text show subtle animation effects as the tone changes (such as slight enlargement when emphasized). It should be understood that the visual design follows three principles: the marking design is simple and clear, and does not interfere with text reading; the visual effect naturally corresponds to the tone and is intuitive and easy to understand; the dynamic effect is soft and smooth, avoiding visual fatigue. The system also provides user customization options, allowing users to adjust the display method of tone marking, sensitivity threshold, and turn off specific types of tone marking to meet the individual needs of different users.
[0125] Specifically, the acoustic feature parameters of the target speech signal are extracted by the audio analysis module, including the fundamental frequency trajectory, short-time energy, and spectral tilt, etc., and at the same time, the text emotion features are obtained by identifying the emotional keywords and semantic tendencies in the text corresponding to the transcribed target speech signal through the text analysis module; the acoustic feature parameters and the text emotion features are input into the multi-modal classification model for tone analysis, and the tone analysis results output by the multi-modal classification model are obtained, such as tone categories and emotional intensity levels, wherein the tone categories at least include statements, emphasis, questions, surprise, hesitation; further, according to the tone categories and emotional intensity levels, corresponding visualized markings are generated on the intelligent glasses display interface to obtain the audio translation content. For example, tone category icons are inserted into the side margin of the text, dynamic rendering effects are applied to the keywords, and text color parameters are adjusted according to emotional intensity; wherein the display parameters of the visualized markings can be dynamically adjusted according to the environmental light intensity and the user's gaze focus. The dynamic rendering effects can include:
[0126] 1) Emphasis tone: text pulse amplification effect, amplification ratio 110%-130%, duration 0.3-0.5 seconds;
[0127] 2) Question tone: blue wavy underline, amplitude 2-4 pixels, frequency 0.5 Hz;
[0128] 3) Surprise tone: text jitter animation, jitter amplitude positively correlated with emotional intensity level.
[0129] It should be understood that personalized setting instructions for visualized markings can also be received from the user, and the icon display position offset, dynamic effect trigger threshold, and disable state of specific tone types can be adjusted according to the instructions.
[0130] The embodiment of the present application marks the emotions, doubts and other tones of the speaker through visual symbols, combines environmental noise reduction and data comparison function triggered by eye tracking, and constructs an immersive interactive experience. At the same time, through visual marking of emotions, users can better understand the emotions and tones behind the translated text, making communication more natural and expressive. In addition, users can directly feel the speaker's intention and emotion from the translated text, improving the experience of using translation services.
[0131] According to the above embodiment, in the case where the timing synchronization score is greater than the preset score, the audio signal and the lip movement feature are input into the pre-trained speech separation model to obtain the target speech signal after separating the background noise output by the speech separation model, and then the method further comprises:
[0132] Step 3401, in the case where the continuous speech duration of the target speech signal is greater than a second preset duration, or the number of characters in the transcribed text of the target speech signal is greater than a preset number of characters, a conclusive sentence is extracted from the target speech signal to generate a summary card containing at least one key point;
[0133] Step 3402, displaying the summary card in a floating form on the display interface of the smart glasses.
[0134] The audio acquisition module is used to acquire the speech content in real time. When it is detected that the continuous speech duration of the target speech signal exceeds a second preset duration (such as 2 minutes), or the number of characters in the transcribed text of the target speech signal exceeds a preset number of characters (such as 150 characters), the summary generation module is triggered. The summary generation module extracts a conclusive sentence from the speech content by using a natural language processing model to generate a summary card containing at least one key point. The conditions that the key points must satisfy can include that the number of characters in each key point is not more than 25, the number of simultaneously displayed key points is not more than 3, and the key points with importance scores higher than a preset threshold are displayed for a prolonged duration. Then, the summary card is displayed in a floating form in a preset area of the display interface of the smart glasses. The preset area can be located at the upper right corner of the user's field of view, and the width occupies 15 degrees of the angle of view and the height occupies 8 degrees of the angle of view.
[0135] It should be understood that the display mode of the summary card can include a rounded rectangular design with a background transparency of 30%-50%, a dynamic light effect set at the edge, and a brightness that is adaptively adjusted according to the intensity of the ambient light. New key points are presented with a fade-in animation, and outdated key points are removed with a fade-out animation.
[0136] For example, when a speaker expresses a lot of content, it is difficult to directly understand the key points. The smart glasses can generate summary points (conclusive sentences extracted by AI) in real time and display them in the form of a floating card on the glasses interface.
[0137] Abstract card generation: When the speech content lasts more than 2 minutes or the number of words exceeds 150, the AI model built-in glasses automatically extracts the core points. For example, the speaker elaborates on "2024 Q1 revenue growth of 25%, mainly due to new product release and channel expansion", the system generates a floating card, the title is "Core data: Q1 revenue growth of 25%", and the driving factors are marked below "New product release / Channel expansion, compared with last year: growth rate increased by 10%".
[0138] Design of floating card: full consideration of user experience and information acquisition efficiency. The card adopts a rounded rectangular design, with a size of 15 degrees wide by 8 degrees high in the viewing angle, located in the upper right corner of the user's field of view, not blocking the central vision, and the card background adopts a translucent design, with a micro-light effect on the edge to improve the recognition. Each card can display up to 3 key points at the same time, each key point does not exceed 25 characters, ensuring that the information is concise and clear. New key points appear with a fade-in effect, and outdated key points disappear with a fade-out effect, and key points with high importance stay longer. Users can also fix specific key points by staring and blinking actions, enhancing the flexibility of interaction.
[0139] The embodiment of the application improves the efficiency of obtaining important information by automatically extracting key points and displaying them in the form of abstract cards. The abstract card is displayed in a floating form, which does not interfere with the user's attention to the speaker or the PPT, while providing key information, enhancing the user experience.
[0140] According to the above embodiment, the content to be played back further includes translation content corresponding to media content displayed by the speaker, and the translation content corresponding to the media content is obtained in the following manner:
[0141] Step 411, according to the image of the media content obtained by the intelligent glasses, identifying the text area and chart area in the media content;
[0142] Step 412, extracting text content from the text area, and performing image semantic understanding on the chart area to obtain chart semantic features;
[0143] Step 413, inputting the target voice signal, the text content and the chart semantic features into a multi-modal encoder for joint encoding to generate context-aware translation content.
[0144] The system captures the visual content of media content (such as presentation documents and PowerPoint presentations) in real time using a camera. A layout analysis model is used to identify text regions, chart regions, and their logical relationships within the document. Optical character recognition (OCR) technology is employed to extract text content from the text regions, while an image semantic understanding model is used to perform image semantic understanding on the chart regions, obtaining chart semantic features. The target speech signal, extracted text content, and chart semantic features are then input into a multimodal encoder for joint encoding to generate context-aware translations. Logical relationships are used as prior weights in the attention mechanism during the encoding process. Then, based on the document layout structure and the user's gaze focus, the translated content is rendered hierarchically on the smart glasses display interface. For example, headings are displayed in a first-level style, data content is associated with second-level visual charts, and explanatory text is presented as floating annotations. It should be understood that the translation results maintain terminological consistency with the document's visual content and rhetorical consistency with the speech semantics.
[0145] It should be understood that the construction of a multimodal encoder may include: a speech branch: extracting acoustic features using a Conformer network; a text branch: an encoder based on the document structure of BERT (Bidirectional Encoder Representations from Transformers); and a vision branch: fusing ChartQA features with ResNet-50. The outputs of each branch are dynamically weighted through a cross-modal attention mechanism.
[0146] It should be understood that layered rendering satisfies the following: the font size of the translated title text is 20%-30% larger than that of the body text; the spatial alignment error between the translated chart data and the original chart is less than 5 pixels; and the transparency of floating annotations is dynamically adjusted (30%-70%) according to the user's viewing distance.
[0147] For example, to better integrate image rendering with translation capabilities, context-aware translation and multimodal calibration, and to achieve concise and hierarchical text display on smart glasses for better user comprehension, this can be achieved in the following ways:
[0148] (1) Using computer vision and deep learning technologies, perform real-time layout analysis on PPT slides, document pages, etc. captured by the smart glasses camera, identify different areas such as titles, body text, pictures, charts, lists, etc., and understand the logical relationships between them.
[0149] (2) Optical character recognition is performed on the identified text area to extract the text content. For handwritten notes or non-standard fonts, handwriting recognition technology will be used to ensure that all visible text can be accurately extracted.
[0150] (3) For images and charts, use image recognition and object detection technologies to identify key elements, objects and data. For example, for bar charts, identify coordinate axes, data series, legends, etc., and extract their numerical information; for complex images, perform scene understanding and semantic annotation.
[0151] (4) Multimodal context-aware translation: The translation model will no longer rely on the text output of speech recognition, but will jointly encode the speaker's speech information, the visual information of the PPT (including layout, pictures, and charts), and the extracted text information. This means that the translation model can understand multimodal contexts such as "the speaker is pointing to a data point in the chart and explaining it." For speeches in specific fields (such as medicine, finance, and technology), the system will load the corresponding domain knowledge graph and terminology database to ensure the accuracy of professional terminology translation. Users can also customize the terminology database. The translation model will be able to preserve the emotional and intonation information in the speaker's speech and present it in an appropriate way in the translation results, for example, through emoticons, interjections, or specific rendering effects.
[0152] (5) Multimodal information association: The extracted visual text and image semantics are associated with the audio content of the speaker at the same time. For example, when the speaker mentions "as shown in the figure", the system can automatically identify and associate it with the relevant pictures or charts on the current PPT.
[0153] (6) Visual Consistency Calibration: The translation result will be calibrated for consistency with the visual content on the PPT. For example, if the PPT displays a product name, the translation result will prioritize the official translation of the product name rather than the literal translation.
[0154] (7) Speech-semantic calibration: The translation result will be calibrated with the speaker's speech and semantics. For example, if the speaker is using a metaphor, the translation result will attempt to preserve the meaning of the metaphor rather than a literal translation.
[0155] (8) User feedback and adaptation: The system will allow users to provide real-time feedback on the translation results (such as "inaccurate translation" or "poor display position"), and will adaptively optimize the translation model and rendering strategy based on the feedback data to form a closed loop of continuous learning.
[0156] In one embodiment, if the speaker's PowerPoint presentation is captured by the glasses' camera, simplified explanations of relevant charts / formulas can be displayed next to the translated text. If the user continues to focus on the numbers during this process, the glasses automatically display historical data comparison charts. This can be achieved in the following ways:
[0157] (1) Image capture module captures PPT content in user's field of view in real time, content recognition engine identifies different types of content such as text, charts, and formulas, semantic understanding module understands the meaning of charts and the concept expressed by formulas, and simplified explanation generator generates concise and easy-to-understand explanation text. These modules work together to convert complex visual information into explanations that users can easily understand.
[0158] (2) PPT area detection can automatically identify the projection screen or display area in the field of view, ensuring accurate capture of target content. Chart type recognition algorithm can distinguish between column chart, line chart, pie chart and other different chart types, providing a basis for subsequent analysis. Chart data extraction algorithm extracts numerical data and relationships from chart images, restoring the data information expressed by the chart. Formula semantic analysis algorithm converts mathematical formulas into natural language descriptions, making professional formulas easy to understand. The correlation matching algorithm then semantically correlates PPT content with presentation content, providing contextually relevant explanations.
[0159] (3) In user interface design: chart explanations are displayed in the form of simple cards near the original chart, formula explanations are displayed in the form of annotations below the formula, and explanation text is limited to 50 characters to ensure brevity and clarity. A subtle high-light border is displayed around the identified PPT elements, different types of content are encoded with different colors (charts - blue, formulas - purple, etc.), and a thin line is displayed between the explanation content and the original content to indicate the correlation. Users can trigger detailed analysis by gazing at a specific PPT element for 3 seconds, adjust the level of detail of the explanation through hand gestures (such as double-finger pinch), or use voice commands (such as "explain this chart") to actively trigger analysis, providing multiple convenient interaction methods.
[0160] (4) Chart and formula simplification: After the camera captures the PPT image, computer vision algorithms are used to identify the content. If a column chart is detected, it is automatically converted into a dynamic progress bar, with the text "2023-2024 market share comparison" displayed next to the translated text, and the data for each year displayed in different color bars, with specific values displayed when hovering over the bars.
[0161] (5) When the user gazes at a number, the system displays subtle visual feedback (such as a slight highlight), and after 3 seconds of continuous gaze, a loading animation is displayed, indicating that analysis is in progress. After analysis is complete, the data chart is displayed with an unfolding animation. The chart size is moderate, with a default width of 20 degrees of viewing angle, and a simple design style is used to reduce visual interference. Key data points and trend lines are highlighted using high-contrast colors, and data sources and timeliness information are displayed in the upper right corner of the chart to ensure that users understand the data context. In terms of interaction, users can continue to gaze at the chart to view more detailed information, and the chart will automatically hide after 3 seconds of moving away from the line of sight. It also supports gesture control to enlarge or reduce the chart, and voice commands to save or share chart data, providing flexible and diverse operation methods.
[0162] The embodiments of the present application provide more comprehensive context information by combining text, charts and voice signals, help users better understand and absorb media content, and meanwhile, utilize a multi-modal encoder to jointly encode text content and chart semantic features, to generate more accurate and natural translation content. In addition, the generated translation content can perceive context, better fit the original text intention and context, and improve the applicability of translation.
[0163] The information display optimization device of the smart glasses provided by the present application is described below, and the information display optimization device of the smart glasses described below can be referred to each other corresponding to the information display optimization method of the smart glasses described above.
[0164] Reference Figure 3 The information display optimization device of the smart glasses provided by the present application includes a data acquisition module 301, a calling module 302, a to-be-played content acquisition module 303 and an optimized display module 304.
[0165] The data acquisition module 301 is configured to acquire multi-modal interaction data of a user in a case that the user uses the smart glasses to participate in a conference; the multi-modal interaction data is interaction data between the user and the conference environment acquired through multiple interaction modes.
[0166] The calling module 302 is configured to call a timeline floating window of the smart glasses to display a plurality of historical record tags if a backtracking function of the smart glasses is triggered according to the multi-modal interaction data.
[0167] The to-be-played content acquisition module 303 is configured to acquire to-be-played content corresponding to a historical record tag selected by the user in response to a selection operation of the user on the timeline floating window.
[0168] The optimized display module 304 is configured to display the to-be-played content in the form of a floating window according to a data type of the to-be-played content.
[0169] The information display optimization device of the smart glasses provided by the embodiment of the application, in the case that a user uses the smart glasses to participate in a conference, collects multi-modal interaction data of the user; the multi-modal interaction data is interaction data between the user and the conference environment collected through multiple interaction modes; if the recall function of the smart glasses is triggered according to the multi-modal interaction data, a time axis floating window of the smart glasses is called, and multiple historical record labels are displayed; in response to a selection operation of the user on the time axis floating window, the content to be played back corresponding to the historical record label selected by the user is acquired; and the content to be played back is displayed in the form of a floating window according to the data type of the content to be played back. The embodiment of the application triggers the recall function through the multi-modal interaction data, and displays the historical record label, so that the user can quickly recall the key information in the conference, and the content to be played back is displayed in the form of a floating window according to the data type, so that the display optimization of the content to be played back is realized, thereby improving the immersive experience of the user participating in the cross-voice conference.
[0170] In one embodiment, the calling module 302 is further configured to:
[0171] In the case that the multi-modal interaction data is voice data, text content and emotional tone information in the voice data are identified; if the text content includes a trigger word of the recall function, or the emotional tone information indicates that the user is in a confused state, the recall function is triggered;
[0172] In the case that the multi-modal interaction data is eye movement data, according to the eye movement data, the gaze direction and gaze scanning information of the user are acquired; if the gaze direction deviates from the main display area for more than a first preset time length, and / or the gaze scanning information is disordered scanning information, the recall function is triggered;
[0173] In the case that the multi-modal interaction data is touch data, according to the temple sliding operation of the smart glasses, the recall function is triggered.
[0174] In one embodiment, the optimization display module 304 is specifically configured to:
[0175] According to the display information of the main display area of the smart glasses, the display parameters of the content to be played back and the display position of the floating window corresponding to the content to be played back are determined;
[0176] According to the data type of the content to be played back, the display mode of the content to be played back is determined;
[0177] In the display position, according to the display mode and the display parameters, the content to be played back in the floating window is rendered, and the rendered content to be played back is displayed.
[0178] In an embodiment, the content to be played back includes audio translation content of a speaker in a conference, and the information display optimization apparatus of the smart glasses further includes a translation module configured to:
[0179] determine a three-dimensional spatial position of the speaker according to visual data acquired by the smart glasses in real time, and extract a video stream of a lip region of the speaker;
[0180] control a beamforming parameter of a microphone array in the smart glasses according to the three-dimensional spatial position, generate a sound pickup beam pointing to the speaker, and collect an audio signal of the speaker;
[0181] extract a lip movement feature of the speaker according to the video stream of the lip region, and determine a time sequence synchronization score of the lip movement feature and the audio signal;
[0182] in a case where the time sequence synchronization score is greater than a preset score, input the audio signal and the lip movement feature into a pre-trained speech separation model to obtain a target speech signal after background noise separation output by the speech separation model;
[0183] translate the target speech signal to obtain the audio translation content.
[0184] In an embodiment, the translation module is further configured to:
[0185] extract an acoustic feature parameter of the target speech signal, and perform emotion keyword detection on a translation text of the target speech signal to obtain a text emotion feature;
[0186] perform tone analysis on the target speech signal according to the acoustic feature parameter and the text emotion feature to obtain a tone analysis result;
[0187] add visual annotations to the translation text of the target speech signal according to the tone analysis result to obtain the audio translation content.
[0188] In an embodiment, the optimization display module 304 is further configured to:
[0189] in a case where a continuous speech time length of the target speech signal is greater than a second preset time length or a transcribed text word number of the target speech signal is greater than a preset word number, extract a conclusive sentence from the target speech signal to generate a summary card containing at least one key point;
[0190] display the summary card in a floating form on a display interface of the smart glasses.
[0191] In one embodiment, the content to be replayed further includes translated content corresponding to the media content presented by the speaker, and the translation module is further configured to:
[0192] Based on the image of the media content acquired by the smart glasses, identify the text and chart areas in the media content;
[0193] Text content is extracted from the text area, and image semantic understanding is performed on the chart area to obtain chart semantic features;
[0194] The target speech signal, the text content, and the semantic features of the chart are input into a multimodal encoder for joint encoding to generate context-aware translated content.
[0195] Figure 4 An example is a schematic diagram of the physical structure of smart glasses, such as... Figure 4 As shown, the smart glasses may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an information display optimization method for the smart glasses. This method includes: collecting multimodal interaction data of the user when the user participates in a meeting using the smart glasses; the multimodal interaction data is interaction data between the user and the meeting environment collected through various interaction methods; if the smart glasses' playback function is triggered based on the multimodal interaction data, then the timeline floating window of the smart glasses is retrieved to display multiple historical record tags; in response to the user's selection operation on the timeline floating window, the content to be played back corresponding to the historical record tag selected by the user is obtained; and the content to be played back is displayed in a floating window according to the data type of the content to be played back.
[0196] Further, the logic instructions in the memory 430 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0197] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the information display optimization method of the smart glasses provided by the above-mentioned methods. The method comprises: in the case that a user uses smart glasses to participate in a meeting, collecting multi-modal interaction data of the user; the multi-modal interaction data is interaction data between the user and the meeting environment collected through multiple interaction modes; if the backtracking function of the smart glasses is triggered according to the multi-modal interaction data, calling a timeline floating window of the smart glasses, and displaying a plurality of historical record tags; in response to a selection operation of the user on the timeline floating window, obtaining the to-be-played content corresponding to the historical record tag selected by the user; and displaying the to-be-played content in the form of a floating window according to the data type of the to-be-played content.
[0198] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the information display optimization method of the smart glasses provided by the above-mentioned methods. The method comprises: in the case that a user uses smart glasses to participate in a meeting, collecting multi-modal interaction data of the user; the multi-modal interaction data is interaction data between the user and the meeting environment collected through multiple interaction modes; if the backtracking function of the smart glasses is triggered according to the multi-modal interaction data, calling a timeline floating window of the smart glasses, and displaying a plurality of historical record tags; in response to a selection operation of the user on the timeline floating window, obtaining the to-be-played content corresponding to the historical record tag selected by the user; and displaying the to-be-played content in the form of a floating window according to the data type of the to-be-played content.
[0199] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0201] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for optimizing information display of smart glasses, characterized by, The application comprises the following technical solutions: In the case that a user participates in a conference using smart glasses, multi-modal interaction data of the user is collected; The multi-modal interaction data is interaction data between the user and the conference environment collected through multiple interaction modes; If a backtracking function of the smart glasses is triggered according to the multi-modal interaction data, a timeline floating window of the smart glasses is called, and multiple historical record labels are displayed; In response to a selection operation of the user on the timeline floating window, content to be played back corresponding to a historical record label selected by the user is obtained; The content to be played back is displayed in the form of a floating window according to a data type of the content to be played back; The backtracking function of the smart glasses is triggered according to the multi-modal interaction data, which comprises the following steps: In the case that the multi-modal interaction data is voice data, text content and emotional tone information in the voice data are identified; if the text content includes a trigger word of the backtracking function, or the emotional tone information indicates that the user is in a confused state, the backtracking function is triggered; In the case that the multi-modal interaction data is eye movement data, a line-of-sight direction and line-of-sight scanning information of the user are obtained according to the eye movement data; if the line-of-sight direction deviates from a main display area for more than a first preset time length, and / or the line-of-sight scanning information is disordered scanning information, the backtracking function is triggered; In the case that the multi-modal interaction data is touch data, the backtracking function is triggered according to a temple sliding operation of the smart glasses. The generation of the timeline floating window comprises the following steps: the conference historical records are segmented according to semantics to form multiple historical record labels of time periods; an abstract text containing core data points and argument logic is generated for each historical record label; and each historical record label is arranged in a three-dimensional display space according to a conference process time sequence, and an importance level of each historical record label is marked. 2.The method of claim 1, wherein, The content to be played back is displayed in the form of a floating window according to a data type of the content to be played back, which comprises the following steps: Display information of a main display area of the smart glasses is used to determine display parameters of the content to be played back and a display position of a floating window corresponding to the content to be played back; A display mode of the content to be played back is determined according to the data type of the content to be played back; In the display position, the content to be played back in the floating window is rendered according to the display mode and the display parameters, and the rendered content to be played back is displayed. 3.The method of claim 1, wherein, The content to be played back includes audio translation content of a speaker in a conference, and the audio translation content is obtained according to the following manner: Three-dimensional space position of the speaker is determined according to visual data obtained by the smart glasses in real time, and a lip region video stream of the speaker is extracted; A beamforming parameter of a microphone array in the smart glasses is controlled according to the three-dimensional space position, a sound pickup beam pointing to the speaker is generated, and an audio signal of the speaker is collected; Lip movement features of the speaker are extracted according to the lip region video stream, and a time sequence synchronization score of the lip movement features and the audio signal is determined. input the audio signal and the lip movement feature into a pre-trained speech separation model when the timing synchronicity score is greater than a preset score, to obtain a target speech signal after separation of background noise output by the speech separation model; translate the target speech signal to obtain the audio translation content.
4. The information display optimization method of smart glasses according to claim 3, characterized in that, The translation of the target speech signal to obtain the audio translation content includes: extracting acoustic feature parameters of the target speech signal, and detecting emotional keywords of the translation text of the target speech signal to obtain text emotional features; performing tone analysis on the target speech signal according to the acoustic feature parameters and the text emotional features to obtain a tone analysis result; adding visual annotations to the translation text of the target speech signal according to the tone analysis result to obtain the audio translation content. 5.The method of claim 3, wherein, After inputting the audio signal and the lip movement feature into a pre-trained speech separation model when the timing synchronicity score is greater than a preset score, to obtain a target speech signal after separation of background noise output by the speech separation model, the method further includes: extracting a conclusive sentence from the target speech signal when the continuous speech duration of the target speech signal is greater than a second preset duration, or the number of characters in the transcribed text of the target speech signal is greater than a preset number of characters, to generate a summary card containing at least one key point; displaying the summary card in a floating form on the display interface of the smart glasses. 6.The method of claim 3, wherein, The to-be-played content further includes translation content corresponding to media content displayed by the speaker, and the translation content corresponding to the media content is obtained in the following manner: recognizing text regions and chart regions in the media content according to images of the media content obtained by the smart glasses; extracting text content from the text regions and performing image semantic understanding on the chart regions to obtain chart semantic features; inputting the target speech signal, the text content, and the chart semantic features into a multi-modal encoder for joint encoding to generate context-aware translation content. 7.An information display optimization apparatus of smart glasses, characterized by, The method includes: a data collection module configured to collect multi-modal interaction data of a user when the user uses smart glasses to participate in a conference; The multi-modal interaction data is interaction data between the user and the conference environment collected through multiple interaction modes; a retrieval module configured to retrieve a timeline floating window of the smart glasses to display a plurality of historical record tags if a backtracking function of the smart glasses is triggered according to the multi-modal interaction data; a to-be-played content acquisition module configured to acquire to-be-played content corresponding to a historical record tag selected by the user in response to a selection operation of the user on the timeline floating window; an optimized display module configured to display the to-be-played content in a floating window form according to a data type of the to-be-played content; and The calling module is further configured to, in a case where the multi-modal interaction data is voice data, identify text content and emotional tone information in the voice data; if the text content includes a trigger word of the backtracking function, or the emotional tone information indicates that the user is in a confused state, trigger the backtracking function; in a case where the multi-modal interaction data is eye movement data, acquire a visual line direction and visual line scanning information of the user according to the eye movement data; if the visual line direction deviates from a main display area for more than a first preset time length, and / or the visual line scanning information is disordered scanning information, trigger the backtracking function; in a case where the multi-modal interaction data is touch data, trigger the backtracking function according to a temple sliding operation of the smart glasses. The generation of the timeline floating window comprises: performing semantic-based topic segmentation on a conference history record to form a plurality of time period history record labels; generating an abstract text containing core data points and argument logic for each history record label; and arranging the history record labels in a three-dimensional display space according to conference process time sequence and marking an importance level of each history record label.
8. An intelligent eyeglass comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The computer program, when executed by the processor, implements the information display optimization method of the smart glasses according to any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the information display optimization method of the smart glasses according to any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent glasses control method, intelligent glasses, storage medium and program product
CN118658487A
Foveated beamforming for augmented reality devices and wearables
US20230071778A1
KR20230079846A