Conference information display method and device, electronic equipment and storage medium

By integrating a camera and microphone into smart glasses to collect handwriting, video, and audio information in multi-person meeting scenarios, and using a pre-trained model for feature extraction and fusion, the problem of low efficiency in handwriting recognition in meeting scenarios by smart glasses is solved, achieving more efficient and accurate handwriting recognition and prediction.

CN121859231APending Publication Date: 2026-04-14LUXSHARE PRECISION TECH(NANJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing smart glasses and other electronic terminal devices suffer from low efficiency and accuracy in handwriting recognition and prediction in meeting scenarios.

Method used

By acquiring current display information in multi-person meeting scenarios, including handwriting information of users wearing and not wearing handwriting, display video information, and meeting audio information, a pre-trained handwriting prediction model is used for recognition and prediction. Feature extraction and fusion are combined with a long short-term memory neural network to improve the accuracy of handwriting recognition and prediction.

Benefits of technology

It improves the efficiency and accuracy of handwriting recognition and prediction, and enhances the ability to display information in complex meeting scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859231A_ABST
    Figure CN121859231A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a conference information display method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring current display information in a multi-person conference scene; obtaining first handwriting information corresponding to the current wearing user, second handwriting information corresponding to the current display information, display video information corresponding to the current display information and conference audio information associated with the current display information; identifying the first handwriting information, and determining a display information processing instruction corresponding to the first handwriting information; based on the display information processing instruction, controlling a device worn by the current wearing user to perform marking processing on the current display information; inputting the second handwriting information, the display video information and the conference audio information into a writing prediction model to obtain a writing prediction result; and displaying the writing prediction result in the current display information. According to the technical scheme, the handwriting recognition and prediction efficiency and accuracy can be improved in a complex multi-person conference scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart device technology, and in particular to a method, apparatus, electronic device, and storage medium for displaying conference information. Background Technology

[0002] With the rapid development of artificial intelligence and wearable devices, products such as smart glasses and smart styluses are gradually entering scenarios such as education, office work, and remote collaboration. However, existing smart glasses and other electronic terminal devices generally require users to wear them and interact with the computer through touchpads, voice, or eye-tracking recognition. This allows the terminal device to send requests to the cloud or paired devices to achieve functions such as taking photos, navigation, and voice assistants. However, for handwriting recognition and prediction in meeting scenarios, the method of collecting image data of the user's handwritten content for recognition and display is usually done by collecting image data of the user's handwritten content. This method of recognition and prediction results in low efficiency and extremely low accuracy in handwriting recognition. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for displaying meeting information, which can improve the efficiency and accuracy of handwriting recognition and prediction in complex multi-person meeting scenarios.

[0004] According to one aspect of the present invention, an embodiment of the present invention provides a method for displaying meeting information, the method comprising:

[0005] Retrieve the current display information in a multi-person meeting scenario;

[0006] Obtain the first handwriting information corresponding to the currently wearing user, the second handwriting information corresponding to the currently displayed information, the display video information corresponding to the currently displayed information, and the meeting audio information associated with the currently displayed information;

[0007] The first handwriting information is identified to determine the display information processing instruction corresponding to the first handwriting information.

[0008] Based on the display information processing instructions, control the device currently worn by the user to perform marking processing on the current display information;

[0009] The second handwriting information, the display video information, and the conference audio information are input into a pre-trained writing prediction model for prediction and recognition to obtain the writing prediction result;

[0010] The writing prediction result is displayed in the current display information.

[0011] According to another aspect of the present invention, embodiments of the present invention also provide a conference information display device, the device comprising:

[0012] The first information acquisition module is used to acquire the currently displayed information in a multi-person meeting scenario;

[0013] The second information acquisition module is used to acquire the first handwriting information corresponding to the current wearer, the second handwriting information corresponding to the currently displayed information, the display video information corresponding to the currently displayed information, and the conference audio information associated with the currently displayed information;

[0014] The handwriting recognition module is used to recognize the first handwriting information and determine the display information processing instruction corresponding to the first handwriting information.

[0015] The tagging processing module is used to control the device currently worn by the user to perform tagging processing on the current display information based on the display information processing instructions;

[0016] The handwriting prediction module is used to input the second handwriting information, the display video information, and the conference audio information into a pre-trained handwriting prediction model for prediction and recognition, and to obtain the handwriting prediction result.

[0017] The prediction result display module is used to display the writing prediction result in the currently displayed information.

[0018] According to another aspect of the present invention, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0019] At least one processor; and

[0020] A memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the meeting information display method according to any embodiment of the present invention.

[0022] According to another aspect of the present invention, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the meeting information display method described in any embodiment of the present invention.

[0023] According to another aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the meeting information display method described in any embodiment of the present invention.

[0024] The above-described technical solution of this invention acquires the current display information in a multi-person conference scenario and identifies the first handwriting information of the current user to determine the display information processing instruction corresponding to the first handwriting information. Based on the display information processing instruction, the device worn by the current user is controlled to perform marking processing on the current display information. Based on this, the second handwriting information, display video information, and conference audio information of the current display information are also input into the writing prediction model to obtain the writing prediction result. The writing prediction result is displayed in the current display information. Based on the identification of the first handwriting information, and based on the second handwriting information and display video information of the current display information, the conference audio information is added to identify and predict the second handwriting information, thereby improving the efficiency and accuracy of handwriting recognition and prediction.

[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A flowchart illustrating a meeting information display method according to an embodiment of the present invention;

[0028] Figure 2 A flowchart illustrating another method for displaying meeting information according to an embodiment of the present invention;

[0029] Figure 3 This is a schematic diagram illustrating a dynamic handwriting prediction process according to an embodiment of the present invention;

[0030] Figure 4 This is a structural block diagram of a conference information display system provided in an embodiment of the present invention;

[0031] Figure 5 This is a structural block diagram of a conference information display device provided in an embodiment of the present invention;

[0032] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] In one embodiment, Figure 1 This is a flowchart of a meeting information display method according to an embodiment of the present invention. This embodiment is applicable to the display of meeting information in complex meeting scenarios. The method can be executed by a meeting information display device, which can be implemented in hardware and / or software. The meeting information display device can be configured in an electronic device. The electronic device in this embodiment can include, but is not limited to, smart wearable devices, such as smart glasses and other smart terminals.

[0036] like Figure 1 As shown, the method for displaying meeting information specifically includes:

[0037] S110: Obtain the current display information in a multi-person meeting scenario.

[0038] The information currently displayed can be understood as the content displayed in a multi-person meeting scenario. This content may include, but is not limited to, handwriting information, video information, and associated meeting audio information of users who are not wearing smart glasses.

[0039] In this embodiment, various smart devices such as smart glasses can integrate devices such as cameras, miniature displays, and microphones, thus enabling the acquisition of current display information in multi-person meeting scenarios.

[0040] S120. Obtain the first handwriting information corresponding to the current user, the second handwriting information corresponding to the currently displayed information, the display video information corresponding to the currently displayed information, and the conference audio information associated with the currently displayed information.

[0041] The first handwriting information can be understood as the handwriting information written by a user currently wearing a smart device (such as smart glasses) using a writing tool; the second handwriting information can be understood as the handwriting information written by one or more other users not currently wearing a smart device using a writing tool. The displayed video information can be understood as the video stream information related to the user generating the second handwriting information; the conference audio information can be understood as the audio information related to the user generating the second handwriting information.

[0042] In this embodiment, both the first and second handwriting information are obtained through writing with a writing tool. This writing tool can include a smart writing pen and a regular writing pen. If it is a smart writing pen, it can include integrated devices such as a pressure sensor, tilt sensor, Bluetooth module, and camera, which can be used to collect the handwriting information of the user and transmit the collected handwriting information to the relevant display device for display via Bluetooth. Thus, the linkage between the smart pen and smart glasses is also realized. It can be understood that the smart pen can collect handwriting information through its built-in sensors. Of course, if it is a regular handwriting tool, it does not have the function of collecting handwriting information.

[0043] In this embodiment, the smart wearable device, such as smart glasses, can acquire the first handwriting information generated by a user currently wearing the smart glasses using a writing tool in a multi-person meeting scenario. Alternatively, if the writing tool is a smart stylus, the first handwriting information can also be acquired through the stylus. In addition, it can acquire the current display information in the multi-person meeting scenario, namely, the second handwriting information generated by a user not wearing smart glasses using a writing tool, the second handwriting information corresponding to the current display information, the display video information, and the meeting audio information associated with the current display information. This is because smart wearable devices can integrate devices such as cameras, miniature displays, and microphones. In this case, handwriting information and video stream information can be acquired through the camera, and audio information can be acquired through the microphone, and then corresponding handwriting prediction can be performed. It should be noted that the video stream information is multi-frame video image information during writing, and each frame can represent image information under static writing conditions. Audio information can be non-verbal events such as writing sounds and tapping sounds during writing, or it can be voice information during writing; this embodiment does not impose any limitations.

[0044] S130. Identify the first handwriting information and determine the display information processing instruction corresponding to the first handwriting information.

[0045] Among them, the information processing instructions can be understood as the execution instructions represented by the content of the first handwriting information. For example, they can be opening meeting minutes or sharing data, and these written contents are some instructional contents.

[0046] In this embodiment, after collecting the first handwriting information corresponding to the current wearer, the integrated big data model can be used to perform real-time semantic analysis on the collected first handwriting information to obtain the analysis result corresponding to the handwriting information. The analysis result represents the complete content to be expressed by the first handwriting information. After obtaining the instruction content, the instruction is executed according to the instruction content. In other embodiments, the instruction content corresponding to the handwriting information can also be identified through keywords or preset trigger words in the handwriting information. For example, an instruction library can be established, and then the text information identified by handwriting can be matched with the instruction library to obtain the corresponding instruction content. Of course, in addition to the above two methods, other methods can also be used to identify the display information processing instructions corresponding to the first handwriting information. This embodiment does not impose specific limitations here.

[0047] S140. Based on the display information processing instruction, control the device currently worn by the user to perform marking processing on the currently displayed information.

[0048] The tagging process includes at least semantic segmentation and tag generation.

[0049] In this embodiment, when performing the labeling process of the currently displayed information, the relevant instruction content in the display information processing instruction can be determined, and then the device worn by the current user can be controlled to perform semantic segmentation processing on the currently displayed information according to the corresponding instruction content to obtain the corresponding semantic segmentation results. Based on this, keywords are extracted from each semantic segmentation result, and corresponding tags are generated based on the extracted keywords. Of course, in addition to the above method, other methods can also be used to perform the labeling process of the currently displayed information. This embodiment does not impose specific limitations here.

[0050] S150. Input the second handwriting information, the display video information, and the conference audio information into the pre-trained writing prediction model for prediction and recognition, and obtain the writing prediction result.

[0051] The pre-trained handwriting prediction model can include, but is not limited to, a long short-term memory neural network. In this embodiment, the long short-term memory neural network is a recurrent neural network that uses a gating mechanism consisting of a forget gate, an input gate, and an output gate, and incorporates cell states for dynamic handwriting prediction. This neural network model is a commonly used neural network model in the prior art, and will not be described in detail here.

[0052] In this embodiment, the displayed video information includes image information of each frame when the handwriting is generated; the conference audio information includes: mixed audio information of multiple users, that is, different voices may be emitted by different users, and it is necessary to distinguish the audio information corresponding to different users from the mixed audio; in this embodiment, the conference audio information can be used to match the timestamps of each image information in the displayed video information, which can be understood as the displayed video information can be timestamped with the image information, that is, in a smart synchronization association state.

[0053] In this embodiment, when performing real-time dynamic handwriting prediction, the handwriting trajectory with time-series characteristics can be obtained by analyzing the second handwriting information, the audio features with audio time-series characteristics can be obtained by analyzing the conference audio information, and the spatial image with spatial characteristics can be obtained by analyzing the display video information. Then, the obtained handwriting trajectory, audio features and spatial image are fused, and the fused features are input into the neural network model to obtain the dynamic handwriting prediction output. It can be understood that different feature processing stages can extract corresponding features through corresponding technical means, thereby performing subsequent feature fusion.

[0054] S160. Display the writing prediction results in the currently displayed information.

[0055] In this embodiment, after obtaining the writing prediction result, the obtained writing prediction result is displayed in the current display information.

[0056] The above-described technical solution of this invention acquires the current display information in a multi-person conference scenario and identifies the first handwriting information of the current user to determine the display information processing instruction corresponding to the first handwriting information. Based on the display information processing instruction, the device worn by the current user is controlled to perform marking processing on the current display information. Based on this, the second handwriting information, display video information, and conference audio information of the current display information are also input into the writing prediction model to obtain the writing prediction result. The writing prediction result is displayed in the current display information. Based on the identification of the first handwriting information, and based on the second handwriting information and display video information of the current display information, the conference audio information is added to identify and predict the second handwriting information, thereby improving the efficiency and accuracy of handwriting recognition and prediction.

[0057] In one embodiment, Figure 2This is a flowchart of another meeting information display method provided by an embodiment of the present invention. Based on the above embodiments, this embodiment identifies the first handwriting information and determines the display information processing instruction corresponding to the first handwriting information; based on the display information processing instruction, controls the device currently worn by the user to perform marking processing on the current display information; and inputs the second handwriting information, display video information, and meeting audio information into a pre-trained writing prediction model for prediction and recognition, and further refines the writing prediction results.

[0058] like Figure 2 As shown, the meeting information display method in this embodiment may specifically include the following steps:

[0059] S210: Obtain the current display information in a multi-person meeting scenario.

[0060] S220: Obtain the first handwriting information corresponding to the current user, the second handwriting information corresponding to the currently displayed information, the display video information corresponding to the currently displayed information, and the conference audio information associated with the currently displayed information.

[0061] S230. Use a large language model to perform semantic parsing on the first handwriting information, obtain the corresponding semantic parsing result, and determine the display information processing instruction corresponding to the first handwriting information based on the semantic parsing result.

[0062] In this embodiment, a large-scale artificial intelligence speech model can be used to perform semantic analysis on the first handwriting information to obtain the corresponding semantic analysis result. Based on the semantic analysis result, the display information processing instruction corresponding to the first handwriting information is determined. This can be understood as the display information processing instruction representing the command issued by the user. The instruction content is then executed, and after the instruction content is executed, the result of the executed instruction is fed back. For example, if the handwriting is semantically analyzed and found to be shared data, the shared data is used as an instruction for execution, and the result is fed back.

[0063] S240. According to the instruction content of the display information processing instruction, control the device currently worn by the user to perform semantic segmentation processing on the currently displayed information to obtain the corresponding semantic segmentation result.

[0064] In this embodiment, by displaying the instruction content of the information processing instruction, the device currently worn by the user can be controlled to perform semantic segmentation processing on the currently displayed information to obtain the corresponding semantic segmentation result. Specifically, for the text content of the instruction, layout analysis can be performed first, and then logical semantic segmentation can be performed by combining the layout information and the text content to obtain the corresponding semantic segmentation result.

[0065] S250. Extract keywords from each semantic segmentation result, generate corresponding tags based on the extracted keywords, and display the tags in a floating format.

[0066] In this embodiment, keywords are extracted for each semantic segmentation result, and all keywords are analyzed to generate corresponding tags. These tags can be displayed in a floating format. Of course, tags can also be displayed in other ways besides this method, and this embodiment does not impose specific limitations. This can be understood as real-time detection of handwriting and commands, semantic segmentation and tag generation of chat content, such as automatically recognizing "task reminders," "data sharing," and "schedules," and displaying them as floating tags to improve users' information acquisition efficiency in complex meeting dialogue scenarios.

[0067] S260. Extract motion trajectory features from the second handwriting information to obtain handwriting trajectory features.

[0068] In this embodiment, the second handwriting information is input into a preset first neural network model for motion trajectory feature extraction, resulting in a handwriting trajectory with time-series characteristics. The preset first neural network model can be a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU). The preset first neural network model can construct a continuous sequence of points from the handwritten handwriting; this can be understood as the direction and speed of a stroke being a sequence of continuous points. For example, for the trajectory coordinate sequence of the stylus, LSTM / GRU can learn the temporal dynamic features such as stroke order, rhythm, and continuity. For instance, it can identify whether "A" or "V" is being written, even though the final static strokes may be similar, the temporal trajectories during writing are different.

[0069] S270. Extract text features from the displayed video information to obtain written text features.

[0070] In this embodiment, each frame of the displayed video information is input into a preset second neural network model to extract text features from the images, thereby obtaining written text features with spatial characteristics. The preset second neural network model can be a Convolutional Neural Network (CNN) used to extract spatial hierarchical structure features. In this embodiment, the CNN can extract spatial features such as the shape, structure, and layout of characters.

[0071] In one embodiment, text feature extraction is performed on the displayed video information to obtain written text features, including: extracting features from the text regions on each video frame in the displayed video information to obtain text region features corresponding to each video frame; extracting inter-frame features from the text region features corresponding to each adjacent video frame to obtain inter-frame text features corresponding to adjacent video frames; and using the text region features and inter-frame text features as written text features.

[0072] In this embodiment, when extracting text features, the text regions on each video frame of the displayed video information are first extracted. Then, inter-frame feature extraction is performed on the text region features corresponding to adjacent video frames. Specifically, feature extraction for each video frame is performed first, followed by inter-frame text feature extraction. The text region features on each video frame include: first extracting features from the text regions on each video frame of the displayed video information to obtain the text region features corresponding to each video frame. The inter-frame text feature extraction includes: performing inter-frame feature extraction on the text region features corresponding to adjacent video frames to obtain the inter-frame text features corresponding to adjacent video frames. Thus, the text region features and inter-frame text features are used as written text features.

[0073] S280. Extract audio features from the conference audio information to obtain the conference audio features.

[0074] In this embodiment, a preset third neural network model is used to extract audio features from the conference audio information to obtain conference audio features with audio temporal characteristics. The preset third neural network model can be a combination of a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) network. In this embodiment, during the analysis of conference audio information, since it is mixed audio information, the mixed audio information from multiple users can first be separated. For the separated audio information, large language model voiceprint recognition technology is used to generate voiceprint tags corresponding to different speaking users, thereby obtaining audio features with audio temporal characteristics.

[0075] In one embodiment, after extracting audio features from the conference audio information to obtain the conference audio features, the method further includes:

[0076] Voiceprint recognition is performed on the audio features of the meeting to determine the audio features of each speaker.

[0077] The system performs writing sound recognition on the meeting audio features to identify writing audio features associated with writing actions.

[0078] In this embodiment, after extracting audio features from the meeting audio information, it is necessary to perform voiceprint recognition and handwriting sound recognition on the meeting audio features to determine the speaker audio features corresponding to different speakers, and to determine the handwriting audio features associated with the handwriting action. This can be understood as follows: the microphone in the smart glasses can recognize near and far sounds, separating the collected mixed audio; large language model voiceprint recognition technology is used to distinguish different speakers in the audio information, generating a unique voiceprint tag for each speaker; simultaneously, through timestamp alignment and semantic analysis, the text content recognized from the image is intelligently synchronized and associated with the corresponding speaker audio stream. In this embodiment, the voiceprint tag can first be input into a CNN for initial audio feature extraction and dimensionality reduction, and then LSTM is used to learn the long-range context in the audio to understand words, sentences, and semantic content. This can be understood as using a combination of CNN and LSTM to extract speech content features from the audio; CNN captures phonemes, LSTM understands words and sentences, and finally, the speech information is aligned and fused with the handwritten note content.

[0079] S290. The handwriting trajectory features, written text features, and conference audio features are fused to obtain the target fused features.

[0080] In this embodiment, after obtaining the handwriting trajectory features, written text features, and conference audio features, the handwriting trajectory features, written text features, and conference audio features are fused to obtain fused features. These fused features are then input into a preset long short-term memory neural network model to obtain a dynamic handwriting prediction output. Specifically, the feature fusion method can be to directly concatenate the handwriting trajectory features, written text features, and conference audio features, or to set corresponding weights for weighted fusion, or to use an attention-based fusion mechanism. This embodiment does not impose specific limitations on these methods.

[0081] In one embodiment, the handwriting trajectory features, written text features, and conference audio features are fused to obtain the target fused features, including:

[0082] The handwriting trajectory features, written text features, speaker audio features, and written audio features are fused to obtain the target fused features.

[0083] In this embodiment, after extracting audio features from the meeting audio information and obtaining the meeting audio features, speaker audio features and written audio features are then extracted. Therefore, when fusing the target fusion features, the speaker audio features and written audio features can be added to the feature fusion process to obtain the target fusion features.

[0084] In one embodiment, the handwriting trajectory features, written text features, speaker audio features, and written audio features are fused to obtain the target fused features, including:

[0085] Semantic recognition is performed based on the features of written text and the speaker's audio features to obtain semantic recognition results;

[0086] Based on the semantic recognition results, semantic association is performed on the written text features and the speaker's audio features to obtain semantic association features;

[0087] The handwriting trajectory features, written text features, speaker audio features, written audio features, and semantic association features are fused to obtain the target fused features.

[0088] In this embodiment, when performing target feature fusion, semantic association features can be added. Specifically, semantic recognition is first performed on the written text features and the speaker's audio features to obtain semantic recognition results. Based on the semantic recognition results, semantic association is performed on the written text features and the speaker's audio features to obtain semantic association features. The handwriting trajectory features, written text features, speaker's audio features, written audio features, and semantic association features are then fused to obtain the target fusion features.

[0089] S2100. Perform handwriting prediction on the target fusion features to obtain the handwriting prediction result.

[0090] In this embodiment, the target fusion features are input into the trained writing prediction model for prediction and recognition to obtain the writing prediction result.

[0091] S2110. Display the writing prediction results in the currently displayed information.

[0092] In one embodiment, to facilitate a better understanding of the output of the dynamic handwriting prediction results, Figure 3 This is a schematic diagram illustrating a dynamic handwriting prediction process according to an embodiment of the present invention. In this embodiment, an example is given using a preset first neural network model of LSTM or GRU, a preset second neural network model of CNN, and a preset third neural network model of CNN+LSTM. Figure 3As shown, the big data model collects data including handwritten note trajectories, images recognized by the camera of smart glasses, and voice information collected by the microphone; then, through big data preprocessing and semantic analysis, it predicts fragments of text or handwriting information. Specifically, the handwriting prediction process mainly involves: displaying the data stream, feature extraction, fusion method, and output prediction. First, feature acquisition is performed, which can include the trajectory (x, y, timestamp), image (handwriting), and audio (pen grip / writing sound). LSTM or GRU is used to analyze the second handwriting information to obtain a handwriting trajectory with time-series features. CNN is used to analyze the video stream information to obtain a spatial image with spatial features. CNN+LSTM is used to analyze the audio information to obtain audio features with audio time-series features. The handwriting trajectory, spatial image, and audio features are fused to obtain fused features, which are then input into a preset long short-term memory neural network model to obtain a dynamic handwriting prediction output. In essence, LSTM / GRU can be used to extract time-series features, CNN can be used to extract spatial features, and 1D-CNN+LSTM can be used to extract audio time-series features. After processing by the artificial neural network, dynamically predicted handwriting information is generated.

[0093] It should be noted that, Figure 3 The feature weight training part is the process of adjusting the weights of each feature during the training of each neural network model. The trained neural network model is then used for the extraction of each feature and the output of dynamic handwriting prediction.

[0094] The above-mentioned technical solution in this embodiment extracts motion trajectory features from the second handwriting information to obtain handwriting trajectory features; extracts text features from the displayed video information to obtain written text features; extracts audio features from the conference audio information to obtain conference audio features; fuses the handwriting trajectory features, written text features, and conference audio features to obtain target fusion features; performs handwriting prediction on the target fusion features to obtain writing prediction results, and displays the writing prediction results in the current displayed information. Furthermore, based on the recognition of the first handwriting information, and on the basis of the second handwriting information and the displayed video information, the conference audio information is added to recognize and predict the second handwriting information, improving the efficiency and accuracy of handwriting recognition and prediction.

[0095] In one embodiment, Figure 4 This is a structural block diagram of a conference information display system according to an embodiment of the present invention. In this embodiment, smart glasses are used as an example of a terminal device for illustration. Figure 4As shown, the system includes smart glasses (the data acquisition layer and intelligent linkage platform can be integrated into the smart glasses), a writing pen (which can include a smart writing pen and a regular writing pen; if it is a smart writing pen, it can include an integrated pressure sensor, tilt sensor, Bluetooth module, and camera, connecting to a display device via Bluetooth to transmit handwriting data to an application), and a smartphone or display device, etc.; the intelligent linkage platform in this embodiment can include a natural language processing model, a handwriting prediction engine, a structured instruction generation system, and a multi-device instruction manager. The multi-device instruction manager can directly display the predicted handwriting results and the executed instructions in real time, and can also call external devices, etc.

[0096] In this embodiment, the data acquisition layer in the smart glasses collects handwriting information written by the user wearing the smart glasses, or handwriting information written by a pen used by another non-smart glasses wearer. If the handwriting information is collected from the user wearing the smart glasses, the intelligent linkage platform determines the corresponding instruction content, executes the instruction content, and then provides feedback on the executed instruction result. If the handwriting information is collected from a pen used by another non-smart glasses wearer, the intelligent linkage platform performs real-time dynamic handwriting prediction.

[0097] In one embodiment, Figure 5 This is a structural block diagram of a meeting information display device according to an embodiment of the present invention. This device is suitable for displaying meeting information in complex meeting scenarios and can be implemented in hardware or software. It can be configured in a terminal device to implement a meeting information display method according to an embodiment of the present invention. Figure 5 As shown, the device includes: a first information acquisition module 510, a second information acquisition module 520, a handwriting recognition module 530, a marker processing module 540, a handwriting prediction module 550, and a prediction result display module 560.

[0098] The first information acquisition module 510 is used to acquire the current display information in a multi-person conference scenario;

[0099] The second information acquisition module 520 is used to acquire the first handwriting information corresponding to the current wearer, the second handwriting information corresponding to the current display information, the display video information corresponding to the current display information, and the conference audio information associated with the current display information;

[0100] The handwriting recognition module 530 is used to recognize the first handwriting information and determine the display information processing instruction corresponding to the first handwriting information.

[0101] The tagging processing module 540 is used to control the device currently worn by the user to perform tagging processing on the current display information based on the display information processing instruction;

[0102] The handwriting prediction module 550 is used to input the second handwriting information, the display video information and the conference audio information into a pre-trained handwriting prediction model for prediction and recognition, and to obtain the handwriting prediction result.

[0103] The prediction result display module 560 is used to display the writing prediction result in the current display information.

[0104] In this embodiment of the invention, a handwriting recognition module acquires the current display information in a multi-person meeting scenario and recognizes the first handwriting information of the currently wearing user to determine the display information processing instruction corresponding to the first handwriting information. A marking processing module, based on the display information processing instruction, controls the device worn by the currently wearing user to perform marking processing on the current display information. Based on this, a handwriting prediction module also inputs the second handwriting information of the current display information, the display video information, and the meeting audio information into a handwriting prediction model to obtain a handwriting prediction result. A prediction result display module displays the handwriting prediction result in the current display information. This invention can recognize and predict the second handwriting information by adding meeting audio information to the first handwriting information, the second handwriting information of the current display information, and the display video information, thereby improving the efficiency and accuracy of handwriting recognition and prediction.

[0105] In one embodiment, the handwriting prediction module 550 includes:

[0106] The handwriting feature extraction unit is used to extract motion trajectory features from the second handwriting information to obtain handwriting trajectory features;

[0107] The text feature unit is used to extract text features from the displayed video information to obtain written text features;

[0108] An audio feature extraction unit is used to extract audio features from the conference audio information to obtain conference audio features;

[0109] The feature fusion unit is used to fuse the handwriting trajectory features, the written text features, and the conference audio features to obtain the target fused features;

[0110] The writing prediction unit is used to perform handwriting prediction on the target fusion features to obtain the writing prediction result.

[0111] In one embodiment, the method further includes: an identification module, configured to perform voiceprint recognition on the meeting audio features after extracting audio features from the meeting audio information to obtain meeting audio features, and determine the speaker audio features corresponding to each speaker from the meeting audio features;

[0112] The meeting audio features are used to perform writing sound recognition, and writing audio features associated with writing actions are determined from the meeting audio features;

[0113] Correspondingly, the feature fusion unit includes:

[0114] The feature fusion subunit is used to fuse the handwriting trajectory features, the written text features, the speaker audio features, and the written audio features to obtain the target fused features.

[0115] In one embodiment, the feature fusion subunit is specifically used for:

[0116] Semantic recognition is performed based on the written text features and the speaker's audio features to obtain semantic recognition results;

[0117] Based on the semantic recognition results, semantic association is performed on the written text features and the speaker audio features to obtain semantic association features;

[0118] The handwriting trajectory features, the written text features, the speaker audio features, the written audio features, and the semantic association features are fused to obtain the target fused features.

[0119] In one embodiment, the text feature unit includes:

[0120] The video frame feature extraction subunit is used to extract features from the text regions on each video frame in the displayed video information to obtain the text region features corresponding to each video frame.

[0121] The inter-frame feature extraction subunit is used to extract inter-frame features from the text region features corresponding to each of the adjacent video frames to obtain the inter-frame text features corresponding to the adjacent video frames.

[0122] The text feature extraction subunit is used to extract the text region features and the inter-frame text features as the written text features.

[0123] In one embodiment, the handwriting recognition module 530 includes:

[0124] The parsing unit is used to perform semantic parsing on the first handwriting information using a large language model, obtain the corresponding semantic parsing result, and determine the display information processing instruction corresponding to the first handwriting information based on the semantic parsing result.

[0125] In one embodiment, the tagging processing module 540 includes:

[0126] The semantic segmentation unit is used to control the device worn by the current user to perform semantic segmentation processing on the currently displayed information according to the instruction content of the display information processing instruction, so as to obtain the corresponding semantic segmentation result;

[0127] The tag generation unit is used to extract keywords from each of the semantic segmentation results, generate corresponding tags based on the extracted keywords, and display the tags in a floating format.

[0128] The meeting information display device provided in the embodiments of the present invention can execute the meeting information display method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0129] In one embodiment, Figure 6 This is a schematic diagram of an electronic device provided for an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0130] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0131] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows terminal device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0132] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as conference information display methods.

[0133] In some embodiments, the conference information display method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the conference information display method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the conference information display method by any other suitable means (e.g., by means of firmware).

[0134] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0135] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable conferencing display device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0136] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0138] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0139] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0140] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0141] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for displaying meeting information, characterized in that, The method includes: Retrieve the current display information in a multi-person meeting scenario; Obtain the first handwriting information corresponding to the currently wearing user, the second handwriting information corresponding to the currently displayed information, the display video information corresponding to the currently displayed information, and the meeting audio information associated with the currently displayed information; The first handwriting information is identified to determine the display information processing instruction corresponding to the first handwriting information. Based on the display information processing instructions, control the device currently worn by the user to perform marking processing on the current display information; The second handwriting information, the display video information, and the conference audio information are input into a pre-trained writing prediction model for prediction and recognition to obtain the writing prediction result; The writing prediction result is displayed in the current display information.

2. The method according to claim 1, characterized in that, The step of inputting the second handwriting information, the display video information, and the conference audio information into a pre-trained handwriting prediction model for prediction and recognition, and obtaining the handwriting prediction result includes: Motion trajectory features are extracted from the second handwriting information to obtain handwriting trajectory features; Text features are extracted from the displayed video information to obtain written text features; Audio features are extracted from the conference audio information to obtain the conference audio features; The handwriting trajectory features, the written text features, and the conference audio features are fused to obtain the target fused features; Handwriting prediction is performed on the target fusion features to obtain the handwriting prediction result.

3. The method according to claim 2, characterized in that, After extracting audio features from the conference audio information to obtain the conference audio features, the method further includes: Voiceprint recognition is performed on the meeting audio features to determine the speaker audio features corresponding to each speaker from the meeting audio features; The meeting audio features are used to perform writing sound recognition, and writing audio features associated with writing actions are determined from the meeting audio features; Accordingly, the process of fusing the handwriting trajectory features, the written text features, and the conference audio features to obtain the target fused features includes: The handwriting trajectory features, the written text features, the speaker audio features, and the written audio features are fused together to obtain the target fused features.

4. The method according to claim 3, characterized in that, The process of fusing the handwriting trajectory features, the written text features, the speaker audio features, and the written audio features to obtain the target fused features includes: Semantic recognition is performed based on the written text features and the speaker's audio features to obtain semantic recognition results; Based on the semantic recognition results, semantic association is performed on the written text features and the speaker audio features to obtain semantic association features; The handwriting trajectory features, the written text features, the speaker audio features, the written audio features, and the semantic association features are fused to obtain the target fused features.

5. The method according to claim 2, characterized in that, The step of extracting text features from the displayed video information to obtain written text features includes: Feature extraction is performed on the text regions in each video frame of the displayed video information to obtain the text region features corresponding to each video frame; Inter-frame feature extraction is performed on the text region features corresponding to each of the adjacent video frames to obtain the inter-frame text features corresponding to the adjacent video frames; The text region features and the inter-frame text features are used as the written text features.

6. The method according to claim 1, characterized in that, The step of identifying the first handwriting information and determining the display information processing instruction corresponding to the first handwriting information includes: The first handwriting information is semantically parsed to obtain the corresponding semantic parsing result, and the display information processing instruction corresponding to the first handwriting information is determined based on the semantic parsing result.

7. The method according to claim 1, characterized in that, The step of controlling the device currently worn by the user to perform marking processing on the current display information based on the display information processing instructions includes: According to the instruction content of the display information processing instruction, the device currently worn by the user is controlled to perform semantic segmentation processing on the currently displayed information to obtain the corresponding semantic segmentation result; Keyword extraction is performed on each of the semantic segmentation results, and corresponding tags are generated based on the extracted keywords, which are then displayed in a floating format.

8. A conference information display device, characterized in that, The device includes: The first information acquisition module is used to acquire the currently displayed information in a multi-person meeting scenario; The second information acquisition module is used to acquire the first handwriting information corresponding to the current wearer, the second handwriting information corresponding to the currently displayed information, the display video information corresponding to the currently displayed information, and the conference audio information associated with the currently displayed information; The handwriting recognition module is used to recognize the first handwriting information and determine the display information processing instruction corresponding to the first handwriting information. The tagging processing module is used to control the device currently worn by the user to perform tagging processing on the current display information based on the display information processing instructions; The handwriting prediction module is used to input the second handwriting information, the display video information, and the conference audio information into a pre-trained handwriting prediction model for prediction and recognition, and to obtain the handwriting prediction result. The prediction result display module is used to display the writing prediction result in the currently displayed information.

9. An electronic device, characterized in that, The terminal device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the conference information display method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the conference information display method according to any one of claims 1-7.