Ssummary generation method, device, intelligent glasses, server and system

Smart glasses automatically collect information through audio and image acquisition units, and combine visual and auditory generation minutes to solve the problems of incomplete information recording and low efficiency in the prior art, achieving efficient and accurate minutes generation.

CN120455186APending Publication Date: 2025-08-08BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510848768.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, information recording in conferences, classrooms and other scenarios is difficult to ensure the integrity and accuracy of information, and the collation efficiency is low.

Method used

The audio and image acquisition unit of the smart glasses automatically collects audio information and target images under the preset conditions, uses automatic speech recognition and image processing technology to generate minutes, and combine visual and auditory information to generate more comprehensive and accurate minutes.

Benefits of technology

It improves the convenience and efficiency of information recording, and the generated minutes can promptly reflect the real situation on site, avoid information loss or deviation, and reduce manual sorting time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455186A_ABST
    Figure CN120455186A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, in particular to a summary generation method and device, intelligent glasses, a server and a system. The method is applied to the intelligent glasses and comprises the following steps: acquiring audio information by using an audio acquisition unit of the intelligent glasses; in the process of collecting the audio information, in response to determining that a preset condition is met, collecting a target image by using an image collection unit of the intelligent glasses; and obtaining a transcriptional text according to the audio information, and generating a summary according to the target image and the transcriptional text. Thus, automatic collection of the audio information and the target image can be realized by using the intelligent glasses, the recording convenience and efficiency can be improved, and through combination of sound and vision, a summary which can more comprehensively and accurately reflect the real situation of the scene can be generated, and the summary generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and in particular to a minutes generation method, device, smart glasses, server, and system. Background Art

[0002] In today's society, meetings, classrooms, interviews, and other scenarios permeate our daily work and study lives, often involving the transmission and exchange of large amounts of important information. Currently, minutes are manually compiled by relevant personnel, but this method struggles to ensure the completeness and accuracy of the information recorded, and is also inefficient. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides a minutes generation method, device, smart glasses, server and system.

[0004] According to a first aspect of an embodiment of the present disclosure, a method for generating minutes is provided, which is applied to smart glasses and includes: Collecting audio information using the audio collection unit of the smart glasses; In the process of collecting the audio information, in response to determining that a preset condition is satisfied, collecting a target image using an image collection unit of the smart glasses; A transcribed text is obtained according to the audio information, and a summary is generated according to the target image and the transcribed text.

[0005] According to a second aspect of an embodiment of the present disclosure, a method for generating minutes is provided, which is applied to a server and includes: receiving audio information and a target image uploaded by the smart glasses, wherein the target image is an image captured by the image capture unit of the smart glasses when the audio capture unit of the smart glasses is collecting the audio information and a preset condition is met; Obtaining a transcribed text based on the audio information, and generating a summary based on the target image and the transcribed text; The minutes are issued so that a device receiving the minutes can display the minutes.

[0006] According to a third aspect of an embodiment of the present disclosure, a minutes generation device is provided, which is applied to smart glasses and is used to implement the minutes generation method provided by the first aspect of the present disclosure.

[0007] According to a fourth aspect of an embodiment of the present disclosure, a minutes generation device is provided, which is applied to a server and is used to implement the minutes generation method provided in the second aspect of the present disclosure.

[0008] According to a fifth aspect of the embodiments of the present disclosure, there is provided a pair of smart glasses, including: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the minutes generation method provided in the first aspect of the present disclosure.

[0009] According to a sixth aspect of an embodiment of the present disclosure, a server is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the minutes generation method provided in the second aspect of the present disclosure.

[0010] According to a seventh aspect of an embodiment of the present disclosure, a minutes generation system is provided, including the smart glasses provided in the fifth aspect of the present disclosure.

[0011] According to an eighth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the minutes generation method provided in the first aspect of the present disclosure or the second aspect of the present disclosure are implemented.

[0012] According to a ninth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the minutes generation method provided in the first aspect of the present disclosure or the second aspect of the present disclosure.

[0013] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects: Smart glasses, as wearable devices, are convenient for users to use in a variety of scenarios. Users no longer need to carry multiple recording devices. Smart glasses can automatically capture audio information and target images during meetings or events, improving the convenience and efficiency of recording. A transcript is then generated based on the audio information, and minutes are generated based on the target image and transcript. This combination of sound and vision allows for the creation of minutes that more comprehensively and accurately reflect the actual situation on site, avoiding the potential for missing or biased information from a single source. This also effectively improves the efficiency of minute generation, eliminating the need for users to spend significant time manually organizing minutes after a meeting or event, allowing for timely access to the minutes.

[0014] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0016] Figure 1 The figure is a flowchart of a method for generating minutes according to an exemplary embodiment.

[0017] Figure 2 The figure is a schematic diagram showing the interaction between smart glasses and a display terminal according to an exemplary embodiment.

[0018] Figure 3 The figure is a schematic diagram showing the interaction among smart glasses, a display terminal and a server according to an exemplary embodiment.

[0019] Figure 4 The figure is a flowchart of a method for generating minutes according to an exemplary embodiment.

[0020] Figure 5 The figure is a flowchart of a method for generating minutes according to an exemplary embodiment.

[0021] Figure 6 The figure is a flowchart of a method for generating minutes according to an exemplary embodiment.

[0022] Figure 7 The figure is a block diagram of a device for generating minutes according to an exemplary embodiment.

[0023] Figure 8 The figure is a block diagram of a device for generating minutes according to an exemplary embodiment.

[0024] Figure 9 The figure is a block diagram of smart glasses according to an exemplary embodiment.

[0025] Figure 10 The figure is a block diagram of a server according to an exemplary embodiment.

[0026] Figure 11 The figure is a block diagram of a system for generating minutes according to an exemplary embodiment. DETAILED DESCRIPTION

[0027] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0028] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0029] Figure 1 This is a flow chart of a method for generating minutes according to an exemplary embodiment. The method can be applied to smart glasses. Smart glasses have an audio acquisition unit and an image acquisition unit, such as Figure 1 As shown, the method may include steps S101 to S103.

[0030] In step S101, audio information is collected using the audio collection unit of the smart glasses.

[0031] For example, the audio collection unit can be a sound pickup device, such as a microphone, which can be installed on the frame or the temple of the smart glasses and can be distributed in an array. The collected audio information can be subjected to noise reduction processing to filter out ambient noise to improve the quality of the audio information.

[0032] In step S102 , during the process of collecting audio information, in response to determining that a preset condition is met, the target image is collected by using the image collection unit of the smart glasses.

[0033] For example, the image acquisition unit may be a camera. The smart glasses may also pre-process the acquired target image, such as performing operations such as denoising and tilt correction, to remove interference factors in the image and improve image clarity and readability.

[0034] During the audio information acquisition process, in response to determining that preset conditions are met, the target image is captured using the image acquisition unit of the smart glasses. This allows for automated and simultaneous capture of the target image during the audio information acquisition process, improving recording convenience and efficiency. Furthermore, by setting preset conditions, unnecessary image capture can be avoided.

[0035] In some possible implementations, the preset condition includes: the current screen includes a preset information presentation interface, and the information presentation interface includes text and / or graphics.

[0036] The information presentation interface may include at least one of a notebook page, a whiteboard, and a presentation. In this way, the smart glasses can automatically capture the target image during the audio collection process, improving the device's automation level. This ensures that the captured target image contains the information required for the minutes generation process, reducing the collection of irrelevant images.

[0037] In some possible implementations, the preset condition includes: a photo button is pressed. In this way, the user's image acquisition intention can be captured by utilizing the photo button state to meet the user's personalized needs.

[0038] In some possible implementations, the area where the image capturing unit captures images is determined according to the gaze point of the target user.

[0039] Eye tracking technology can be used to determine the target user's gaze point. For example, the area of image capture can be centered around the gaze point. This allows the user's focus to be accurately captured, improving the relevance and effectiveness of image capture, reducing irrelevant information interference, and ultimately enhancing the user experience.

[0040] In step S103, a transcribed text is obtained according to the audio information, and a summary is generated according to the target image and the transcribed text.

[0041] In some possible implementations, automatic speech recognition (ASR) technology is used to transcribe audio information to generate a transcript. For example, deep learning models (such as Whisper or Conformer) can be used to generate the transcript. Subsequently, the target image and the transcript can be combined to generate a summary.

[0042] In the above technical solution, smart glasses serve as a wearable device, making them convenient for users to use in various scenarios. Users no longer need to carry multiple recording devices. While attending meetings or events, they can use smart glasses to automatically capture audio information and target images, thus improving the convenience and efficiency of recording. A transcript is then generated based on the audio information, and minutes are generated based on the target image and transcript. This combination of sound and vision allows for the generation of minutes that more comprehensively and accurately reflect the actual situation on site, avoiding information gaps or biases that may arise from a single information source. This also effectively improves the efficiency of minute generation, eliminating the need for users to spend significant time manually organizing minutes after a meeting or event, allowing for timely access to the minutes.

[0043] In some possible implementations, the smart glasses execute step S101 when any of the following conditions is met: receiving a voice command instructing content recording; It is determined that the target user performs a target operation on a preset area of the smart glasses.

[0044] For example, if a user issues a voice command to "start recording," the audio information can be collected using the smart glasses' audio acquisition unit. This allows for quick audio information collection via voice commands without the user touching the smart glasses, improving operational convenience and user experience. Furthermore, voiceprint recognition can be performed on the received voice command to determine whether the user issuing the voice command is the legitimate owner of the smart glasses, thereby preventing erroneous audio information collection.

[0045] For example, the preset area can be a touch area on the smart glasses, and the corresponding target operation can be a preset gesture operation, such as double-clicking the touch area or touching the touch area for a preset duration. As another example, the preset area can be a button on the smart glasses, and the corresponding target operation can be pressing the button. This allows accurate determination of the user's recording intent in a noisy environment, thereby enabling the capture of audio information.

[0046] It should also be noted that the step of starting audio information collection strictly follows the privacy policy. This function is officially launched when the user explicitly authorizes it and the system confirms that the current environment meets security standards.

[0047] Combining the above-mentioned methods to enable the collection of audio information can better adapt to the needs and usage habits of different users, thereby improving the intelligence of smart glasses and the user experience.

[0048] In some possible implementations, the processing unit of the smart glasses may be used to perform the steps of obtaining a transcribed text based on the audio information and generating a minutes based on the target image and the transcribed text. In this case, the minutes generation method provided by the present disclosure may further include: The minutes are displayed using the smart glasses, or the minutes are sent to a terminal bound to the smart glasses using a communication unit of the smart glasses, so that the minutes are displayed using the terminal.

[0049] For example, the processing unit can be a chip within the smart glasses. If the smart glasses have voice broadcast and / or visualization functions, the minutes can be displayed visually or broadcasted by voice to facilitate user access to the minutes. Alternatively, the communication unit of the smart glasses can be used to send the minutes to a terminal connected to the smart glasses, such as a smartphone or smart tablet, and the minutes can be displayed on the terminal connected to the smart glasses for the user to view the minutes.

[0050] Figure 2 This is a schematic diagram of interaction between smart glasses and a display terminal (a terminal bound to the smart glasses) according to an exemplary embodiment. Figure 2, we can more clearly understand the process of generating and displaying the minutes of the present disclosure when the smart glasses execute the step of obtaining the transcribed text according to the audio information and generating the minutes according to the target image and the transcribed text. Figure 2 As shown, the method may include steps S201 to S206.

[0051] In step S201, the smart glasses receive a voice command of "start recording" to start the recording function and collect audio information.

[0052] In step S202 , during the process of collecting audio information, the smart glasses collect a target image when a preset condition is met.

[0053] In step S203, the smart glasses transcribe the audio information to obtain a transcribed text.

[0054] In step S204 , the smart glasses generate minutes based on the target image and the transcribed text.

[0055] In step S205, the smart glasses send the minutes.

[0056] In step S206, the display terminal receives the minutes sent by the smart glasses and displays the minutes.

[0057] In this way, there is no need to upload and wait for server responses, which can reduce the time required for minutes generation and reduce the risk of data leakage. In addition, it can also ensure the offline availability of the minutes generation function.

[0058] In some possible implementations, the smart glasses may be used to perform the following steps to obtain a transcript based on the audio information and generate a summary based on the target image and the transcript: The communication unit of the smart glasses is used to send the audio information and the target image to the server, which is used to obtain a transcribed text based on the audio information, and generate and issue a minutes based on the target image and the transcribed text, so that the device receiving the minutes can display the minutes.

[0059] The device that receives the minutes includes smart glasses and / or a terminal bound to the smart glasses.

[0060] Figure 3 This is a schematic diagram showing the interaction between smart glasses, a display terminal (a terminal bound to the smart glasses) and a server according to an exemplary embodiment. Figure 3 , we can more clearly understand the generation and display process of the minutes of the present disclosure when the server executes the step of obtaining the transcription text according to the audio information and generating the minutes according to the target image and the transcription text. Figure 3 As shown, the method may include steps S301 to S312.

[0061] In step S301, the smart glasses receive a voice command of "start recording" and send the voice command to the server.

[0062] In step S302, the server understands the received voice instruction of "start recording" and generates an instruction to start the recording function.

[0063] In step S303, the server sends an instruction to start the recording function to the smart glasses.

[0064] In step S304, if an instruction to start the recording function is received, the smart glasses start the recording function and collect audio information.

[0065] In step S305, the smart glasses receive a transcription instruction or a stop recording instruction and upload the audio information to the server.

[0066] When the smart glasses receive a transcription command, they upload the audio information to the server, allowing users to choose when to upload the audio information based on their needs. When the smart glasses receive a stop recording command, they also upload the audio information to the server. This automatic upload simplifies the user interaction process and ensures the integrity of the recorded content.

[0067] In step S306, the server transcribes the audio information to obtain a transcribed text.

[0068] In step S307, the server sends the transcribed text to the display terminal.

[0069] In step S308 , during the audio information collection process, the smart glasses collect the target image when the preset conditions are met.

[0070] In step S309 , the smart glasses receive the minutes generation instruction and upload the target image to the server.

[0071] In step S310 , the server generates a summary based on the target image and the transcribed text.

[0072] In step S311, the server sends the minutes to the presentation terminal.

[0073] In step S312, the presentation terminal presents the minutes.

[0074] In this way, the server transcribes the audio information to obtain a transcribed text, and generates minutes based on the target image and the transcribed text. This can fully utilize the powerful computing power of the server to accurately and reliably analyze and process the target image and audio information, thereby reducing the operating burden of the smart glasses.

[0075] In some possible implementations, the target image has corresponding time information, and the transcribed text has corresponding time information. Both the smart glasses and the server can use the following method to generate the target image and transcribed text: Figure 4 The following method is used to generate the minutes based on the target image and the transcribed text: In step S401 , the target image and the transcribed text are temporally aligned according to the time information corresponding to the target image and the time information corresponding to the transcribed text.

[0076] In step S402, a summary is generated using the alignment result.

[0077] For example, when each target image is captured, a corresponding timestamp can be generated. When generating a transcribed text using ASR technology, the start and end timestamps of each text segment can be recorded. Based on the timestamp of the target image and the timestamp of the transcribed text, the target image can be time-aligned with the corresponding text segment. For example, if a target image is captured at the 10th minute, and the content from the 10th minute to the 11th minute in the transcribed text is "Next, let's look at this picture, which shows...", then the target image will be aligned with the text segment.

[0078] In this way, through time alignment, it can be ensured that the target image and the corresponding audio content are accurately matched, avoiding the misalignment problem between the target image and the transcribed text content, making the minutes content clearer and more accurate.

[0079] In some possible implementations, the Figure 5 Steps S4021 to S4023 shown implement step S402.

[0080] In step S4021, natural language processing is performed on the transcribed text to obtain the first key information.

[0081] For example, natural language processing (NLP) technology can be used to process the transcript. For example, NLP technology can be used to perform semantic analysis and extract key content such as decision-making information and task assignment information. Another example is that NLP technology can be used to separate speakers to distinguish the speech records of different speakers and improve readability. The first key information can be sorted according to the timeline corresponding to the transcript. The first key information can include at least one of decision-making information, task assignment information, summary information, and speech records determined based on the transcript.

[0082] In step S4022, character recognition is performed on the target image to obtain second key information.

[0083] For example, optical character recognition (OCR) technology can be used to process the target image to obtain text information in the target image, and the extracted text information can be determined as the second key information. The second key information and the corresponding target image can have the same timestamp.

[0084] In step S4023, the first key information and the second key information in the same time dimension are fused and optimized to generate a text minutes.

[0085] For example, the timestamp of the first key information and the timestamp of the second key information can be extracted. Aligning the first and second key information based on their timestamps ensures a precise match between the image and voice content. The first and second key information within the same time dimension can be integrated, and the integrated information can be optimized to remove duplicate and redundant content. NLP technology can then be used to further optimize the deduplicated content to make it more consistent with natural speech expression.

[0086] In this way, through information fusion optimization, the information in the target image is used to supplement the information in the audio information, which can reduce information omissions and misunderstandings, making the generated minutes more rich and complete. Through fusion optimization processing, the text minutes can be made more concise and fluent, with higher readability.

[0087] In some possible implementations, the written minutes include at least one of decision information, task assignment information, summary information, and speech records.

[0088] The decision information, task assignment information, summary information, and speech records in the text minutes are obtained based on the transcribed text and the target image.

[0089] In one embodiment, the speech record contains the identity information of the corresponding speaker, and the identity information is determined by at least one of the following methods: Use voiceprint recognition technology to analyze the speaker's voiceprint characteristics and determine the speaker's identity information; Using face recognition technology, the first state image of the speaker when speaking is analyzed to determine the speaker's identity information.

[0090] For example, a voiceprint sample of each participant can be collected in advance, and the voiceprint features in the audio information can be analyzed using voiceprint recognition technology. The determined voiceprint features and voiceprint samples can be compared to determine the identity information of the speaker corresponding to the speech record.

[0091] For another example, facial image samples of each participant can be collected in advance. During the audio information collection process, a first-state image of the speaker can be collected. This first-state image may include an image of the speaker's face. Facial recognition technology can be used to analyze the facial features in the collected first-state image and compare it with the facial image samples to determine the speaker's identity information corresponding to the speech record.

[0092] As another example, the results of voiceprint recognition and facial recognition can be combined to improve the accuracy and reliability of identity recognition. For example, if the voiceprint recognition and facial recognition results match, it is confirmed that they are the same speaker and the confirmed identity information is output; if they do not match, a second verification can be performed.

[0093] In this way, voiceprint recognition and / or facial recognition technology can be used to accurately determine the speaker's identity. The presence of the corresponding speaker's identity information in the speech record facilitates subsequent review and tracing of key speech content, especially in multi-person meetings, allowing the speech of a specific speaker to be quickly located.

[0094] In one embodiment, the speech record also contains the corresponding speaker's emotional information, and the emotional information is determined by at least one of the following methods: Utilize audio information to analyze the speaker's emotional characteristics and determine the speaker's emotional information; The second state image of the speaker when speaking is analyzed for expression and action to determine the speaker's emotional information.

[0095] For example, emotional features such as intonation, speaking speed, and volume can be extracted from audio information to determine the speaker's emotional information (e.g., calm, excited, anxious, etc.). Emotional features refer to acoustic characteristics related to the speaker's emotional state.

[0096] As another example, during the audio information collection process, a second-state image of the speaker can be captured. This second-state image may include a full-body image of the speaker. Image analysis techniques can be used to analyze facial expressions (such as smiles, frowns, and glares) and body movements (such as gestures and body posture) in the second-state image. Subsequently, by combining expression recognition algorithms and motion analysis models, the speaker's emotional information can be determined.

[0097] As another example, the results of speech analysis and image analysis can be combined to improve the accuracy and reliability of emotion recognition. For example, if speech analysis shows that the speaker's tone is high and the speech rate is fast, and image analysis shows that the speaker's facial expression is tense and the gestures are frequent, the speaker's emotional state can be comprehensively judged as excited.

[0098] In this way, using speech analysis and / or image analysis, we can accurately identify the speaker's emotional state. By including the speaker's emotional information in the speech record, the minutes can more comprehensively and vividly reflect the meeting's true atmosphere and the speaker's attitude, providing valuable reference for meeting summary and improvement.

[0099] In some possible implementations, such as Figure 6 As shown, when generating a text minutes, steps S4024 and S4025 can be used to generate a multimodal minutes.

[0100] In step S4024, the target location is determined in the text minutes.

[0101] In one embodiment, the target location can be determined in the text minutes by: According to the time information corresponding to the target image, the target position is determined in the text minutes.

[0102] For example, the timestamp of the target image can be matched to the timeline of the text minutes to determine the target's location within the text minutes. This time alignment ensures a precise match between the target image and the text minutes, avoiding information misalignment and improving the traceability and accuracy of the minutes.

[0103] In yet another embodiment, the target location may be determined by: The position in the text minutes that has the highest similarity to the second key information is determined as the target position.

[0104] For example, the semantic similarity between the text minutes and the second key information can be determined, and the location in the text minutes with the highest semantic similarity to the second key information can be determined as the target location. In this way, by calculating the similarity, the target image is inserted into the most relevant part of the text minutes, ensuring a precise match between the target image and the text content, better associating it with the context, and allowing users to quickly reference highly relevant images while reading the text, thereby better understanding the content of the minutes.

[0105] In step S4025, the target image is inserted into the corresponding target position to obtain a multimodal minutes.

[0106] In this way, in the above-mentioned technical solution for generating multimodal minutes, the two forms of images and text are combined, which can more intuitively display key information, enhance the expressiveness and comprehensibility of the minutes, and avoid information omission.

[0107] In some possible implementations, the minutes can be exported visually, for example, to PDF / Word format; or further to Markdown format, suitable for team collaboration tools (such as Notion and Confluence).

[0108] Figure 7 1 is a block diagram of a minutes generating device 500 according to an exemplary embodiment, wherein the minutes generating device 500 is applied to smart glasses. Figure 7 , the minutes generating device 500 may include: An audio collection module 501 is configured to collect audio information using the audio collection unit of the smart glasses; An image acquisition module 502 is configured to acquire a target image using an image acquisition unit of the smart glasses in response to determining that a preset condition is satisfied during the process of acquiring the audio information; The first processing module 503 is configured to obtain a transcribed text according to the audio information, and generate a summary according to the target image and the transcribed text.

[0109] In the above technical solution, smart glasses serve as a wearable device, making them convenient for users to use in various scenarios. Users no longer need to carry multiple recording devices. While attending meetings or events, they can use smart glasses to automatically capture audio information and target images, thus improving the convenience and efficiency of recording. A transcript is then generated based on the audio information, and minutes are generated based on the target image and transcript. This combination of sound and vision allows for the generation of minutes that more comprehensively and accurately reflect the actual situation on site, avoiding information gaps or biases that may arise from a single information source. This also effectively improves the efficiency of minute generation, eliminating the need for users to spend significant time manually organizing minutes after a meeting or event, allowing for timely access to the minutes.

[0110] In some possible implementations, the preset conditions include: The current screen includes a preset information presentation interface, and the information presentation interface includes text and / or graphics.

[0111] In some possible implementations, the area where the image acquisition unit captures images is determined according to the gaze point of the target user.

[0112] In some possible implementations, the audio collection module 501 is configured to collect audio information using the audio collection unit of the smart glasses when any of the following conditions is met: receiving a voice command instructing content recording; It is determined that the target user performs a target operation on a preset area of the smart glasses.

[0113] In some possible implementations, the first processing module 503 is also used to transcribe the audio information using the processing unit of the smart glasses to obtain a transcribed text, and generate minutes based on the target image and the transcribed text, and then use the smart glasses to display the minutes, or use the communication unit of the smart glasses to send the minutes to a terminal bound to the smart glasses, so that the minutes can be displayed using the terminal.

[0114] In some possible implementations, the first processing module 503 is further used to use the communication unit of the smart glasses to send the audio information and the target image to a server, so that the server obtains a transcribed text based on the audio information, and generates and issues the minutes based on the target image and the transcribed text, so that the device that receives the minutes displays the minutes, and the device that receives the minutes includes the smart glasses, and / or a terminal bound to the smart glasses.

[0115] In some possible implementations, the target image has corresponding time information, and the transcribed text has corresponding time information; the first processing module 503 is configured to generate a minutes based on the target image and the transcribed text in the following manner: Temporally aligning the target image and the transcribed text according to time information corresponding to the target image and time information corresponding to the transcribed text; The minutes are generated using the alignment results.

[0116] In some possible implementations, the first processing module 503 is configured to generate the minutes using the alignment result in the following manner: Performing natural language processing on the transcribed text to obtain first key information; Performing character recognition on the target image to obtain second key information; The first key information and the second key information in the same time dimension are fused and optimized to generate a text minutes, which includes at least one of decision information, task assignment information, summary information, and speech records.

[0117] In some possible implementations, the first processing module 503 is further configured to determine a target position in the text minutes; and insert the target image into the corresponding target position to obtain a multimodal minutes.

[0118] In some possible implementations, the first processing module 503 is configured to determine the target location in the text minutes by: Determine the target position in the text minutes according to the time information corresponding to the target image; or The position in the text minutes that has the highest similarity to the second key information is determined as the target position.

[0119] In some possible implementations, the speech record contains identity information of the corresponding speaker, and the identity information is determined by at least one of the following methods: Analyzing the speaker's voiceprint characteristics using voiceprint recognition technology to determine the identity information; The face recognition technology is used to analyze the first state image of the speaker when speaking to determine the identity information.

[0120] In some possible implementations, the speech record further includes emotional information of the corresponding speaker, and the emotional information is determined by at least one of the following methods: Analyzing the emotional characteristics of the speaker using the audio information to determine the emotional information; The expression and action of the speaker in the second state image when speaking are analyzed to determine the emotional information.

[0121] Figure 8 1 is a block diagram of a minutes generating device 600 according to an exemplary embodiment, wherein the minutes generating device 600 is applied to a server. Figure 8 , the minutes generating device 600 may include: A receiving module 601 receives audio information and a target image uploaded by the smart glasses, wherein the target image is an image captured by the image capture unit of the smart glasses when the audio capture unit of the smart glasses is collecting the audio information and a preset condition is met; A second processing module 602 is configured to obtain a transcribed text based on the audio information, and generate a summary based on the target image and the transcribed text; The sending module 603 is used to send the minutes so that the device that receives the minutes can display the minutes.

[0122] In the above technical solution, minutes are generated using a target image captured when preset conditions are met, and a transcribed text obtained by transcribing audio information. This not only improves the efficiency of minute generation, but also enables efficient integration of visual and auditory information, more comprehensively and accurately capturing key information, reducing omissions and misunderstandings, and making the generated minutes richer and more complete. A device in the server transcribes the audio information to generate a transcribed text, and generates minutes based on the target image and transcribed text. This fully utilizes the server's powerful computing power to accurately and reliably analyze and process the target image and audio information, reducing the operational burden on the smart glasses.

[0123] In some possible implementations, the preset conditions include: The current screen includes a preset information presentation interface, and the information presentation interface includes text and / or graphics.

[0124] In some possible implementations, the area where the image acquisition unit captures images is determined according to the gaze point of the target user.

[0125] In some possible implementations, the target image has corresponding time information, and the transcribed text has corresponding time information; the second processing module 602 is configured to generate a minutes based on the target image and the transcribed text in the following manner: Temporally aligning the target image and the transcribed text according to time information corresponding to the target image and time information corresponding to the transcribed text; The minutes are generated using the alignment results.

[0126] In some possible implementations, the second processing module 602 is configured to generate the minutes using the alignment result in the following manner: Performing natural language processing on the transcribed text to obtain first key information; Performing character recognition on the target image to obtain second key information; The first key information and the second key information in the same time dimension are fused and optimized to generate a text minutes, which includes at least one of decision information, task assignment information, summary information, and speech records.

[0127] In some possible implementations, the second processing module 602 is further configured to determine a target position in the text minutes; and insert the target image into the corresponding target position to obtain a multimodal minutes.

[0128] In some possible implementations, the second processing module 602 is configured to determine the target location in the text minutes by: Determine the target position in the text minutes according to the time information corresponding to the target image; or The position in the text minutes that has the highest similarity to the second key information is determined as the target position.

[0129] In some possible implementations, the speech record contains identity information of the corresponding speaker, and the identity information is determined by at least one of the following methods: Analyzing the speaker's voiceprint characteristics using voiceprint recognition technology to determine the identity information; The face recognition technology is used to analyze the first state image of the speaker when speaking to determine the identity information.

[0130] In some possible implementations, the speech record further includes emotional information of the corresponding speaker, and the emotional information is determined by at least one of the following methods: Analyzing the emotional characteristics of the speaker using the audio information to determine the emotional information; The expression and action of the speaker in the second state image when speaking are analyzed to determine the emotional information.

[0131] In some possible implementations, the device that receives the minutes includes the smart glasses and / or a terminal bound to the smart glasses.

[0132] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0133] Figure 9 FIG. 1 is a block diagram of a smart glasses according to an exemplary embodiment. Figure 9 , the smart glasses 800 may include one or more of the following components: a first processing component 802 , a first memory 804 , a first power supply component 806 , a multimedia component 808 , an audio component 810 , a first input / output interface 812 , a sensor component 814 , and a communication component 816 .

[0134] The first processing component 802 generally controls the overall operation of the smart glasses 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The first processing component 802 may include one or more first processors 820 to execute instructions to perform all or part of the steps of the above-described minutes generation method. In addition, the first processing component 802 may include one or more modules to facilitate interaction between the first processing component 802 and other components. For example, the first processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the first processing component 802.

[0135] The first memory 804 is configured to store various types of data to support operations on the smart glasses 800. Examples of such data include instructions for any application or method operating on the smart glasses 800, contact data, phone book data, messages, pictures, videos, etc. The first memory 804 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0136] The first power supply component 806 provides power to the various components of the smart glasses 800. The first power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the smart glasses 800.

[0137] The multimedia component 808 includes a screen that provides an output interface between the smart glasses 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the smart glasses 800 are in an operating mode, such as a capture mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and the rear-facing camera may include a fixed optical lens system or have focal length and optical zoom capabilities.

[0138] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the smart glasses 800 are in an operating mode, such as a call mode, a recording mode, and a speech recognition mode. The received audio signals may be further stored in the first memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0139] The first input / output interface 812 provides an interface between the first processing component 802 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0140] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the smart glasses 800. For example, the sensor assembly 814 can detect the open / closed state of the smart glasses 800, the relative positioning of components, such as the display and keypad of the smart glasses 800. The sensor assembly 814 can also detect changes in the position of the smart glasses 800 or a component of the smart glasses 800, the presence or absence of user contact with the smart glasses 800, the orientation or acceleration / deceleration of the smart glasses 800, and changes in the temperature of the smart glasses 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0141] The communication component 816 is configured to facilitate wired or wireless communication between the smart glasses 800 and other devices. The smart glasses 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0142] In an exemplary embodiment, the smart glasses 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned minutes generation method.

[0143] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is further provided, such as a first memory 804 including instructions. The instructions may be executed by the first processor 820 of the smart glasses 800 to perform the above-described minutes generation method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0144] Figure 10FIG1 is a block diagram of a server 1900 according to an exemplary embodiment. Figure 10 Server 1900 includes a second processing component 1922, which further includes one or more processors and a memory resource represented by a second memory 1932 for storing instructions executable by second processing component 1922, such as an application. The application stored in second memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, second processing component 1922 is configured to execute the instructions to perform the above-described minutes generation method.

[0145] The server 1900 may further include a second power supply component 1926 configured to perform power management of the server 1900, a wired or wireless network interface 1950 configured to connect the server 1900 to the network, and a second input / output interface 1958. The server 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2000. TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or similar.

[0146] The present disclosure also provides a minutes generation system, which includes the smart glasses 800 provided in the above embodiment.

[0147] In some possible implementations, such as Figure 11 As shown, the minutes generation system also includes the server 1900 provided in the above embodiment.

[0148] In some possible implementations, the minutes generation system further includes a terminal bound to the smart glasses, and the terminal is used to receive and display minutes.

[0149] In another exemplary embodiment, the present disclosure further provides a computer program product, which includes a computer program that can be executed by a programmable device, and has a code portion for executing the above-mentioned minutes generation method when executed by the programmable device.

[0150] In another exemplary embodiment, the present disclosure further provides a computer-readable storage medium having computer program instructions stored thereon, which implement the steps of the minutes generation method provided by the present disclosure when the program instructions are executed by a processor.

[0151] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented through electronic hardware, computer software, or a combination of both. Whether such functions are implemented through hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0152] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.

[0153] It should be understood that, unless otherwise expressly specified or limited, the terms "join," "attach," "install," "connect," "connect," "fix," etc. used in the embodiments of the present disclosure should be understood in a broad sense. For example, they can be fixedly connected, detachably connected, or integrated; they can be mechanically connected, electrically connected, or communicable with each other; they can be directly connected, or indirectly connected through an intermediate medium, and they can be internally connected between two elements or an interactive relationship between two elements, unless otherwise expressly limited. For those skilled in the art, the specific meanings of the above terms in this article can be understood according to specific circumstances.

[0154] Although terms such as "first", "second" and "third" may be used herein to describe various components, parts, regions, layers or sections, these components, parts, regions, layers or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer or section from another component, part, region, layer or section. Therefore, without departing from the teachings of each example, the first component, part, region, layer or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer or section. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" can explicitly or implicitly include at least one such feature. In the description herein, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise clearly and specifically defined.

[0155] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then "X applies to A or B" satisfies any of the aforementioned instances. Furthermore, the articles "a" and "an," as used in this application and the appended claims, are generally understood to mean "one or more," unless otherwise specified or clear from the context to refer to the singular form.

[0156] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. With particular regard to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. In addition, although particular features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and advantageous for any given or particular application. Furthermore, to the extent that the terms "include," "have," "have," "have," or variations thereof are used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0157] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

[0158] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating minutes, characterized in that: Applied to smart glasses, including: Collecting audio information using the audio collection unit of the smart glasses; In the process of collecting the audio information, in response to determining that a preset condition is satisfied, collecting a target image using an image collection unit of the smart glasses; A transcribed text is obtained according to the audio information, and a summary is generated according to the target image and the transcribed text.

2. The method for generating minutes according to claim 1, wherein: The preset conditions include: The current screen includes a preset information presentation interface, and the information presentation interface includes text and / or graphics.

3. The method for generating minutes according to claim 1, wherein: The area where the image acquisition unit acquires images is determined according to the gaze point of the target user.

4. The method for generating minutes according to claim 1, wherein: When any one of the following conditions is met, audio information is collected by using the audio collection unit of the smart glasses: receiving a voice command instructing content recording; It is determined that the target user performs a target operation on a preset area of the smart glasses.

5. The method for generating minutes according to claim 1, wherein: The method further comprises: When the audio information is transcribed using the processing unit of the smart glasses to obtain a transcribed text, and minutes are generated based on the target image and the transcribed text, the minutes are displayed using the smart glasses, or the minutes are sent to a terminal bound to the smart glasses using the communication unit of the smart glasses, so that the minutes can be displayed using the terminal.

6. The method for generating minutes according to claim 1, wherein: The step of obtaining a transcribed text based on the audio information and generating a summary based on the target image and the transcribed text includes: The communication unit of the smart glasses is used to send the audio information and the target image to a server, so that the server obtains a transcribed text based on the audio information, and generates and distributes the minutes based on the target image and the transcribed text, so that the device that receives the minutes can display the minutes. The device that receives the minutes includes the smart glasses and / or a terminal bound to the smart glasses.

7. The method for generating minutes according to claim 5 or 6, characterized in that: The target image has corresponding time information, and the transcribed text has corresponding time information; generating a minutes based on the target image and the transcribed text includes: Temporally aligning the target image and the transcribed text according to time information corresponding to the target image and time information corresponding to the transcribed text; The minutes are generated using the alignment results.

8. The method for generating minutes according to claim 7, wherein: The generating of the minutes by utilizing the alignment results includes: Performing natural language processing on the transcribed text to obtain first key information; Performing character recognition on the target image to obtain second key information; The first key information and the second key information in the same time dimension are fused and optimized to generate a text minutes, which includes at least one of decision information, task assignment information, summary information, and speech records.

9. The method for generating minutes according to claim 8, wherein: The method further comprises: identifying a target location in the written minutes; The target image is inserted into the corresponding target position to obtain a multimodal summary.

10. The method for generating minutes according to claim 9, characterized in that: Determining the target location in the text minutes includes: Determine the target position in the text minutes according to the time information corresponding to the target image; or The position in the text minutes that has the highest similarity to the second key information is determined as the target position.

11. The method for generating minutes according to claim 8, wherein: The speech record contains identity information of the corresponding speaker, and the identity information is determined by at least one of the following methods: Analyzing the speaker's voiceprint characteristics using voiceprint recognition technology to determine the identity information; The face recognition technology is used to analyze the first state image of the speaker when speaking to determine the identity information.

12. The method for generating minutes according to claim 8, wherein: The speech record also contains the corresponding speaker's emotional information, and the emotional information is determined by at least one of the following methods: Analyzing the speaker's emotional characteristics using the audio information to determine the emotional information; The expression and action of the speaker in the second state image when speaking are analyzed to determine the emotional information.

13. A method for generating minutes, characterized in that: Applicable to servers, including: receiving audio information and a target image uploaded by the smart glasses, wherein the target image is an image captured by the image capture unit of the smart glasses when the audio capture unit of the smart glasses is collecting the audio information and a preset condition is met; Obtaining a transcribed text based on the audio information, and generating a summary based on the target image and the transcribed text; The minutes are issued so that a device receiving the minutes can display the minutes.

14. The method for generating minutes according to claim 13, wherein: The preset conditions include: The current screen includes a preset information presentation interface, and the information presentation interface includes text and / or graphics.

15. The method for generating minutes according to claim 13, wherein: The area where the image acquisition unit acquires images is determined according to the gaze point of the target user.

16. The method for generating minutes according to claim 13, wherein: The target image has corresponding time information, and the transcribed text has corresponding time information; generating a minutes based on the target image and the transcribed text includes: Temporally aligning the target image and the transcribed text according to time information corresponding to the target image and time information corresponding to the transcribed text; The minutes are generated using the alignment results.

17. The method for generating minutes according to claim 16, wherein: The generating of the minutes by utilizing the alignment results includes: Performing natural language processing on the transcribed text to obtain first key information; Performing character recognition on the target image to obtain second key information; The first key information and the second key information in the same time dimension are fused and optimized to generate a text minutes, which includes at least one of decision information, task assignment information, summary information, and speech records.

18. The method for generating minutes according to claim 17, wherein: The method further comprises: identifying a target location in the written minutes; The target image is inserted into the corresponding target position to obtain a multimodal summary.

19. The method for generating minutes according to claim 18, wherein: Determining the target location in the text minutes includes: Determine the target position in the text minutes according to the time information corresponding to the target image; or The position in the text minutes that has the highest similarity to the second key information is determined as the target position.

20. The method for generating minutes according to any one of claims 13 to 19, characterized in that: The device that receives the minutes includes the smart glasses and / or a terminal bound to the smart glasses.

21. A minutes generating device, characterized in that: Applied to smart glasses, used to implement the minutes generation method according to any one of claims 1 to 12.

22. A minutes generating device, characterized in that: Applied to a server, used to implement the minutes generation method described in any one of claims 13-20.

23. A pair of smart glasses, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the minutes generation method according to any one of claims 1 to 12.

24. A server, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the minutes generation method described in any one of claims 13-20.

25. A minutes generation system, characterized in that: Including the smart glasses as described in claim 23.

26. The minutes generating system according to claim 25, characterized in that: The minutes generating system further comprises the server as claimed in claim 24.

27. The minutes generating system according to claim 25 or 26, characterized in that: The minutes generation system further includes a terminal bound to the smart glasses, and the terminal is used to receive and display the minutes.

28. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the steps of the minutes generation method according to any one of claims 1 to 12, or the steps of the minutes generation method according to any one of claims 13 to 20.

29. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the minutes generation method according to any one of claims 1 to 12, or implements the steps of the minutes generation method according to any one of claims 13 to 20.

Citation Information

Cited By

  • Multi-mode intelligent semantic understanding and abstract generation system and method based on HDMI (High Definition Multimedia Interface) stream

    CN121029981A