Intelligent conference auxiliary system and method for generating conference record

By introducing image interception and analysis functions into the intelligent conference auxiliary system, identifying the image content displayed in the conference, the problem that the existing system cannot record visual information is solved, and the generated conference records are more detailed and comprehensive.

CN120201152APending Publication Date: 2025-06-24MEGAFORCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410306076.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-08
Filing Date
2024-03-18
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing intelligent conference assistance systems are unable to effectively analyze and record the image content, especially text and charts, in the meeting, resulting in a lack of visual information on the meeting records.

Method used

An intelligent conference assistance system is designed, including an image intercepting device and an image analysis device. The image interceptor device intercepts the images displayed during the meeting. The image analysis device recognizes the text and chart content in the image through the image analysis program to generate a meeting record that records the image content.

Benefits of technology

Effective analysis and recording of image content in the meeting is realized. The generated conference records not only contain voice content, but also text and chart content on the image, thereby more comprehensively reflecting the conference content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201152A_ABST
    Figure CN120201152A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent conference auxiliary system and a method for generating a conference record. The intelligent conference auxiliary system comprises an image interception device and an image analysis device. The image capture device is configured to capture an image displayed by the interactive device during a conference. The image analysis device is coupled to the image capture device and configured to perform an image analysis program on the image to generate a first conference record recording image content, whereby non-audio image content is not ignored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an auxiliary system, and particularly to an intelligent conference auxiliary system and a method for generating a conference record. Background Art

[0002] Basically, an intelligent conference auxiliary system is a meeting assistant that can automatically generate a conference record. However, existing intelligent conference auxiliary systems only generate a conference record that records multiple speech contents based on voice data. However, humans are visual learners.

[0003] Therefore, a conference presenter can use an image containing text and / or charts to explain information to conference participants. That is, existing intelligent conference auxiliary systems ignore non-audio image content, resulting in it being unable to be reflected in the conference record. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an intelligent conference auxiliary system and a method for generating a conference record in view of the deficiencies of the prior art, which can analyze an image containing text and / or charts to generate a conference record that records the image content.

[0005] To solve the above technical problem, one of the technical solutions adopted by the present invention is to provide an intelligent conference auxiliary system, including an image capture device and an image analysis device. The image capture device is configured to capture an image displayed by an interactive device during a conference. The image analysis device is coupled to the image capture device and is configured to execute an image analysis program on the image to generate a first conference record that records the image content.

[0006] Preferably, the image content includes one or a combination of text content and chart content on the image, and the image analysis program includes a text recognition program and a chart recognition program.

[0007] Preferably, the image analysis device includes an optical character recognition circuit and a chart recognition circuit. The optical character recognition circuit is configured to execute a text recognition program on the image to generate a first event record that records the text content. The chart recognition circuit is configured to execute a chart recognition program on the image to generate a second event record that records the chart content.

[0008] Preferably, the text content includes one or a combination of printed text content and handwritten text content on the image, and the optical character recognition circuit executes a text recognition program on the image to recognize the printed text content and the handwritten text content on the image.

[0009] Preferably, the chart content includes one or a combination of pie chart content, line chart content, and bar chart content on the image, and the chart recognition circuit executes a chart recognition program on the image to recognize the pie chart content, line chart content, and bar chart content on the image.

[0010] Preferably, the image analysis program further includes a dynamic display recognition program, and the image content includes one or a combination of text content, chart content, and dynamic display content on the image.

[0011] Preferably, the image analysis device further includes a dynamic display recognition circuit. The dynamic display recognition circuit is configured to execute a dynamic display recognition program on the image to generate a third event record recording the dynamic display content. The image analysis device is further configured to generate a first meeting record recording the image content according to the first event record, the second event record, and the third event record.

[0012] Preferably, the intelligent meeting assistance system further includes a sound input device and a sound analysis device. The sound input device is configured to generate sound data during the meeting. The sound analysis device is coupled to the sound input device and is configured to execute a sound analysis program on the sound data to generate a second meeting record recording a plurality of speech contents.

[0013] Preferably, the image capture device is further configured to sequentially capture a plurality of images displayed by the interactive device during the meeting, and the image analysis device generates a first meeting record respectively recording the plurality of image contents.

[0014] Preferably, the image analysis device is further configured to add a timestamp to each image content recorded in the first meeting record, and the sound analysis device is further configured to add a timestamp to each speech content recorded in the second meeting record.

[0015] Preferably, the intelligent meeting assistance system further includes an artificial intelligence processing circuit. The artificial intelligence processing circuit is coupled to the image analysis device and the sound analysis device, and receives the first meeting record and the second meeting record. The artificial intelligence processing circuit is configured to integrate the first meeting record and the second meeting record, and input the integrated meeting record into a natural language processing and machine learning model for analysis and generate a third meeting record.

[0016] To solve the above technical problems, another technical solution adopted by the present invention is to provide a method for generating a meeting record, including the following steps: configuring an image capture device to capture an image displayed by an interactive device during a meeting; and configuring an image analysis device to execute an image analysis program on the image to generate a first meeting record recording the image content.

[0017] Preferably, the image content includes one or a combination of text content and chart content on the image, and the image analysis program includes a text recognition program and a chart recognition program.

[0018] Preferably, the steps of configuring the image analysis device to execute the image analysis program on the image include: configuring an optical character recognition circuit to execute the text recognition program on the image to generate a first event record recording the text content; and configuring a chart recognition circuit to execute the chart recognition program on the image to generate a second event record recording the chart content.

[0019] Preferably, the image analysis program further includes a dynamic display recognition program, and the image content includes one or a combination of text content, chart content, and dynamic display content on the image.

[0020] Preferably, the steps of configuring the image analysis device to execute the image analysis program on the image further include: configuring a dynamic display recognition circuit to execute the dynamic display recognition program on the image to generate a third event record recording the dynamic display content; and configuring the image analysis device to generate a first meeting record recording the image content according to the first event record, the second event record, and the third event record.

[0021] Preferably, the method further includes the following steps: configuring a sound input device to generate sound data during the meeting; and configuring a sound analysis device to execute a sound analysis program on the sound data to generate a second meeting record recording multiple speech contents.

[0022] Preferably, the image capture device is further configured to sequentially capture multiple images displayed by the interactive device during the meeting, and the image analysis device generates a first meeting record respectively recording the multiple image contents.

[0023] Preferably, the image analysis device is further configured to add a timestamp to each image content recorded in the first meeting record, and the sound analysis device is further configured to add a timestamp to each speech content recorded in the second meeting record.

[0024] Preferably, the method further includes the following steps: configuring an artificial intelligence processing circuit to integrate the first meeting record and the second meeting record, and inputting the integrated meeting record into a natural language processing and machine learning model for analysis and generating a third meeting record.

[0025] To enable a further understanding of the features and technical content of the present invention, please refer to the following detailed description of the present invention and the accompanying drawings. However, the provided drawings are only for reference and illustration, and are not used to limit the present invention. Description of the Drawings

[0026] Figure 1It is a functional block diagram of the intelligent conference assistance system according to an embodiment of the present invention.

[0027] Figure 2 It is a step flowchart of the method for generating a meeting record according to an embodiment of the present invention.

[0028] Figure 3 It is a step flowchart of the method for configuring an image analysis device to execute an image analysis program on an image according to an embodiment of the present invention.

[0029] Figure 4A It is a functional block diagram of the optical character recognition circuit according to an embodiment of the present invention.

[0030] Figure 4B It is a functional block diagram of the chart recognition circuit according to an embodiment of the present invention.

[0031] Figure 4C It is a functional block diagram of the dynamic display recognition circuit according to an embodiment of the present invention.

[0032] Figure 5 It is a functional block diagram of the speech recognition circuit according to an embodiment of the present invention.

[0033] Figure 6A It is a schematic diagram of the intelligent conference assistance system according to an embodiment of the present invention for displaying speech content in a copy mode.

[0034] Figure 6B It is a schematic diagram of the intelligent conference assistance system according to an embodiment of the present invention for displaying speech content in a summary mode.

[0035] Figure 6C It is a schematic diagram of the intelligent conference assistance system according to an embodiment of the present invention for displaying a meeting summary in an abstract mode. Detailed implementation manners

[0036] The following are specific embodiments to illustrate the implementation manners of the present invention regarding the "intelligent conference assistance system and method for generating a meeting record". Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the concept of the present invention. Additionally, the drawings of the present invention are only for simple schematic illustration and are not drawn according to actual dimensions, hereby declared in advance. The following implementation manners will further elaborate on the related technical content of the present invention, but the disclosed content is not intended to limit the protection scope of the present invention. In addition, the term "or" used herein should, depending on the actual situation, possibly include any one or a combination of more of the related listed items.

[0037] Please refer to Figure 1 and Figure 2 ,Figure 1 is a functional block diagram of the intelligent conference assistance system according to an embodiment of the present invention, Figure 2 and is a step flowchart of the method for generating a meeting record according to an embodiment of the present invention. As Figure 1 shown, the intelligent conference assistance system 1 of this embodiment includes an image capture device 11 and an image analysis device 12. Specifically, the image capture device 11 is configured to capture the image 4 displayed by the interactive device 2 during the meeting. In addition, the image analysis device 12 is coupled to the image capture device 11 and is configured to execute an image analysis program on the image 4 to generate a first meeting record M1 that records the image content ( Figure 1 not shown).

[0038] For example, the interactive device 2 may be a touch screen in the meeting environment 3 and is configured to display a slide containing text and / or charts during the meeting, so that the meeting presenter 5 can explain information to the meeting participants ( Figure 1 not shown). In this case, the image 4 captured by the image capture device 11 may be a slide containing text and / or charts, and the image analysis device 12 can execute an image analysis program on the slide containing text and / or charts to generate a first meeting record M1 that records the image content of the slide. However, the present invention is not limited to the above example.

[0039] As Figure 2 shown, according to the above content, the method for generating a meeting record of this embodiment is executed by the intelligent conference assistance system 1 and includes the following steps.

[0040] Step S11: Configure the image capture device to capture the image displayed by the interactive device during the meeting. Specifically, the image capture device 11 may be implemented by hardware (such as a lens and an imaging medium) in combination with software and / or firmware. However, the present invention does not limit the specific implementation manner of the image capture device 11.

[0041] Step S12: Configure the image analysis device to execute an image analysis program on the image to generate a first meeting record that records the image content. Similarly, the image analysis device 12 may be implemented by hardware (such as a central processing unit and a memory) in combination with software and / or firmware. However, the present invention also does not limit the specific implementation manner of the image analysis device 12.

[0042] Furthermore, the image content recorded in the first meeting record M1 includes one or a combination of the text content and the chart content on the image 4, and the image analysis program executed by the image analysis device 12 includes a text recognition program and a chart recognition program. Therefore, as Figure 1 shown, the image analysis device 12 may include an optical character recognition circuit 121 and a chart recognition circuit 122.

[0043] The optical character recognition circuit 121 is configured to execute a character recognition program on the image 4 to generate a first event record recording the character content. However, Figure 1 the first event record is not shown. Additionally, the chart recognition circuit 122 is configured to execute a chart recognition program on the image 4 to generate a second event record recording the chart content. Similarly, Figure 1 the second event record is not shown either. That is to say, please refer to Figure 3 , Figure 3 which is a flowchart of the steps for the image analysis device configured in an embodiment of the present invention to execute an image analysis program on an image.

[0044] As Figure 3 shown, according to the above content, step S12 of this embodiment may include the following steps.

[0045] Step S121: Configure the optical character recognition circuit to execute a character recognition program on the image to generate a first event record recording the character content. Specifically, the characters on the image 4 may include one or a combination of printed characters typeset using modern computer fonts and handwritten characters written by the meeting presenter 5 on the interactive device 2 (e.g., a touch screen). Therefore, the character content recorded in the first event record may include one or a combination of the printed character content and the handwritten character content on the image 4, and the optical character recognition circuit 121 executes a character recognition program on the image 4 to recognize the printed character content and the handwritten character content on the image 4, thereby generating a first event record recording the character content.

[0046] Step S122: Configure the chart recognition circuit to execute a chart recognition program on the image to generate a second event record recording the chart content. Specifically, the charts on the image 4 may include one or a combination of a pie chart, a line chart, and a bar chart. Therefore, the chart content recorded in the second event record may include one or a combination of the pie chart content, the line chart content, and the bar chart content on the image 4, and the chart recognition circuit 122 executes a chart recognition program on the image 4 to recognize the pie chart content, the line chart content, and the bar chart content on the image 4, thereby generating a second event record recording the chart content.

[0047] Similarly, the pie chart, line chart, and / or bar chart on Image 4 can be created not only using a charting application on a modern computer but also drawn by the meeting presenter 5 on the interactive device 2 (e.g., touch screen). That is to say, according to the above, the intelligent meeting assistance system 1 and the method for generating meeting records in this embodiment can also intercept and record the content written and drawn by the meeting presenter 5 on the interactive device 2. Therefore, the generated meeting records can more accurately reflect all the behaviors in the meeting.

[0048] On the other hand, the meeting presenter 5 can also interact with the meeting participants using the animation function on Image 4 (e.g., slides). The animation function can, for example, arrange various trajectories and / or various transition effects for text content and chart content to perform actions such as moving, rotating, appearing, and / or disappearing. Therefore, the image analysis program executed by the image analysis device 12 can also include an animation recognition program for identifying the above actions, and the image content recorded in the first meeting record M1 can include one or a combination of the text content, chart content, and animation content on Image 4. Thus, as Figure 1 shown, the image analysis device 12 can also include an animation recognition circuit 123.

[0049] The animation recognition circuit 123 is configured to execute an animation recognition program on Image 4 to generate a third event record recording the animation content. In some embodiments, the above animation-related content is usually the key points of the meeting that the meeting presenter 5 wants to emphasize. Therefore, in the method for generating meeting records provided by the present invention, the identified animation content is regarded as a key part of the meeting and presented in the third event record (e.g., by assigning weights). However, Figure 1 the third event record is not shown either. In addition, the image analysis device 12 can also be configured to generate a first meeting record M1 recording the image content according to the first event record, the second event record, and the third event record. That is to say, as Figure 3 shown, step S12 of this embodiment can also include the following steps.

[0050] Step S123: Configure the animation recognition circuit to execute an animation recognition program on the image to generate a third event record recording the animation content.

[0051] Step S124: Configure the image analysis device to generate a first meeting record recording the image content according to the first event record, the second event record, and the third event record.

[0052] It should be noted that the present invention does not limit the specific form of the image content recorded in the first meeting record M1. For example, the first meeting record M1 of this embodiment can record the text content, chart content, and dynamic display content on image 4 in text form. In addition, the first meeting record M1 of this embodiment can also record the chart content on image 4 in the form of re-creating a chart. However, the present invention is not limited to the above examples.

[0053] On the other hand, as Figure 1 shown, the intelligent conference assistance system 1 of this embodiment may further include a sound input device 13 and a sound analysis device 14. The sound input device 13 is configured to generate sound data during the meeting, but Figure 1 the sound data is not shown. In addition, the sound analysis device 14 is coupled to the sound input device 13 and is configured to execute a sound analysis program on the sound data to generate a second meeting record M2 recording a plurality of speech contents. That is, as Figure 2 shown, the method of this embodiment may further include the following steps.

[0054] Step S13: Configure the sound input device to generate sound data during the meeting. Specifically, the sound input device 13 may be configured to execute a recording program during the meeting to generate sound data. In addition, the sound input device 13 may be implemented by hardware (for example, a microphone) in combination with software and / or firmware. However, the present invention does not limit the specific implementation manner of the sound input device 13.

[0055] Step S14: Configure the sound analysis device to execute a sound analysis program on the sound data to generate a second meeting record recording a plurality of speech contents. Specifically, the sound analysis program executed by the sound analysis device 14 may include a speech recognition program and a voice signature identification program. Therefore, the sound analysis device 14 may include a speech recognition circuit 141 and a voice signature identification circuit 142 to respectively execute the speech recognition program and the voice signature identification program, so that the sound analysis device 14 generates a second meeting record M2 recording a plurality of speech contents.

[0056] Similarly, both the speech recognition circuit 141 and the voice signature identification circuit 142 may be implemented by hardware (for example, a central processing unit and a memory) in combination with software and / or firmware. However, the present invention does not limit the specific implementation manners of the speech recognition circuit 141 and the voice signature identification circuit 142. Since the operating principles of the speech recognition circuit 141 and the voice signature identification circuit 142 are well known to those skilled in the art, details of the sound analysis device 14 will not be elaborated further.

[0057] Further, the meeting presenter 5 may use more than one image 4 (e.g., multiple slides) during the meeting to explain information to the meeting participants. Therefore, the image capturing device 11 of this embodiment may also be configured to sequentially capture a plurality of images 4 displayed by the interactive device 2 during the meeting, and the image analysis device 12 generates a first meeting record M1 that separately records the contents of the plurality of images. It should be understood that the contents of the plurality of images recorded in the first meeting record M1 respectively correspond to the plurality of images 4 captured by the image capturing device 11.

[0058] In this embodiment, the image capturing device 11 may be configured to sequentially capture a plurality of images 4 according to a sampling frequency, but the present invention is not limited thereto. In other embodiments, the image capturing device 11 may also be configured to capture a new image 4 when the image 4 is updated (e.g., the meeting presenter 5 changes to another slide to explain information to the meeting participants, or the meeting presenter 5 writes and / or draws a chart on the interactive device 2), but the present invention is not limited thereto either.

[0059] It should also be understood that each image content may be associated with at least one voice content recorded in the second meeting record M2. Therefore, in order to establish the association and order between the plurality of image contents and the plurality of voice contents, the image analysis device 12 may also be configured to add a timestamp to each image content recorded in the first meeting record M1, and the sound analysis device 14 may also be configured to add a timestamp to each voice content recorded in the second meeting record M2.

[0060] According to the above, since each image content recorded in the first meeting record M1 may include one or a combination of text content, chart content, and dynamic display content, adding a timestamp to each image content may also refer to adding a timestamp to the text content, chart content, and / or dynamic display content of each image content. Therefore, the intelligent meeting assistance system 1 of this embodiment may also include a clock circuit 15.

[0061] The clock circuit 15 is coupled to the image analysis device 12 and the sound analysis device 14, and is used to generate timestamps. That is, the first meeting record M1 and the second meeting record M2 generated by the image analysis device 12 and the sound analysis device 14 can both use the timestamps generated by the clock circuit 15 to record the appearance time of each image content and each voice content.

[0062] Furthermore, the intelligent conference assistance system 1 of this embodiment may further include an artificial intelligence processing circuit 16. The artificial intelligence processing circuit 16 is coupled to the image analysis device 12 and the sound analysis device 14, and receives the first meeting record M1 and the second meeting record M2. Specifically, the artificial intelligence processing circuit 16 is configured to integrate the first meeting record M1 and the second meeting record M2, and input the integrated meeting record into a natural language processing (NLP) and machine learning model for analysis and generation of a comprehensive third meeting record M3.

[0063] In other words, the intelligent conference assistance system 1 of this embodiment can combine text recognition, chart recognition, and speech recognition to generate the first meeting record M1 and the second meeting record M2 that respectively record the image content and the speech content, and integrate the first meeting record M1 and the second meeting record M2 through timestamps. Then, the intelligent conference assistance system 1 of this embodiment can use natural language processing and machine learning models to further analyze the integrated meeting record.

[0064] Basically, through natural language processing and machine learning models, the artificial intelligence processing circuit 16 can better understand the image content and speech content recorded in the first meeting record M1 and the second meeting record M2, and can also correct and supplement the content with recording errors and incomplete recordings, and then extract and organize the context information, so as to generate a comprehensive third meeting record M3. That is to say, the third meeting record M3 will record all the image content and speech content, and all the image content and speech content will be reflected in the third meeting record M3 in sequence. According to the above content, as Figure 2 shown, the method of this embodiment may further include the following steps.

[0065] Step S15: Configure the artificial intelligence processing circuit to integrate the first meeting record and the second meeting record, and input the integrated meeting record into a natural language processing and machine learning model for analysis and generation of a third meeting record.

[0066] The following will further illustrate how the artificial intelligence processing circuit 16 corrects and supplements the content with recording errors and incomplete recordings. However, the present invention is not limited to the following examples. For example, a certain image content recorded in the first meeting record M1 may be the handwritten text content of "U8B", but in fact, the meeting presenter 5 wrote the handwritten text content of "USB" on the interactive device 2 at this time. That is to say, there is a text recognition error, resulting in the first meeting record M1 recording the wrong text content.

[0067] Next, according to the time stamp, when integrating the first meeting record M1 and the second meeting record M2, the artificial intelligence processing circuit 16 can find at least one voice content associated with the foregoing image content, and the at least one voice content is that the meeting presenter 5 is explaining the Universal Serial Bus to the meeting participants. Therefore, through natural language processing and machine learning models, the artificial intelligence processing circuit 16 can understand that the foregoing handwritten text content should be "USB" instead of "U8B", so that the artificial intelligence processing circuit 16 can also correct the foregoing handwritten text content.

[0068] On the other hand, another image content recorded in the first meeting record M1 may be the pie chart content of "the support rate of candidates". That is to say, at this time, there is a pie chart reflecting "the support rate of candidates" on Image 4, but the chart recognition circuit 122 may not be able to successfully identify which candidate's support rate each sector on the pie chart represents, resulting in the first meeting record M1 recording incomplete pie chart content.

[0069] Similarly, according to the time stamp, the artificial intelligence processing circuit 16 can find at least one voice content associated with the foregoing image content, and the at least one voice content is that the meeting presenter 5 is explaining which candidate's support rate each sector on the pie chart represents to the meeting participants. Therefore, through natural language processing and machine learning models, the artificial intelligence processing circuit 16 can also supplement the foregoing incomplete pie chart content.

[0070] Furthermore, the artificial intelligence processing circuit 16 can also be configured to output a real-time third meeting record M3. Therefore, the intelligent meeting assistance system 1 of this embodiment may further include an output device 17. The output device 17 is coupled to the artificial intelligence processing circuit 16 and can be configured to store and / or display the third meeting record M3. For example, the output device 17 can be a smart phone, a laptop computer, an external storage device, or a set-top box, but the present invention is not limited thereto. Similarly, the present invention does not limit the specific forms of the image content and voice content recorded in the third meeting record M3.

[0071] Next, the following is to illustrate the implementation manners of the optical character recognition circuit 121, the chart recognition circuit 122, and the dynamic display recognition circuit 123 through specific specific embodiments, but the present invention is not limited thereto. Please refer to Figures 4A to 4C , Figure 4A is the functional block diagram of the optical character recognition circuit of the embodiment of the present invention, Figure 4B is the functional block diagram of the chart recognition circuit of the embodiment of the present invention, Figure 4C is the functional block diagram of the dynamic display recognition circuit of the embodiment of the present invention.

[0072] AsFigure 4A As shown, the optical character recognition circuit 121 may include a Picture Selection circuit 1211, a Picture Segmentation circuit 1212, a Picture Reconstruct circuit 1213, and an optical character recognition engine 1214, to generate a first event record for recording text content. However, Figure 4A the first event record is not shown either. Since the application principle of character recognition is well-known to those skilled in the art, details of the Picture Selection circuit 1211, the Picture Segmentation circuit 1212, the Picture Reconstruct circuit 1213, and the optical character recognition engine 1214 will not be elaborated further.

[0073] It should be noted that since the image analysis device 12 can be configured to add time stamps to the text content, chart content, and / or dynamic display content of each image content, the optical character recognition circuit 121 may further include a time stamp and event combination circuit 1215. The time stamp and event combination circuit 1215 is coupled to the optical character recognition engine 1214 and is configured to add a time stamp to the text content recorded in the first event record.

[0074] As Figure 4B shown, the chart recognition circuit 122 may include a Picture Selection circuit 1221, a Picture Segmentation circuit 1222, and a chart recognition engine 1223, to generate a second event record for recording chart content. However, Figure 4B the second event record is not shown either. Since the application principle of chart recognition is well-known to those skilled in the art, details of the Picture Selection circuit 1221, the Picture Segmentation circuit 1222, and the chart recognition engine 1223 will not be elaborated further.

[0075] It should be noted that since the chart recognition circuit 122 recognizes the pie chart content, line chart content, and bar chart content on the image 4, the chart recognition engine 1223 may further include a pie chart recognition engine 12231, a line chart recognition engine 12232, and a bar chart recognition engine 12233. Since the application principles of pie chart recognition, line chart recognition, and bar chart recognition are also well-known to those skilled in the art, details of the pie chart recognition engine 12231, the line chart recognition engine 12232, and the bar chart recognition engine 12233 will not be elaborated further.

[0076] Similarly, since the image analysis device 12 can be configured to add timestamps to the text content, chart content, and / or dynamic display content of each image content, the chart recognition circuit 122 may further include a timestamp and event combination circuit 1224. The timestamp and event combination circuit 1224 is coupled to the chart recognition engine 1223 and is configured to add timestamps to the chart content recorded in the second event record.

[0077] As Figure 4C shown, the dynamic display recognition circuit 123 may include an image segmentation circuit 1231 and a dynamic display recognition engine 1232 for generating a third event record recording the dynamic display content. However, Figure 4C the third event record is not shown either. Since the application principle of dynamic display recognition is already well-known to those skilled in the art, details of the image segmentation circuit 1231 and the dynamic display recognition engine 1232 will not be elaborated further.

[0078] Similarly, since the image analysis device 12 can be configured to add timestamps to the text content, chart content, and / or dynamic display content of each image content, the dynamic display recognition circuit 123 may further include a timestamp and event combination circuit 1233. The timestamp and event combination circuit 1233 is coupled to the dynamic display recognition engine 1232 and is configured to add timestamps to the dynamic display content recorded in the third event record. It should be noted again that the present invention does not limit the specific implementation manners of the optical character recognition circuit 121, the chart recognition circuit 122, and the dynamic display recognition circuit 123.

[0079] On the other hand, the following uses specific embodiments to illustrate the implementation manner of the speech recognition circuit 141, but the present invention is not limited thereto. Please refer to Figure 5 , Figure 5 which is a functional block diagram of the speech recognition circuit according to an embodiment of the present invention.

[0080] As Figure 5 shown, the speech recognition circuit 141 may include a sound preprocessing circuit 1411, a speech recognition engine 1412, a cloud recognition unit 1413, and a local recognition unit 1414. Since the application principle of speech recognition is already well-known to those skilled in the art, details of the sound preprocessing circuit 1411, the speech recognition engine 1412, the cloud recognition unit 1413, and the local recognition unit 1414 will not be elaborated further.

[0081] Similarly, the speech recognition circuit 141 may further include a timestamp and event combination circuit 1415. The timestamp and event combination circuit 1415 is coupled to the speech recognition engine 1412 and is configured to add timestamps to the speech content.

[0082] Furthermore, the intelligent conference assistance system 1 can display speech content, text content, and chart content in a Transcript mode or a Recap mode. Please refer to Figure 6A , Figure 6A which is a schematic diagram of the intelligent conference assistance system according to an embodiment of the present invention displaying speech content in the Transcript mode. As Figure 6A shown, each speech content may include a sentence. Therefore, in the Transcript mode, multiple sentences (e.g., Sent1 to Sent5) can be displayed on the screen of the output device 17 in real time and arranged in chronological order.

[0083] In addition, in the case of multiple speakers, the speech recognition circuit 141 can also identify the speaker of each sentence. For example, each sentence in this embodiment may correspond to speaker Sp1 or Sp2, and speakers Sp1 and Sp2 can be the conference presenter 5 and a certain conference participant, but the present invention is not limited thereto. Therefore, as Figure 6A shown, in the Transcript mode, the intelligent conference assistance system 1 can also display the corresponding speaker for each sentence.

[0084] Next, please refer to Figure 6B , Figure 6B which is a schematic diagram of the intelligent conference assistance system according to an embodiment of the present invention displaying speech content in the Recap mode. As Figure 6B shown, in the Recap mode, the intelligent conference assistance system 1 can display all the sentences of the same speaker. In addition, as mentioned above, the sound analysis device 14 can also be configured to add timestamps to each speech content. Therefore, sentences Sent1 to Sent5 in this embodiment may respectively correspond to timestamps St1 to St5, and in the Recap mode, the intelligent conference assistance system 1 can also display the timestamp of each sentence.

[0085] Furthermore, the artificial intelligence processing circuit 16 can also edit, collate, customize, and optimize the meeting record to generate a more concise quick summary. Therefore, in the Recap mode, the intelligent conference assistance system 1 can also display the quick summary QS generated by the artificial intelligence processing circuit 16 to help the user review the entire meeting process.

[0086] On the other hand, compared with the quick summary, the artificial intelligence processing circuit 16 can also generate a more complete meeting summary. Therefore, the intelligent conference assistance system 1 can also display the meeting summary generated by the artificial intelligence processing circuit 16 in a summary mode. Please refer to Figure 6C , Figure 6C which is a schematic diagram of the intelligent conference assistance system according to an embodiment of the present invention displaying the meeting summary in the summary mode.

[0087] AsFigure 6C As shown, the artificial intelligence processing circuit 16 can edit, collate, and optimize the meeting minutes to generate a meeting summary including a title MS1, abstract content MS2, keywords MS3, pain points MS4, action items MS5, and a chart MS6. Therefore, in the summary mode, the title MS1, abstract content MS2, keywords MS3, pain points MS4, action items MS5, and chart MS6 of the meeting summary can be displayed on the screen of the output device 17. In this embodiment, the chart MS6 of the meeting summary will take a pie chart as an example, but the present invention is not limited thereto.

[0088] Furthermore, the output device 17 can provide a graphical user interface including a check box B1, allowing the user to decide whether to display the keywords MS3, pain points MS4, action items MS5, and chart MS6 of the meeting summary. Additionally, the graphical user interface provided by the output device 17 may further include a button B2 with the text "Convert Chart" thereon, and in response to the button B2 being pressed, the user can convert the type of the chart MS6. Similarly, the graphical user interface provided by the output device 17 may further include buttons B3, B4, and B5 with the text "Edit", "Email", and "Store" thereon, and in response to the button B3, B4, or B5 being pressed, the user can edit the meeting summary, send the meeting summary by email, or store the meeting summary.

[0089] In summary, one beneficial effect of the present invention is that the intelligent meeting assistance system and the method for generating meeting minutes provided by the present invention can generate meeting minutes recording the image content through technical means of "capturing the images displayed by the interactive device during the meeting" and "performing an image analysis program on the images".

[0090] Furthermore, the intelligent meeting assistance system and the method for generating meeting minutes provided by the present invention can record the image content other than the voice content, especially can capture and record the content written and drawn by the meeting presenter on the interactive device, so the generated meeting minutes can more accurately reflect all the behaviors in the meeting. Additionally, the intelligent meeting assistance system and the method for generating meeting minutes provided by the present invention can integrate the voice content, text content, and chart content to generate a comprehensive meeting minutes, and through natural language processing and machine learning models, can also correct and supplement the content with recording errors and incomplete records to generate higher-quality meeting minutes.

[0091] The content disclosed above is only the preferred feasible embodiment of the present invention, and does not limit the protection scope of the claims of the present invention. Therefore, all equivalent technical changes made by using the content of the specification and drawings of the present invention are included in the protection scope of the claims of the present invention.

Claims

1. An intelligent conference assistance system, characterized in that: The intelligent conference auxiliary system includes: an image capture device configured to capture an image displayed by an interactive device during a meeting; and An image analysis device is coupled to the image capture device and is configured to execute an image analysis program on the image to generate a first meeting record recording an image content.

2. The intelligent conference assistance system according to claim 1, characterized in that: The image content includes one or a combination of a text content and a graphic content on the image, and the image analysis program includes a text recognition program and a graphic recognition program.

3. The intelligent conference assistance system according to claim 2, characterized in that: The image analysis device comprises: an optical character recognition circuit configured to execute the text recognition procedure on the image to generate a first event record recording the text content; and A chart recognition circuit is configured to execute the chart recognition procedure on the image to generate a second event record recording the contents of the chart.

4. The intelligent conference assistance system according to claim 3, characterized in that: The text content includes one or a combination of a printed text content and a handwritten text content on the image, and the optical character recognition circuit executes the text recognition program on the image to recognize the printed text content and the handwritten text content on the image.

5. The intelligent conference assistance system according to claim 3, characterized in that: The chart content includes one or a combination of a pie chart content, a line chart content and a bar chart content on the image, and the chart recognition circuit executes the chart recognition program on the image to recognize the pie chart content, the line chart content and the bar chart content on the image.

6. The intelligent conference assistance system according to claim 3, characterized in that: The image analysis program also includes a dynamic display recognition program, and the image content includes one or a combination of the text content, the graphic content and a dynamic display content on the image.

7. The intelligent conference assistance system according to claim 6, characterized in that: The image analysis device also includes: a dynamic display recognition circuit configured to execute the dynamic display recognition procedure on the image to generate a third event record recording the dynamic display content; The image analysis device is further configured to generate the first meeting record recording the image content according to the first event record, the second event record and the third event record.

8. The intelligent conference assistance system according to claim 7, characterized in that: The intelligent conference auxiliary system also includes: a voice input device configured to generate voice data during the conference; and A sound analysis device is coupled to the sound input device and is configured to execute a sound analysis program on the sound data to generate a second meeting record recording a plurality of voice contents.

9. The intelligent conference assistance system according to claim 8, characterized in that: The image capture device is also configured to sequentially capture the multiple images displayed by the interactive device during the meeting, and the image analysis device generates the first meeting record that records the contents of the multiple images respectively.

10. The intelligent conference assistance system according to claim 9, characterized in that: The image analysis device is further configured to add a timestamp to each of the image contents recorded in the first meeting record, and the sound analysis device is further configured to add the timestamp to each of the voice contents recorded in the second meeting record.

11. The intelligent conference assistance system according to claim 10, characterized in that: The intelligent conference auxiliary system also includes: An artificial intelligence processing circuit is coupled to the image analysis device and the sound analysis device, and receives the first meeting record and the second meeting record, wherein the artificial intelligence processing circuit is configured to integrate the first meeting record and the second meeting record, and input the integrated meeting record into a natural language processing and machine learning model for analysis and generation of a third meeting record.

12. A method for generating meeting minutes, characterized in that: The method comprises: configuring an image capture device to capture an image displayed by an interactive device during a meeting; and An image analysis device is configured to execute an image analysis program on the image to generate a first meeting record recording an image content.

13. The method according to claim 12, characterized in that The image content includes one or a combination of a text content and a graphic content on the image, and the image analysis program includes a text recognition program and a graphic recognition program.

14. The method according to claim 13, characterized in that The step of configuring the image analysis device to perform the image analysis procedure on the image comprises: configuring an optical character recognition circuit to perform the text recognition procedure on the image to generate a first event record recording the text content; and A chart recognition circuit is configured to execute the chart recognition procedure on the image to generate a second event record recording the contents of the chart.

15. The method according to claim 14, characterized in that The image analysis program also includes a dynamic display recognition program, and the image content includes one or a combination of the text content, the graphic content and a dynamic display content on the image.

16. The method according to claim 15, characterized in that The step of configuring the image analysis device to perform the image analysis procedure on the image further comprises: configuring a dynamic display recognition circuit to execute the dynamic display recognition procedure on the image to generate a third event record recording the dynamic display content; and The image analyzing device is configured to generate the first meeting record recording the image content based on the first event record, the second event record and the third event record.

17. The method according to claim 16, characterized in that The method further comprises: configuring a sound input device to generate sound data during the conference; and A sound analysis device is configured to execute a sound analysis program on the sound data to generate a second meeting record recording a plurality of voice contents.

18. The method according to claim 17, characterized in that The image capture device is also configured to sequentially capture the multiple images displayed by the interactive device during the meeting, and the image analysis device generates the first meeting record that records the contents of the multiple images respectively.

19. The method according to claim 18, characterized in that The image analysis device is further configured to add a timestamp to each of the image contents recorded in the first meeting record, and the sound analysis device is further configured to add the timestamp to each of the voice contents recorded in the second meeting record.

20. The method according to claim 19, characterized in that The method further comprises: An artificial intelligence processing circuit is configured to integrate the first meeting record and the second meeting record, and the integrated meeting record is input into a natural language processing and machine learning model to analyze and generate a third meeting record.