Information processing system
Patent Information
- Application Number
- CN202610319348.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-16
- Publication Date
- 2026-09-22
AI Technical Summary
1. 在复杂或专业性较强的议题下,人类与会者难以及时、系统地综合大量信息生成高质量决策意见,容易产生信息筛选不充分、分析不全面的问题
终端在一种实施形态中呈现表单式界面,允许用户分别输入会议议题、背景描述和讨论要点。终端在另一实施形态中提供实时语音输入功能。
Smart Images

Figure CN122802486A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] In existing technologies, meeting decision-making processes typically rely on human participants making judgments based on their own experience and limited information, which presents the following problems: 1. On complex or highly specialized topics, human participants often struggle to synthesize large amounts of information in a timely and systematic manner to generate high-quality decision-making opinions, which can easily lead to problems such as insufficient information filtering and incomplete analysis.
[0004] 2. Traditional meeting minutes are mostly limited to manual writing or simple recording, lacking automatic transcription and structured processing of audio data, making it difficult to efficiently retrieve and reuse meeting information.
[0005] 3. Existing meeting support systems generally only record content without analyzing the emotional state of participants. They cannot comprehensively consider factors such as participants' emotional fluctuations and attitudes in decision-making, which is not conducive to improving the rationality and acceptability of meeting decisions.
[0006] 4. When disseminating conference outcomes and innovative practices to the outside world, companies often rely on manual editing and organization, lacking an external communication mechanism based on systematic data and automatically generated opinions, making it difficult to fully explore the value of conference activities for brand promotion and marketing.
[0007] Therefore, it is necessary to provide a new system that can automatically generate prompts based on meeting topics to instruct generative artificial intelligence models to generate opinions, automatically acquire and transcribe meeting audio data, analyze the emotions of meeting participants and incorporate these emotional factors into the opinion generation process, and support the external release of generated opinions and meeting outcomes, so as to improve the quality and efficiency of meeting decision-making, improve meeting information management, and enhance the effectiveness of corporate external communication and marketing. Summary of the Invention
[0008] To address the aforementioned issues, the present invention provides an information processing system comprising a processor, wherein the processor is configured to: 1. A prompt message is generated based on the meeting's agenda to instruct a generative artificial intelligence model to generate opinions. Specifically, the processor acquires meeting agenda information and background data related to the agenda items, and organizes elements such as the agenda title, agenda description, problems to be solved, and risk points to be addressed into structured prompt text, which is used as input to the generative artificial intelligence model so that the model can output analytical opinions and suggestions for the agenda item.
[0009] 2. Acquire audio data and convert it into text data. The processor receives audio data generated during the meeting through communication with the terminal or audio acquisition device, calls the speech recognition module to recognize the audio data, converts the recognition results into text data, and associates and stores them according to time sequence and speaker information, thereby forming a searchable and analyzable meeting transcript.
[0010] 3. Analyze the emotions of meeting participants and generate opinions that take these emotions into account. The processor performs emotion analysis on the audio data and / or the text data, for example, by identifying the emotional states of participants (such as agreement, doubt, and concern) through features such as speech characteristics, speech rate, tone changes, and text word usage tendencies. The emotion analysis results are then provided as additional input parameters to the generative artificial intelligence model, enabling the model to comprehensively consider the participants' emotional tendencies when generating meeting opinions, thereby outputting opinions that are more in line with the actual discussion atmosphere and more easily accepted.
[0011] According to a preferred embodiment of the present invention, the processor is configured to either input meeting content through an information management department or automatically acquire meeting content using speech recognition technology. By setting up an input interface for the information management department to manually input meeting summaries, agenda descriptions, and supplementary information, combined with real-time speech content acquired through automatic speech recognition, the system can more comprehensively and accurately grasp meeting information, providing a rich data foundation for subsequent opinion generation and sentiment analysis.
[0012] According to another preferred embodiment of the present invention, the processor is configured to: publish the generated opinions and meeting outcomes externally, so as to generate buzz and achieve marketing effects regarding the innovative initiative. Specifically, the processor organizes and summarizes the opinions output by the generative artificial intelligence model and the meeting conclusions optimized by sentiment analysis, generates content suitable for external dissemination, and publishes the content through the company's official website, social media platforms, or other external communication channels according to preset strategies or user instructions, thereby showcasing the company's innovative practices in introducing artificial intelligence into decision-making, enhancing brand image, and creating new marketing opportunities. Through the above structure and process, the present invention can effectively solve the problems of insufficient meeting decision support, lack of consideration of emotional factors, and underutilization of the value of external dissemination of meeting outcomes in the prior art.
[0013] "System" refers to a combination of at least one processor and optional memory, communication interface and input / output devices, used to perform the functions described in this invention in a conference scenario, and can be implemented in the form of a server, cloud platform, local device or any combination thereof.
[0014] A “processor” refers to a hardware or virtual computing unit that can execute program instructions to process data and perform specific functions, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a neural network accelerator chip, or a virtual processing resource deployed on a server or in the cloud.
[0015] A "meeting" is an activity organized to discuss one or more topics, exchange information, and make decisions. It includes board meetings, management meetings, project review meetings, strategic seminars, etc., and can be in-person meetings, remote meetings, or a combination of both.
[0016] An “issue” refers to a specific matter or problem that is pre-set or raised on the spot during a meeting and needs to be discussed and / or concluded. Each issue typically includes a title, background information, and the goal to be addressed.
[0017] "Generative artificial intelligence models" refer to artificial intelligence models trained based on machine learning and / or deep learning techniques that can automatically generate text, speech, or other forms of output content based on input prompts, including but not limited to large language models and multimodal generative models.
[0018] "Prompt information" refers to the combination of input data and instructions used to guide generative artificial intelligence models to generate specific types of output content. It usually includes a description of the meeting topic, background information, analysis requirements, output format requirements, etc., and can also be called prompt words or prompt text.
[0019] "Audio data" refers to digital sound signals that carry the voice information of participants and are collected during a meeting, including voice records collected by microphones, recording devices or terminals and stored or transmitted in digital form.
[0020] “Text data” refers to data that represents the content of a meeting speech in text form, obtained from audio data through speech recognition or manual transcription, as well as meeting transcripts directly input or edited by users.
[0021] "Meeting participants" refers to persons who actually attend the meeting and may speak, express opinions, or vote, including but not limited to directors, senior executives, employees, advisors, and other invited persons.
[0022] "Emotion" refers to the psychological and attitudinal state characteristics exhibited by meeting participants during speaking or communication, such as agreement, disagreement, doubt, worry, excitement, calmness, etc. It can be identified and classified through voice features, word choice tendencies, tone intensity, etc.
[0023] "Analyzing the emotions of meeting participants" refers to the process of inferring the emotional category and intensity of participants at a certain time or on a certain topic by extracting features and recognizing patterns from audio and / or text data, and representing the results as structured information that can be used for subsequent processing.
[0024] "Generating opinions" refers to the process by which a generative artificial intelligence model outputs analytical, suggestive, or evaluative text about a specific topic or meeting content, based on meeting topics, meeting information, sentiment analysis results of meeting participants, and prompts. This text can then be post-processed by a processor to form conclusions or suggestions that can be used for decision-making.
[0025] "Meeting information" refers to all kinds of data and content related to the meeting, including meeting time, participants, list of topics, background information, real-time or historical speech content, transcripts and related attachments.
[0026] "Information management department" refers to a department or position within an organization that is responsible for functions such as document management, meeting information organization, system operation and maintenance, including but not limited to the information systems department, general affairs department, secretariat, or a dedicated meeting management team.
[0027] "Speech recognition technology" refers to the technology that automatically converts human speech signals into corresponding text data, including end-to-end speech recognition models, acoustic model and language model combination systems, and various speech-to-text systems running on local devices or cloud services.
[0028] "External release" refers to the act of providing opinions generated by generative artificial intelligence models and content compiled based on meeting outcomes to the external public or specific external parties through online platforms, media channels, or other information dissemination methods, in the form of press releases, reports, announcements, promotional materials, etc.
[0029] "Topicability" refers to the ability to generate attention, discussion, and dissemination among external audiences, including but not limited to gaining high exposure and discussion in media reports, social networks, industry forums, and other scenarios.
[0030] "Marketing effectiveness" refers to the positive impact of releasing content related to AI-involved decision-making activities on aspects such as increased brand awareness, improved corporate image, acquisition of potential customers, and increased business opportunities. It can be evaluated through indicators such as page views, inquiries, conversion rates, and the number of media reports. Attached Figure Description
[0031] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0032] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0033] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0034] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0035] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0036] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0037] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0038] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0039] Figure 9 This represents an emotion map that maps multiple emotions.
[0040] Figure 10This represents an emotion map that maps multiple emotions.
[0041] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0042] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0043] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0044] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0045] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.
[0046] First, let me explain the terminology used in the following instructions.
[0047] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0048] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0049] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0050] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0051] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0052] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0053] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0054] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0055] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0056] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0057] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0058] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0059] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0060] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0061] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0062] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0063] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0064] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0065] In modern meeting decision-making processes, with the dramatic increase in the amount of meeting information and the diversification of participants, relying solely on manual recording and summarizing of meeting information is no longer sufficient to grasp the key points of discussion in a timely and comprehensive manner and to form high-quality decision opinions. Traditional computer-aided conferencing systems mostly offer only simple speech-to-text transcription, text recording, or static template-based report generation functions. Technically, they suffer from the following shortcomings: First, the processing flow often remains at the level of linear storage and display of raw text or audio, lacking the ability to perform deep semantic understanding, summary extraction, and structured data processing of meeting information, resulting in limited effectiveness of subsequent automated decision support. Second, existing systems generally treat the invocation of generative AI models as a "black box" of simple question-and-answer interfaces, failing to establish an internal mechanism for constructing prompts optimized for generative AI models, thus failing to fully utilize the model's reasoning and generation capabilities. Third, traditional systems almost never model or perform data processing on the emotional states of meeting participants, failing to integrate emotional information as one of the input dimensions into the opinion generation process. This makes the generated results difficult to adapt to the actual meeting atmosphere and the psychological state of the participants, affecting the adoptability and persuasiveness of decision recommendations. Furthermore, the generated results are often presented in plain text form, lacking structured organization and data formats oriented towards downstream applications (such as subsequent queries, reuse, and external publication), which is not conducive to efficient indexing, retrieval, and automated reuse within the computer system.
[0066] Therefore, it is necessary to design a new information processing system from the perspective of computer technology, which performs multi-stage data processing and computation on raw meeting information in the same computing environment. This includes multi-source meeting information acquisition, natural language preprocessing, summarization and keyword extraction, automatic construction of prompt statements, invocation of generative artificial intelligence models combined with emotional information, and structured and external publishing of output results. This will improve the overall processing capability and performance of computers in supporting meeting decisions in terms of system architecture and algorithm flow, and improve the efficiency and quality of human-computer collaborative decision-making.
[0067] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0068] In this invention, the server includes: a device for acquiring and recording information related to meeting topics and content, and unifying the meeting information into text-based meeting information; a device for performing natural language processing algorithms on the text-based meeting information to perform summarization and key word extraction processing, thereby generating summary information and key word information of the meeting information, and automatically constructing prompt statements for instructing a generative artificial intelligence model to generate meeting opinions based on the summary information and key word information and meeting topic information; a device for acquiring voice data from a voice acquisition device and converting the voice data into text data using a voice recognition algorithm, and integrating the text data into the text-based meeting information; and a device for... The system comprises: a device for analyzing the emotional information of meeting participants using sentiment inference algorithms based on the meeting information and the aforementioned voice data, and reflecting this emotional information in the prompt statements and / or generated meeting opinions; a device for inputting the prompt statements into the generative artificial intelligence model, enabling the model to perform data processing and computation to generate meeting opinion information, organizing the generated opinion information into structured information, and sending it to a terminal device to facilitate multi-faceted and rapid decision-making by meeting participants; and a device for generating meeting outcome information based on the generated opinion information and the meeting information when needed, and outputting it to an external information providing device, which then provides the meeting outcome information to external entities through an information publishing device. This establishes a complete technical chain within the computer, from multi-source meeting information acquisition, semantic summarization and keyword extraction, sentiment feature acquisition, prompt statement construction optimized for generative artificial intelligence models, to structured decision opinion generation and external publication. This allows the server to perform deeper data processing and computation when processing meeting information and generating decision opinions, significantly improving the processing efficiency, relevance, and usability of the generated results of the meeting decision support system, and enhancing the overall technical performance of human-machine collaborative decision-making.
[0069] A "system" refers to a comprehensive computer implementation scheme consisting of multiple functional devices or modules connected and working together through communication methods to collect, process, generate, and output conference information.
[0070] "Information processing device" refers to a computing device including at least one processor and memory, which is hardware or a combination of hardware and software used to execute program instructions to process input data and output processing results.
[0071] A "server" refers to an information processing device in a network that provides data processing and data storage services to one or more terminal devices. It completes meeting information processing, model invocation, and result output by executing predetermined programs.
[0072] "Terminal device" refers to an electronic device operated by a user to access a server and display processing results, including but not limited to mobile terminals, desktop terminals, or other computing devices with communication and display functions.
[0073] "Meeting topics" refer to the themes that are pre-set or determined in real time during the meeting and require discussion and decision-making.
[0074] "Meeting information" refers to various types of data related to the meeting, including but not limited to meeting topics, meeting discussion content, participant information, and meeting background information.
[0075] "Text-based meeting information" refers to meeting information represented as a sequence of characters, including text data directly entered by the user and text data converted from voice data.
[0076] "Natural Language Processing Algorithms" refers to a set of algorithms used to analyze and process natural language text, including methods such as word segmentation, part-of-speech tagging, syntactic analysis, text summarization, and keyword extraction.
[0077] "Summary processing" refers to the process of compressing and summarizing raw meeting information to generate short text information that represents the main content.
[0078] "Important word extraction and processing" refers to the process of identifying and selecting semantically or statistically representative words or phrases from textual meeting information to represent the key points of the meeting information.
[0079] "Summary information" refers to a summary text extracted from the original meeting information through summarization processing, used to concisely express the core content of the meeting information.
[0080] "Important word information" refers to a data set consisting of one or more important words or phrases, obtained through important word extraction and processing.
[0081] "Prompt statements" refer to input text constructed to instruct generative artificial intelligence models to perform specific generative tasks, and include at least some or all of the following: summary information, key word information, and conference topic information.
[0082] "Generative artificial intelligence models" refer to artificial intelligence models built on machine learning algorithms and large-scale training data that can automatically generate new text information based on input prompts.
[0083] "Speech acquisition device" refers to a hardware device used to acquire sound signals that occur during a meeting and convert them into audio data that can be processed by computing devices.
[0084] “Voice data” refers to audio signal data obtained by a voice acquisition device and represented in digital form.
[0085] "Speech recognition algorithm" refers to a pattern recognition algorithm or model used to convert speech data into corresponding text data.
[0086] “Text data” refers to data in the form of characters or strings that are recognized and output by speech recognition algorithms from speech data.
[0087] "Participant in the meeting" refers to one or more individuals or entities that participate in the discussion at the meeting, including but not limited to speakers and audience members.
[0088] "Emotion inference algorithm" refers to a computational method or model that analyzes and infers the emotional state or attitude of meeting participants based on voice data, text data, or related features.
[0089] "Emotional information" refers to data obtained by emotion inference algorithms that represent the emotional state or sentiment of meeting participants, including but not limited to emotion category and emotion intensity.
[0090] "Opinion information" refers to textual information such as suggestions, insights, or solutions related to the meeting topics, generated by a generative artificial intelligence model based on meeting information and prompts.
[0091] "Data processing" refers to the process of performing operations such as format conversion, filtering, aggregation, reorganization, or enhancement on raw data to facilitate subsequent calculations or generation.
[0092] "Data computation" refers to the computational process performed on data using algorithms or models in electronic computing devices, including statistical computation, logical reasoning, probabilistic inference, and vector operations.
[0093] "Structured information" refers to data that has been organized and structured with clearly defined fields or hierarchical structures, so that computers can store, retrieve, and further process it.
[0094] "Information management device" refers to a computing device or system used for inputting, editing and managing meeting information, which can be operated by the management department to maintain meeting-related data.
[0095] "Input device" refers to human-computer interaction equipment used to input meeting information or instructions into information processing equipment, including keyboards, touch screens, mice, voice input interfaces, etc.
[0096] "Acquisition device" refers to a data acquisition component or module used to automatically or semi-automatically collect meeting information from the external environment, including functional units that work in conjunction with voice acquisition devices and voice recognition algorithms.
[0097] "External information providing device" refers to an external computing device or service system used to receive conference outcome information from the system output and further store, display or distribute it.
[0098] "Information dissemination device" refers to a device or system used to publicly or selectively disseminate information about the results of a meeting to external entities, including online publishing platforms and message push systems.
[0099] "Meeting outcome information" refers to data compiled from generated opinion information and meeting information, used to represent meeting conclusions, decision results, or key outputs.
[0100] "External subjects" refers to entities outside the system that receive information about the conference outcomes, including external users, external systems, or the public.
[0101] "Business results" refers to the positive outcomes in areas such as business development, brand influence, or market performance resulting from the external release and utilization of conference outcomes information.
[0102] In the following implementations, the subject is limited to "server", "terminal", or "user", and the processing content is not described in the form of program code.
[0103] I. System Overall Composition and Hardware Structure Example In one implementation, the server employs a computing platform with multi-core central processing units and graphics processing units, such as a rack-mounted computer based on a multi-core CPU and at least one GPU supporting general-purpose parallel computing. In another implementation, the server is deployed in a cloud computing environment, consisting of virtual machines or container clusters. In both environments, the server runs an operating system, such as a Unix-like operating system, and application server software, such as web server software and application frameworks. In some implementations, the server also includes a relational database management system for storing meeting information and generated results.
[0104] In one embodiment, the terminal is a smart terminal running a mobile operating system; in another embodiment, it is a personal computer running a desktop operating system. The terminal communicates with the server via a network. The terminal includes a display device, an input device, and an optional audio acquisition device.
[0105] When using the system, users interact with the server through the terminal, input meeting information, view the comments generated by the server, and ask follow-up questions when necessary.
[0106] II. Server-side program modules and data structures In one implementation, the server uses multiple functional modules to implement the system described in this invention. In different implementations, these modules can be deployed as independent services or integrated into a single application. The logical functions are described below.
[0107] 1. Meeting Information Management Module The server maintains multiple data structures in the meeting information management module. In one implementation, the server creates a meeting record object for each meeting. A meeting record includes at least the following fields: meeting identifier, meeting agenda text, original meeting information text, preprocessed meeting information text, summary information text, list of key words, prompt text, sentiment vector, generated opinion information text, structured opinion data, and metadata. The server stores these records in the database as tables or collections. The field structure is fixed, allowing for rapid retrieval and indexing through key-value and field combinations.
[0108] 2. Text Preprocessing and Natural Language Processing Module The server employs text cleaning, word segmentation, and syntactic analysis algorithms in its natural language processing module. In one implementation, the server utilizes open-source natural language processing libraries to implement these functions; for example, it segments Chinese text using word segmentation tools and cleans up non-textual symbols using regular expressions. In another implementation, the server can use an encoder based on a pre-trained language model to vectorize the text, segmenting long texts into fixed-dimensional feature vectors for subsequent summarization and keyword extraction.
[0109] The server employs a specific data structure during the summary generation process. It segments the original text into a sentence list. The server calculates a weight for each sentence, which can include features such as sentence position in the text, sentence length, and keyword coverage. In one implementation, the server uses a graph-based ranking algorithm, treating sentences as graph nodes, constructing edges based on similarity, and iteratively calculating scores for important sentences to select a few sentences to form the summary. This algorithm, implemented internally by the computer through multiple matrix multiplications and iterative updates, is more stable than simple sentence selection based on position or manual rules, reducing the error of missing important information in the summary.
[0110] During keyword extraction, the server combines word segmentation results with word frequency statistics. In one implementation, the server calculates the TF-IDF value of each candidate word, filters stop words and overly short words, and selects several keywords after sorting them by score. In another implementation, the server uses pre-trained word vectors or context vectors as features to cluster or score the candidate words for relevance, in order to more accurately identify terms that are meaningful for meeting decisions. In this way, the server internally forms a set of key feature words for constructing prompts, reducing irrelevant words entering the model input, thereby reducing input redundancy in the generative artificial intelligence model and improving computational efficiency.
[0111] 3. Prompt Statement Construction Module In the prompt statement construction module, the server organizes summary information, key word information, and meeting agenda information into a structured text template. In one implementation, the server pre-stores multiple templates for different types of meetings, and these templates contain placeholder fields. During execution, the server populates the summary placeholders with summary information, concatenates the list of key words by delimiters and populates the keyword placeholders, and populates the meeting agenda text into the agenda placeholders, forming a complete and readable prompt statement.
[0112] Here is an example of a prompt statement built by the server in one implementation: "As a corporate strategy consultant, please provide multi-faceted opinions on the 'New Product Launch Strategy' based on the following meeting information."
[0113] Conference Summary: This conference focuses on new product launch strategies in the SME market, discussing issues such as competitor status, target customer groups, and budget constraints.
[0114] Key keywords: new product; market launch; target customers; competitive analysis; pricing strategy; promotion channels; budget constraints.
[0115] Please provide clearly structured and actionable recommendations covering four aspects: target market selection, pricing strategy, distribution channels, and key risks. In another implementation, when a user asks an additional question, the server constructs an enhanced prompt statement based on the existing summary and previously generated results, for example: "Regarding the previous suggestions on 'new product market launch strategy,' please elaborate further: If we prioritize medium-sized manufacturing customers, what low-cost but high-impact marketing actions should be taken in the first three months? Please provide at least five actionable suggestions, and explain the expected effects and potential risks of each suggestion." Through this prompt statement construction process, the server internally implements a preprocessing mechanism that is not a simple forwarding mechanism. It summarizes the original meeting information into a structured context that the model can use efficiently, effectively reducing input redundancy, thereby reducing the computational burden and communication data volume of the generative artificial intelligence model and improving processing speed.
[0116] 4. Speech Processing and Sentiment Analysis Module The server receives audio data uploaded by the terminal in its speech processing module. In one implementation, the server invokes a speech recognition algorithm to convert the audio signal into text data. In another implementation, the server also extracts acoustic features, such as fundamental frequency, energy, formant distribution, and speech rate, for emotion inference.
[0117] In the sentiment analysis module, the server uses multimodal features to estimate the sentiment of meeting participants. In one implementation, the server employs a combination of convolutional neural networks (CNNs) and recurrent neural networks (RNNs). The CNNs process the time-spectral features of the audio, while the RNNs process the time-series features of the text. The server concatenates or weights the intermediate vectors from both, outputting the sentiment category distribution and sentiment intensity score through one or more fully connected layers. The server represents the output as a sentiment information vector and records it in the meeting minutes.
[0118] When generating subsequent suggestions, the server incorporates emotional information as an additional condition into the prompts. For example, the server can add text describing the current meeting atmosphere to the prompts, such as, "The atmosphere in the meeting is rather tense, and some participants are quite sensitive to risks. Please appropriately emphasize risk control and robust strategies in your suggestions." In this way, the server enables the generative AI model to consider emotional constraints at the initial input level, thereby automatically adjusting the tone and risk level of the solutions when outputting suggestions, which helps improve the fit between the generated suggestions and the actual decision-making environment.
[0119] 5. Generative Artificial Intelligence Model Invocation and Internal Computation In one implementation, the server deploys the generative AI model on a remote inference service and invokes it via a network interface. In another implementation, the server deploys the model to run on a local GPU. In both implementations, the generative AI model can employ a Transformer-based neural network to encode and decode prompts. The server sets generation parameters during invocation, including temperature, maximum output length, and sampling strategy.
[0120] Instead of a simple single-turn question-and-answer interaction, the server compresses the meeting context into a high-information-density input by embedding summaries, keywords, and sentiment information within the prompts. Internally, the server utilizes a self-attention mechanism to assign higher attention weights to selected high-weight sentences and keywords in the summary. This allows the model to focus on generating relevant information, thereby reducing irrelevant content output.
[0121] The server uses a large-scale corpus and a specialized dataset of meeting decision-making texts when training or selecting generative AI models. During offline training, the server uses cross-entropy loss as the error function and employs a gradient descent-based weight update algorithm. In some implementations, the server performs data augmentation on the training data, such as synonym substitution, order perturbation, or summary expansion of the meeting texts, to improve the model's robustness to different expressions. These training methods enable the model to more accurately understand and reason about meeting scenario prompts at runtime, thereby improving the accuracy and consistency of generated opinions.
[0122] Before invoking the generative AI model, the server performs context pruning and vector caching management. For multi-turn conference dialogues, the server uses a sliding window strategy, retaining only the most relevant historical summaries and keywords to avoid increased inference overhead due to excessively long input lengths. The server locally caches recently used embedding vectors and reuses some intermediate results when adjacent requests have high similarity, thereby further reducing the computational burden and improving response speed.
[0123] 6. Feedback Information Post-processing and Structured Module After receiving the opinion text output by the model, the server performs post-processing using rules and statistical methods. The server segments the text into paragraphs and entries, categorizing the content into groups such as "strategic recommendations," "risk analysis," and "action plans" based on predefined subheading patterns or automatic clustering. The server stores these categories as key-value pairs, allowing the terminal to selectively display or collapse portions of the content as needed.
[0124] In one implementation, the server also performs consistency and redundancy checks on the generated text. For example, the server can calculate the similarity between paragraphs, and if multiple paragraphs are found to be highly repetitive, they are merged or deleted, thereby reducing the burden of redundant information on the user's reading experience. In another implementation, the server uses rule-based filters to remove parts containing inappropriate language, ensuring that the output conforms to preset specifications in terms of both technology and expression.
[0125] Through this structuring process, the server not only generates natural language text, but also forms structured data that can be used for subsequent retrieval, statistics, and automated reporting, enabling the system to reuse this information in subsequent meetings or cross-project analyses.
[0126] III. Interaction methods and technical effects between terminals and users In one implementation, the terminal presents a form-based interface, allowing users to input meeting topics, background descriptions, and key discussion points. In another implementation, the terminal provides real-time voice input functionality.
[0127] When using the terminal, users can directly enter the following text prompt as an additional question: "As a corporate strategy consultant, please provide a comprehensive opinion on the 'New Product Launch Strategy' based on the following meeting information."
[0128] Meeting Information: 1. We plan to launch a SaaS product for small and medium-sized enterprises (SMEs) within six months for sales lead management.
[0129] 2. The main competitors are Company A and Company B. Company A has an advantage in price, while Company B has more comprehensive functions.
[0130] 3. Our budget is limited, and we hope to control costs in the early stages of marketing.
[0131] Please provide specific suggestions regarding target market selection, pricing strategy, distribution channels, and potential risks. After receiving user input, the terminal submits the text along with information such as the summary generated by the server. Internally, the server integrates the user input into its prompt message building module, making the generated prompt messages more tailored to the user's specific needs.
[0132] When displaying structured feedback information returned by the server, the terminal presents each category of content as a collapsible list. In one implementation, the terminal also maps structured fields to charts or key point lists, allowing users to quickly browse crucial recommendations. In this way, the terminal not only acts as a passive display device but also works in conjunction with the server's structured output design to achieve coordinated optimization of interface interaction and data organization.
[0133] IV. Explanation of Technological Improvements and Causal Relationships Through the aforementioned modular design and specific data structures, the server implements a deep processing workflow within the computer that differs from traditional "record-playback" conferencing systems. By extracting summaries and keywords, the server compresses long text conference information into high-information-density summaries, reducing input length and directly decreasing the number of tokens that generative AI models need to process. This shortens inference time under the same hardware resources, resulting in an overall improvement in processing speed.
[0134] The server constructs prompts containing summary information, keywords, and sentiment information, enabling the generative AI model to obtain a highly structured context at the input stage, rather than directly facing a lengthy and loosely structured full text. This input organization method changes the distribution of attention within the model, allowing the algorithm to focus more on decision-related segments, reducing the waste of attention on redundant content. This technically improves the relevance and accuracy of the generated results and reduces the probability of generating irrelevant or off-topic content.
[0135] The server uses multimodal sentiment fusion to map audio and text features onto a unified sentiment vector space, explicitly reflecting the desired meeting atmosphere in prompts. This sentiment perception mechanism enables the system to automatically adjust risk profiles and wording strategies when generating suggestions, demonstrating unique processing rules that distinguish it from simple automated document systems, thus achieving dynamic adaptation to the meeting context.
[0136] The server performs structured post-processing on the generated text to create key-value data objects, enabling different outputs from the same meeting to be quickly retrieved, compared, and reused subsequently. Compared to traditional methods that can only be searched by full text, this structured organization significantly reduces the computational load during queries and provides standardized input for subsequent cross-meeting data mining, thereby improving overall data management efficiency and computational utilization.
[0137] By employing these specific algorithmic processes, data structure designs, and generative artificial intelligence models in synergy, this system transforms tasks involving significant human intervention in meeting recording and analysis into a reproducible and scalable technical processing chain that can be executed repeatedly within the server. This achieves a comprehensive improvement in computer processing power and model inference efficiency, rather than simply automating manual steps.
[0138] use Figure 11 The processing flow is explained.
[0139] Step 1: Users input meeting information using a terminal.
[0140] Users can enter meeting topics and content in the meeting application interface on the terminal, or describe the key points of the meeting discussion via voice. Input can include text descriptions (such as background information on "new product market launch strategy", competitor information, budget constraints, etc.) and additional notes.
[0141] The terminal takes raw text and / or voice data input by the user as input, performs basic format checks on the text (such as length and whether it is empty), and performs local encoding on the voice data (such as compression to a specific audio format). The output is a packaged meeting information request, which includes at least: meeting agenda text, raw meeting information text, and optional audio data.
[0142] Step 2: The terminal sends a meeting information request to the server.
[0143] After the user confirms the input, the terminal sends a meeting information request to the server via a network communication protocol. The input is the meeting information object containing text and audio generated in step 1.
[0144] The terminal serializes the object, organizing the text and tagging information into structured data, and appends the audio data in binary form. Before sending, the terminal adds a session identifier and a user identifier so the server can identify the source. The output is a conference information request message transmitted over the network to the server.
[0145] Step 3: The server receives and stores the original meeting information.
[0146] The server receives a conference information request from the terminal at the communication interface. The input is the request message received by the network layer, which contains text conference information, audio data, and relevant identifiers.
[0147] The server parses the request, extracting the text content, audio data, and session information, and performs data type conversion and validity checks (such as checking the existence of required fields). The server then inserts the original meeting information into the storage system, creating a new meeting record. The output is a meeting record object generated in storage, containing the original meeting information, and its unique identifier.
[0148] Step 4: The server preprocesses the text conference information.
[0149] The server reads the raw text information of the corresponding meeting record from storage. The input is the raw meeting information text field from the meeting record.
[0150] The server performs data cleaning operations on the text, including removing extra spaces, standardizing punctuation, deleting control characters and meaningless symbols, and segmenting sentences based on linguistic features. The server uses a word segmentation algorithm to divide the text into word sequences and performs basic filtering on the segmentation results (removing stop words, numerical noise, etc.). These operations are completed through string processing, regular expression matching, and simple statistical calculations. The output is the preprocessed text data and the corresponding word segmentation list, which the server writes back to the preprocessing field of the meeting minutes.
[0151] Step 5: The server generates summary information.
[0152] The server takes preprocessed text data as input and calls a text summarization algorithm. The input consists of cleaned meeting information text and its sentence segmentation results.
[0153] The server calculates sentence importance scores based on features such as sentence length, position, and similarity to other sentences, and then ranks the sentences using graph sorting or other statistical methods. During this process, the server performs data operations through vectorization and similarity calculations, transforming sentences into vectors and constructing a similarity matrix. After obtaining the ranking results, the server selects the top-scoring sentences as a summary. The output is a short summary text, which the server saves to the summary field of the meeting minutes.
[0154] Step 6: The server extracts important keyword information.
[0155] The server takes the preprocessed word segmentation list as input and performs a keyword extraction process. The input consists of a sequence of words and their frequencies in the text.
[0156] The server first counts the frequency of each word and calculates its weight in the current meeting information, for example, using TF-IDF or word vector-based weights. The server filters candidate words, removing those that are too short or lack meaningful semantics, and then selects a number of high-weight words as important words based on their weights. This process involves frequency statistics, weight calculation, and sorting operations. The output is a list of important words, which the server records in the keyword field of the meeting minutes.
[0157] Step 7: The server extracts text data from the audio data.
[0158] The server takes the stored audio data as input and calls the speech recognition processing module. The input is either a stream of audio data or an audio file collected during the meeting.
[0159] The server uses a speech recognition algorithm to convert audio signals into feature sequences, extracts acoustic features, matches them to a language model, and gradually outputs corresponding text segments. The server then concatenates all segments and performs basic cleaning (such as removing noise markers) to form complete text data. The output is text-based conference information converted from audio; the server incorporates this into the text-based conference information and updates the preprocessed text fields.
[0160] Step 8: The server performs sentiment analysis on the meeting information.
[0161] The server takes text-based meeting information and optional audio features as input and performs sentiment inference. The input includes preprocessed meeting information text and acoustic features extracted from the audio (such as fundamental frequency, energy, speech rate, etc.).
[0162] The server encodes the text using a sentiment classification model to obtain a text sentiment vector; simultaneously, it inputs acoustic features into an acoustic sentiment model to obtain an audio sentiment vector. The server performs a weighted fusion operation on these two vectors to obtain a comprehensive sentiment vector, and outputs the sentiment category and intensity through a classification or regression layer. The output is sentiment information representing the emotional state of the meeting participants, which the server writes into the sentiment field of the meeting record.
[0163] Step 9: The server builds prompt statements for generative artificial intelligence models.
[0164] The server takes meeting agenda information, summary information, key words information, and sentiment information as input and generates prompts. The input consists of the agenda text, summary text, keyword list, and sentiment information from the meeting minutes.
[0165] The server selects a prompt template corresponding to the meeting scenario, fills the summary text into the summary position, fills the keyword list into the keyword position using separators, and adds emotional states as needed, either in textual form or as constraints. The server constructs the prompt statement through string concatenation and placeholder replacement. The output is a complete prompt statement text, which the server saves in the prompt statement field of the meeting record.
[0166] Step 10: The server calls a generative artificial intelligence model to generate opinion information.
[0167] The server takes the pre-constructed prompt as input and sends it to the generative artificial intelligence model. The input consists of the prompt text and model call parameters (such as maximum generation length, temperature, etc.).
[0168] The server takes the prompts as the input sequence to the model. Internally, the model performs vectorization and attention calculations through an encoder-decoder structure or a decoder structure to gradually generate opinion text. The server receives the text sequence output by the model, parses its format, and combines consecutive tokens into natural language sentences. The output is the original opinion information text generated by the model, which the server records in the opinion information field of the meeting minutes.
[0169] Step 11: The server performs post-processing and structuring on the generated opinion information.
[0170] The server takes the generated opinion information text as input and processes and structures it. The input is the original opinion information text.
[0171] The server segments the text using paragraph separators, bullet points, or numbering to identify the main content of different sections. Based on preset rules or key phrases, the server categorizes these sections into groups such as "strategic recommendations," "risk analysis," and "action plans." The server then organizes these categories and corresponding text into structured data objects, such as fields or lists. The output is structured opinion information, which the server writes into the structured fields of the meeting minutes.
[0172] Step 12: The server sends structured opinion information to the terminal.
[0173] The server takes structured opinion information and relevant metadata as input and generates a response message. The input includes structured opinion fields from the meeting minutes, the original opinion text, and the meeting identifier, etc.
[0174] The server packages this data into a response, including the complete comment text and structured data categorized by type. The server then sends the response to the corresponding terminal via a network protocol. The output is a comment information response message sent to the terminal, containing all the necessary data for display and subsequent interaction.
[0175] Step 13: The terminal displays feedback information and supports subsequent user interactions.
[0176] The terminal takes the structured feedback information returned by the server as input, parses it, and displays it. The input consists of response data containing feedback text, structured fields, and meeting identifiers.
[0177] The terminal displays the overall opinion text as the main content, while mapping structured data to different blocks on the interface, such as presented in columns like "Strategic Recommendations," "Risk Analysis," and "Action Plan." The terminal allows users to expand or collapse different sections of content and supports selecting specific sections for copying, annotation, or saving. When a user needs to add a question, the terminal resends the newly entered text along with the original meeting identifier to the server, forming a new meeting information request. The output consists of opinion information presented intuitively on the terminal interface, as well as user input data that can be used for the next round of requests.
[0178] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0179] In modern production management and meeting decision-making scenarios, widely used computer-aided systems typically handle only a single task, such as statistical analysis of production line operating data or recording and transcribing meeting audio. These systems have the following limitations at the computer technology level: (1) Traditional production management systems mostly use preset rules or simple algorithms, only displaying or alarming the collected sensor data. They lack the ability to construct high-dimensional features and perform deep modeling on large-scale time-series running data, and cannot automatically generate the optimal production plan under complex constraints at the computation process level. This results in the underutilization of computing resources and rigid algorithm paths.
[0180] (2) Existing conference support systems are mostly based on speech recognition, which converts audio into text and then performs simple keyword extraction. They lack the technical path to encode unstructured features such as emotional information and structured business data in a unified manner and input them into the same artificial intelligence model. This makes it impossible for computers to process "human decision intentions" and "production system status" simultaneously in a single model reasoning process, resulting in a broken overall system reasoning link and low model calling efficiency.
[0181] (3) In many systems, the way artificial intelligence models are called is through fixed interface calls. The input parameters and instruction texts are manually written or statically configured. There is a lack of a mechanism to automatically generate prompts based on real-time running data and meeting context and dynamically construct the model input vector. This makes it difficult for the model's reasoning ability to adapt to different scenarios at the computer level, affecting the model output quality and the efficiency of computing resource utilization.
[0182] (4) Existing production control systems generally separate “plan generation” and “equipment instruction issuance” into different software modules. They lack an integrated computing process management from raw sensor data to generative artificial intelligence model reasoning, and then to automatic generation and feedback monitoring of control instructions. This results in multiple manual interventions and multiple data transfers, leading to redundant calculations, repeated storage, and increased system complexity and delay.
[0183] Therefore, there is an urgent need for a technical solution that, at the computer technology level, unifies the preprocessing of multi-source data such as production operation data, conference audio, and emotional information; automatically generates prompts and input features for calling generative artificial intelligence models based on this unified encoded data; simultaneously outputs conference opinions and candidate production plans in a single inference process; and automatically feeds back user interaction results and control instructions to the production object, thereby improving the data processing flow, model calling method, and control link, and thus enhancing overall computing efficiency, resource utilization, and system responsiveness.
[0184] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0185] In this invention, the server includes: a device for automatically generating prompts for instructing a generative artificial intelligence model to perform inference based on meeting topics and the operating status of production objects; a device for collecting operating data from an information acquisition device and converting the operating data into standardized time series data and constructing model input features; a device for converting voice data and / or image data related to meeting participants into text data and emotional information, and uniformly encoding them with the prompts and the input features before inputting them into the generative artificial intelligence model; a device for simultaneously outputting opinions on meeting topics and candidate production plans including production line time allocation, target objects, target quantities, and resource allocation based on a single generation process of the generative artificial intelligence model; a device for sending candidate production plans to an information prompting device and receiving user correction operations to generate a final production plan; and a device for automatically generating executable instruction information for a control device based on the final production plan and controlling the operation of production objects through the control device. This enables the formation of an integrated data processing and control chain within the computer, encompassing the acquisition of multi-source operational and conference data, the construction of unified features and the automatic generation of prompt statements, and the automatic issuance of generative artificial intelligence model inference and control commands. This reduces manual configuration and data transfer between multiple modules, enhances the adaptability and inference quality of model calls, improves computational efficiency and resource utilization in the production plan generation and execution process, and achieves substantial improvements in data processing workflows and intelligent control mechanisms in the field of computer technology.
[0186] A "system" refers to a computer implementation that consists of multiple functional devices, processing units, and storage units interconnected through communication means to achieve overall functions such as data acquisition, data processing, model reasoning, and control command generation and issuance.
[0187] A "server" is an information processing device equipped with a processor, memory, and communication interface, used to execute program instructions and provide data processing and control services to the outside world.
[0188] "Production objects" refer to the controlled objects that are monitored and controlled in the production environment, including industrial operating entities such as production lines, processing equipment, conveying devices and related facilities.
[0189] "Information acquisition device" refers to a device configured on or around the production object to collect operational status information, including sensors, controllers, monitoring terminals and their interface modules.
[0190] "Operational data" refers to various types of data collected by information acquisition devices that reflect the operating status of the production object, including time information, equipment status, output, speed, energy consumption, temperature, pressure, and material status.
[0191] "Standardized time series data" refers to a multidimensional data set that is suitable as input for time series data, formed by arranging the original operational data in chronological order and processing it through cleaning, alignment, normalization, or standardization.
[0192] "Input features" refer to numerical feature data obtained by feature extraction and encoding from running data and / or other relevant information, which are used to input into generative artificial intelligence models for inference operations.
[0193] "Generative artificial intelligence models" refer to artificial intelligence models that can automatically generate new structured or unstructured data (such as text or plans) based on input features and prompts. These models include those built using algorithms such as deep learning, probabilistic generation, sequence modeling, or reinforcement learning.
[0194] "Prompt statements" refer to textual information used to instruct generative artificial intelligence models on what type of generative task, constraints, and objectives to perform, serving as instruction inputs or control signals during model inference.
[0195] "Meeting topics" refer to the themes discussed and decided upon in a meeting setting, including production strategies, resource allocation, plan adjustments, and risk responses.
[0196] "Meeting participants" refers to human subjects who actually attend the meeting and speak, discuss, or make decisions, including managers, technicians, and operators.
[0197] "Voice data" refers to audio signals acquired by audio acquisition equipment that reflect the content and voice characteristics of the participants' speeches.
[0198] "Image data" refers to image or video frame data acquired by image acquisition equipment that reflects the facial expressions, postures, or meeting environment of meeting participants.
[0199] “Text data” refers to text information represented in the form of characters or strings, which is obtained by processing speech data through speech recognition or by extracting text from image data.
[0200] "Emotional information" refers to qualitative or quantitative information that indicates the emotional state or attitude of meeting participants, obtained from the analysis of voice data, image data, and / or text data.
[0201] "Generative processing" refers to the computational process by which a generative artificial intelligence model performs inference operations based on input features and prompts, thereby outputting opinions, plans, or other generated results.
[0202] "Candidate production plan" refers to one or more production plan proposals generated by a generative artificial intelligence model that have not yet been finally confirmed by the user, including the time arrangement, target objects, quantity indicators and resource allocation information for each production line.
[0203] "Information prompting device" refers to a device used to receive information sent by a server and display or output it to the user, including terminal equipment, display equipment and human-computer interaction interface.
[0204] "Correction operation" refers to the interactive operations performed by users on candidate production plans, such as adding, deleting, modifying, and adjusting the plan content through information prompt devices.
[0205] "Final production plan" refers to a deterministic production plan scheme that is based on candidate production plans, modified by the user, and confirmed by the system, and serves as the basis for generating subsequent control instructions.
[0206] "Control device" refers to a device used to receive instruction information from the server and control the production object, including programmable controllers, actuator control units and their interface modules.
[0207] "Instruction information" refers to a set of control instructions generated by the server based on the final production plan, which can be recognized and executed by the controllable device to drive the production object to operate in a predetermined manner.
[0208] "External information providing device" refers to an information processing device or platform located outside the system, used to receive opinions, production plans and evaluation information output by the system and to display, publish or further process them.
[0209] "Evaluation information" refers to performance indicators, effect assessments, comparative analyses, or statistical results related to the generated production plan and its execution results, which are used to reflect the quality of the plan and its execution status.
[0210] In a preferred embodiment of the present invention, the server serves as the core information processing device, the terminal serves as the human-computer interaction device, and the user serves as the decision-making participant. The three work together to realize a generative artificial intelligence processing flow that integrates meeting decision-making and production control.
[0211] In terms of hardware, servers can employ computer devices equipped with multi-core central processing units and graphics processing units, such as general-purpose server hardware platforms with graphics processing units capable of parallel computing. In terms of software, they can run general-purpose operating systems (such as server operating systems based on UNIX-like architectures) and install interpreted execution environments (such as the Python runtime environment), deep learning frameworks (such as TensorFlow), and relational database management systems (such as general-purpose relational databases). Furthermore, servers can use data analysis libraries (such as Pandas and NumPy) to process structured and time-series data and extract features.
[0212] The server works collaboratively through multiple functional modules, which may include: a data acquisition module, a feature construction module, a prompt statement generation module, a generative artificial intelligence model invocation module, a candidate production plan generation module, a user interaction management module, and a control command generation module. These modules can exchange data through shared data structures in memory and persistent tables in the database.
[0213] In terms of data acquisition and storage, the server can connect to information acquisition devices deployed on production objects via industrial communication interfaces. These devices can include various sensors and programmable control units. The server periodically reads operational data, such as equipment status, current output, production cycle time, downtime, energy consumption, temperature, pressure, and material inventory, using industrial communication protocols. The server writes the collected raw data into a time-series data table in the database, using timestamps as the primary key, and indexes the timestamp and production line identifier fields during the write process to improve subsequent query efficiency.
[0214] In terms of feature construction, the server uses a data analysis software library to preprocess the raw operational data. The server performs missing value imputation, outlier removal, unit standardization, and time alignment operations on the raw data to generate standardized time series data. The server can calculate sliding window statistical characteristics for each production line, such as average output, variance, maximum load, and downtime frequency within a fixed time window. The server can then correlate these statistics with order information, shift plans, and other data to construct multidimensional input features. The server organizes the constructed features into multidimensional arrays, for example, using a tensor structure of "batch size × time step × feature dimension," and stores them in memory in numerical form for direct reading by the generative artificial intelligence model's calling module.
[0215] In terms of prompt generation, the server automatically generates corresponding text instructions based on the current meeting topic and the operational status of the production object. The server can read keywords for the current topic from the meeting topic management table, such as "capacity balance," "delivery guarantee," and "energy consumption optimization," and combine these keywords with operational data analysis results to construct natural language prompts. For example, the server can generate the following Chinese prompt: "You are a production planning optimization expert. Based on the current real-time load and historical output data of each production line, as well as the following order requirements, please generate the optimal production plan for the next 24 hours, with the following priorities: 1) Meet order delivery dates; 2) Maximize equipment utilization; 3) Reduce product changeover times. Please output the product models and planned output for each production line in each time period, along with a brief text description." The server can also generate the following prompt when a user wants to adjust the night shift load: "While ensuring all orders are delivered on time, please regenerate the production plan for the next 12 hours, keeping the equipment utilization rate for the night shift below 75%, and try to schedule high-intensity tasks for the day and afternoon shifts. Please explain the main reasons for the adjustments." In constructing and calling generative artificial intelligence models, the server can use deep learning frameworks to implement multi-layer neural network models. The server can employ a sequence generation model with an encoder-decoder structure. The encoder can include several transformation layers with multi-head attention mechanisms to jointly encode time-series features and prompt embedding vectors. The decoder can generate production plan sequences and corresponding text opinions. The server encodes the terms in the prompts into vectors through embedding layers, and then captures the dependencies between different constraints in the text through positional encoding and multi-head attention mechanisms. Simultaneously, the server maps standardized time-series data to a feature space and fuses it with the text embeddings in a unified latent space.
[0216] During the model training phase, the server can collect historical production plan data, execution result data, and meeting minutes data in advance to construct training samples. Based on supervised learning methods, the server updates the model weights by minimizing the plan generation error and an auxiliary loss function. The server can define the loss function as a multinomial sum, including the difference between the production plan and the historical best solution, a penalty term for resource constraint violation, and a penalty term for delivery delay. The server iteratively updates the weights using stochastic gradient descent-like optimization algorithms or their variants. During training, the server can employ data augmentation strategies, such as adding noise perturbation to a portion of the input time series data, to improve the model's robustness to sensor noise and incomplete data.
[0217] During the inference phase, the server uses a pre-trained generative AI model. The server feeds input features generated from current running data, along with prompts, into the model. The model output includes a sequence of production tasks across multiple time segments, along with information for each task such as the production line number, start time, end time, product category, target output, and estimated resource consumption. The server parses the model output into structured candidate production plans and stores them in an intermediate data structure. Simultaneously, it generates text-based meeting comments and explanations, such as why a particular production line is scheduled to perform a specific product task at a specific time, and how resource conflicts are automatically avoided.
[0218] In this invention, the terminal serves as an information prompting device. It can be a general-purpose computing device with a display screen and input unit, such as a tablet, laptop, or industrial all-in-one machine. The terminal can run a graphical user interface program and retrieve candidate production plans and generate opinions from the server via a network communication interface. The terminal can display the candidate plans graphically, for example, using a timeline chart or Gantt chart to show the task allocation for each production line, while simultaneously displaying the opinions and rationale output by the generative artificial intelligence model in text form. The terminal can provide visual editing functions, allowing users to modify fields such as task time, production line allocation, and target output through drag-and-drop, clicking, and form input.
[0219] Users can view candidate production plans and related comments on the terminal, and can make adjustments based on their own experience or additional information. If a user finds that the night shift workload in the AI-generated plan is too high, they can manually adjust some tasks to the day or afternoon shift, or enter a new prompt through the interface to request the server to regenerate a plan that meets additional constraints. For example, a user can enter the following prompt: "Please minimize the production load of the night shift as much as possible without affecting the timely delivery of key customer orders, and concentrate maintenance tasks during periods of lower equipment utilization." After the user completes the correction, the terminal packages the modification results and the new prompt statement into request data and sends it to the server. The server updates the input feature values and prompt statement accordingly, and calls the generative artificial intelligence model again to perform reasoning and generate a new candidate production plan.
[0220] After the final plan is confirmed, the server generates instruction information for the control device based on the final production plan. The server can map each production task to a set of control instructions that the control device can recognize, such as setting target output parameters, production formula numbers, start and stop times, etc. The server writes the control instructions into specific registers or interfaces of the control device through industrial communication protocols, and the control device then drives the actual actuators to perform operations such as starting, stopping, speed adjustment, and product switching on the production object. In this way, the system of the present invention directly transforms the output of the generative artificial intelligence model into specific physical control behaviors, realizing a closed loop from data analysis to equipment control.
[0221] The improvements of this invention at the computer technology level are reflected in the following aspects.
[0222] The server encodes multi-source data (including sensor time-series data, order structured data, conference voice and text, and sentiment information) into unified high-dimensional feature vectors through a unified feature construction module. These features are then jointly processed within the same generative artificial intelligence model, reducing the overhead of multiple format conversions and redundant calculations between different modules in traditional systems. Because features are processed within a unified numerical space, the server can use vectorized operations for parallel computation on the graphics processing unit, significantly improving processing speed.
[0223] The server automatically generates prompts, mapping real-time operational status and meeting topics into natural language commands. This avoids fixed rule configurations or manually written call parameters, allowing model call parameters to dynamically adjust with time and context. Because the prompts contain specific optimization goals and constraints, generative AI models can adaptively optimize for the current scenario during inference, thereby improving the accuracy and usability of the output solution. This method of controlling model behavior based on prompts manifests internally as dynamic adjustment of the model's input layer data distribution, improving the effective search space for model inference and reducing ineffective inference computations.
[0224] The server employs a neural network architecture with an attention mechanism. This not only models the correlations between distant data points in the time series but also weights the relationships between different constraints and production features in the prompts. This allows the model to focus on the data components that significantly impact the objective function during inference. Compared to traditional schemes based on fixed weights or simple linear models, this attention mechanism reduces unnecessary redundant computations through dynamic weight allocation, improves the signal-to-noise ratio of useful features, and thus enhances inference accuracy and reduces errors.
[0225] When training a generative AI model, the server introduces a multi-objective loss function, enabling the model to simultaneously consider multiple metrics such as production plan deviations, resource constraint violations, and time delays during the learning process. The server assigns weight coefficients to different loss components during weight updates, thereby searching the parameter space for a solution that optimizes overall system performance. Compared to algorithms that optimize only a single metric, this multi-objective training method implements a more complex error propagation path within the computer, which is beneficial for obtaining a model with balanced performance under various constraints and reduces the number of subsequent manual corrections.
[0226] In terms of data management, the server employs dedicated indexes and partitioning strategies for time-series data tables, enabling rapid retrieval of historical and real-time data within the required time window during plan generation, significantly reducing database query latency. The server can also utilize a caching mechanism to store recently used feature vectors and intermediate calculation results in high-speed storage, allowing for reuse when the generative AI model is called multiple times in a short period, thereby reducing redundant preprocessing operations and improving overall computational efficiency.
[0227] Because this invention integrates meeting decisions and production control into a unified data flow and model inference chain, the server no longer needs to manually transmit decision results through multiple separate software systems. Instead, it directly drives production objects through standardized data structures and a unified instruction generation module. This integrated design eliminates multiple manual inputs and operations in traditional processes, reduces human errors and delays in data transmission and conversion, and is executed automatically within the computer, contributing to high-precision and high-speed production control.
[0228] The terminal presents candidate plans and AI opinions through a visual interface and provides interactive editing functions, allowing users to intervene to a limited extent while maintaining the advantages of the system's automatic generation. In this invention, the user's corrections are accurately recorded by the terminal and uploaded to the server. The server can use this correction data for subsequent model retraining or fine-tuning, enabling the generative AI model to gradually adapt to the preferences of specific factories and users. This online improvement mechanism based on actual operational feedback technically allows the model parameters to continuously approach the optimal solution in the real operating environment, thereby improving long-term prediction accuracy and resource utilization.
[0229] In summary, this invention, through the collaboration of servers, terminals, and users, and the organic combination of generative artificial intelligence models and prompt statements, realizes an integrated technical solution within a computer, encompassing multi-source data acquisition, feature construction, prompt statement generation, model inference, control command generation, and device control. This not only achieves specific physical control of production objects but also improves upon traditional computer processing models in terms of processing speed, inference accuracy, data management, and computational efficiency. Thus, it provides an implementation model that goes beyond simply automating business processes, but substantially improves computer technology itself.
[0230] use Figure 12 The processing flow is explained.
[0231] Step 1: The server collects operational data from information acquisition devices on the production objects. Inputs include raw signal data from sensors and control devices (e.g., voltage values, count values, status bits, etc.) along with corresponding timestamps and production line identifiers. The server reads these signals via industrial communication protocols, parses discrete register values into physically meaningful numerical values (e.g., output, temperature, pressure, equipment status, material inventory), and appends a uniformly formatted timestamp to each record. The server performs preliminary filtering on the parsed data, removing obvious errors, and outputs the results as structured raw operational data to memory and the database.
[0232] Step 2: The server preprocesses and standardizes the structured raw operational data to generate standardized time series data. The input is the structured raw operational data output from step 1. The server uses a data analysis software library to impute missing values, detect outliers (e.g., based on set thresholds or statistical ranges), and unify units, mapping data from different sampling frequencies to a unified time axis using a time alignment algorithm. The server then segments the data according to set time windows (e.g., every 5 minutes or every 15 minutes), and calculates statistics such as mean, maximum, minimum, and variance for the data within each window, normalizing or standardizing the numerical features, and finally outputting standardized time series data organized as "sample number × time step × feature dimension".
[0233] Step 3: The server constructs the input features for the generative AI model. The input consists of the standardized time-series data output from step 2, as well as structured business data related to orders and shifts. The server associates data such as order delivery dates, order quantities, product categories, and priorities with the time-series operational data of each production line, merging them into a unified feature space using key-value pairs (such as production line identifiers and time intervals). The server generates numerical codes or embedding vectors for each type of discrete attribute, maintains continuous attributes as floating-point features, and concatenates all features into a high-dimensional feature vector. The server performs alignment and sorting on different feature dimensions, forming an input feature tensor that can be directly input into the generative AI model and then outputting it.
[0234] Step 4: The server collects and processes meeting-related data, including audio and image data. Inputs are audio and video streams obtained through the meeting capture device. The server uses speech recognition algorithms to convert the audio signals into text data, extracts speech features such as pitch, energy, and speech rate using feature extraction algorithms, and infers sentiment labels or sentiment intensity values using sentiment analysis models. Simultaneously, the server extracts facial expression features and head posture features from image or video frames, generating supplementary sentiment information using image sentiment recognition algorithms. The server combines the text data obtained from speech recognition with the sentiment information, outputting meeting feature data that includes the meeting text content and corresponding sentiment annotations.
[0235] Step 5: The server generates prompts based on the meeting agenda and operational status. Inputs include meeting agenda information, the input feature overview from step 3, and the meeting feature data from step 4. The server first reads the current meeting topic (e.g., "Production scheduling optimization for the next 24 hours") from the meeting agenda management data and analyzes the current operational status overview (e.g., capacity utilization distribution, urgency of key orders, equipment failure status). The server combines this with the emotional information of meeting participants (e.g., sensitivity to risk, preference for a particular solution) and generates natural language prompts according to predefined templates or rules. The server then combines agenda keywords, optimization objectives, constraints, and explanatory text into complete sentences, outputting prompt text instructing the generative AI model on its tasks.
[0236] Step 6: The server encodes the input features, prompts, and meeting features into the model input. The input consists of the input features from step 3, the meeting feature data from step 4, and the prompt text from step 5. The server segments the prompts into words or characters using a text encoding module, mapping each word to an embedding vector, and adds positional encoding to the text sequence. Simultaneously, the server encodes the meeting text content into another type of embedding, and sentiment information is concatenated as numerical features into the corresponding time step or global vector. The server aligns and concatenates the above text embeddings and sentiment features with the input features along the feature dimension or time dimension to form a unified multimodal input tensor, which is then output as the final input data structure of the generative artificial intelligence model.
[0237] Step 7: The server invokes a generative AI model to perform reasoning, generating candidate production plans and meeting opinions. The input is the multimodal input tensor output from step 6. The server feeds this tensor into the trained generative AI model, which includes a multi-layer attention network and a decoder structure. The server performs forward propagation computation on the graphics processing unit, calculating the attention weights, linear transformations, and non-linear activations of each layer, ultimately obtaining the probability distribution of the production task sequence and text opinions. The server uses a sampling or greedy decoding strategy to generate specific production task sequences (including production line number, start and end times, product category, target output, resource requirements, etc.) and meeting opinion texts (e.g., explanations of the rationality of the plan and risk points) from the output distribution. The server organizes these results into candidate production plans and meeting opinion outputs.
[0238] Step 8: The server sends candidate production plans and meeting comments to the terminal. The input is the candidate production plan and meeting comment data structure generated in step 7. The server encodes the plan data into a structured data format (e.g., fielded data objects) via the application interface and sends it to the terminal, simultaneously sending the meeting comments in text format. Before sending, the server compresses or batch-packages the data to reduce communication load and appends a version number and timestamp to the communication protocol. The server output includes a task list for each production line, corresponding time interval, product category, target output, resource allocation suggestions, and comment explanations.
[0239] Step 9: The terminal receives and displays candidate production plans and meeting comments. The input is the data sent by the server in step 8. The terminal parses the structured data, mapping production tasks to visualization components (such as timeline views and list views), and presents them on the screen in graphical and textual form. The terminal draws bar elements of different colors and lengths based on task time intervals and production line attributes, expanding with detailed information when the user touches or clicks. Simultaneously, the terminal displays the text of meeting comments on the side of the interface or in a pop-up area to help the user understand the rationale behind the generative AI model. The terminal output is a visual display of the status on the user interface and an interactive interface that responds to user actions.
[0240] Step 10: Users review and revise candidate production plans on the terminal. Input consists of the candidate production plans and meeting comments shown in step 9. Based on their experience and on-site conditions, users perform specific operations through the interface components provided on the terminal, such as dragging task bars to change task start and end times, moving tasks from one production line to another, modifying target output values, or manually inserting maintenance tasks. Users can also enter new prompts in text input boxes, instructing the server to consider new constraints (such as limiting night shift load or prioritizing specific orders) when regenerating the plan. All user edits and inputs constitute the revision operation, and the output is the revised production plan data and the newly added prompt text.
[0241] Step 11: The terminal sends the revised production plan and new prompt statement to the server. The input is the result of the correction operation completed by the user in step 10. The terminal converts all modification records into difference data or complete plan data structure, packages them together with the new prompt statement into a request message, and sends it to the server through the network interface. The terminal can perform basic data validation, such as checking time format and numerical range, to ensure the integrity and legality of the sent data. The terminal output is a request data packet containing the revised plan and prompt statement.
[0242] Step 12: The server receives the revised plan and prompts, updates the plan, and optionally re-invokes the generative AI model. The input is the request data packet sent in step 11. The server first updates the candidate production plan records in the database, marking the revised plan as the user's adjustment scheme. Based on the new prompts and the latest operating data, the server reconstructs the input features if necessary, and again invokes the generative AI model to perform inference and generate new candidate production plans. The server compares the original plan with the revised plan, calculates and evaluates key differences (e.g., changes in load balance, changes in delivery time satisfaction), and outputs the finally confirmed plan as the final production plan.
[0243] Step 13: The server generates executable instructions for the control devices based on the final production plan. The input is the final production plan data output from step 12. The server maps each production task to specific control parameters, including start time, stop time, target output, equipment operating mode, and recipe number. The server encodes these parameters into registers and writes them into instructions or command frames according to the control device's protocol specifications, forming an instruction information sequence. The server sorts and timestamps the instruction information to ensure the control devices execute in the correct timing, and the final output is a set of control instructions for each control device.
[0244] Step 14: The server issues commands to the control devices and monitors execution feedback. The input is the set of control commands generated in step 13. The server sends the commands to each control device through the industrial communication interface, triggering the production objects to perform corresponding physical operations. The server then continuously collects operational data, compares the actual execution with the final production plan, and calculates deviation indicators (such as the ratio of actual output to target output, the difference between actual start time and planned start time, etc.). When the server detects that the deviation exceeds a preset threshold, it can record the event, generate alarm information, and reconstruct the input feature quantities and prompt statements as needed, triggering generative artificial intelligence model reasoning for local adjustments to form new adjustment suggestions. The server output includes the equipment execution status, deviation information, and possible correction suggestions to maintain the consistency between the production process and the production plan.
[0245] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0246] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0247] In the field of conference support, existing technologies typically only provide simple recording functions or speech recognition-based text transcription functions. Servers mostly just convert speech data into text data for users to read later. This approach has the following technical shortcomings: First, servers lack automated processing capabilities for long-term conference data, failing to establish a low-latency, continuously operating data pipeline between audio acquisition, speech recognition, and intelligent analysis, resulting in the inability to provide usable analysis results in a timely manner during the conference. Second, servers usually only linearly store the text data obtained through speech recognition, without structuring it according to semantic features such as the conference topic and decision focus. This leads to low processing efficiency and unstable context for subsequent generative AI models, thereby reducing the relevance and interpretability of the model's output. Third, when providing input to generative AI models, servers often simply send the raw text content to the model, lacking a mechanism to dynamically construct prompts based on the conference agenda, and cannot address specific issues such as "cut-off" prompts. Fourth, after receiving the model output, the server generally returns it to the terminal only in plain text form, without structuring or partitioning the response results. This makes it impossible to effectively distinguish between different information dimensions such as "meeting key points," "time setting suggestions," and "risk and resource suggestions," thereby increasing the user's understanding cost and hindering rapid decision-making assistance within a limited time. Fifth, the server usually does not link the sentiment analysis results with the input control or output structuring process of the generative artificial intelligence model, causing the system to be unable to dynamically adjust the expression and focus of suggestions according to the emotional state of meeting participants, thus reducing the system's adaptability to complex meeting atmospheres.
[0248] Therefore, it is necessary to propose an improved computer implementation scheme that enables the server to form a collaboratively optimized data processing pipeline across multiple stages, including audio acquisition, speech recognition, text processing, prompt statement construction, generative artificial intelligence model invocation, and structured output of results. This would enhance the automation, real-time performance, relevance, and usability of the meeting analysis and decision support process, thereby improving the overall performance of the meeting data processing system at the computer technology level.
[0249] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0250] In this invention, the server includes: a module for receiving conference voice information acquired by a terminal via a voice input function and transmitted through a communication path, converting the conference voice information into text information using a speech recognition function, and integrating it into conference record information in chronological or speaking order; a module for automatically extracting segments related to predetermined topics or decision-making objects from the conference record information, dynamically generating prompt statements containing instruction text according to the topics, and combining the prompt statements with the segments to form input information for input to a generative artificial intelligence model; a module for inputting the input information into the generative artificial intelligence model to perform semantic parsing processing, generating response information indicating meeting opinions or proposals based on the parsing results, and performing structured processing on the response information, dividing it into meeting key point information, time setting proposal information, and suggestion information related to risks or resources; and a module for performing sentiment estimation processing on the conference record information when necessary to obtain sentiment information, and using the sentiment information as a control condition to adjust the input composition of the generative artificial intelligence model or the structuring method of the response information. This allows for the formation of an end-to-end data processing pipeline on the server side, encompassing voice acquisition and access, text data generation and organization, generative AI model invocation based on prompts, and structured result output. This enhances the automation and real-time performance of meeting data processing, strengthens the relevance and readability of model outputs to specific decision-making issues, and leverages emotional information to improve the system's adaptability to the meeting atmosphere, thereby improving the overall performance of the meeting support system from a computer technology perspective.
[0251] "System" refers to an integral set of hardware and software units that work together organically to perform conference-related data acquisition, processing and output, consisting of at least one information processing device, at least one terminal and communication infrastructure for exchanging data between the information processing device and the terminal.
[0252] "Terminal" refers to an information processing device that is directly operated by the user at the meeting site and has sound input and display functions, including but not limited to portable computing devices, fixed computing devices, or other electronic devices with audio acquisition and communication capabilities.
[0253] "Voice input function" refers to the ability of a terminal to convert analog sound signals in the environment into audio data that can be processed by digital circuits using acoustic sensors and their accompanying drivers.
[0254] "Recording function" refers to the ability of a terminal to store, cache, or temporarily store the collected audio data in a predetermined format for subsequent transmission or processing.
[0255] "Conference voice information" refers to the collection of audio content, such as voices, dialogues, and discussions, made by participants during a meeting, which is collected and represented as digital audio data by the terminal's audio input function.
[0256] "Communication path" refers to the wired or wireless data transmission channel used to transmit conference voice information, text information and control information between the terminal and the information processing device, including network links, communication protocol stacks and their management mechanisms.
[0257] "Information processing device" refers to a computing device that has a processor, memory and network interface and is capable of executing programs to process, store and forward received data, including server equipment or computing infrastructure with the same functions.
[0258] "Speech recognition function" refers to a software algorithm or service interface that runs on an information processing device and can accept audio data as input and output corresponding text information, including speech-to-text processing capabilities based on statistical models or deep learning models.
[0259] “Text information” refers to text data that is generated by processing conference audio information using speech recognition functions, and is represented by a sequence of symbols.
[0260] "Chronological order" refers to the order in which the corresponding text information is arranged according to the chronological order in which the audio information occurs during the actual meeting.
[0261] "Speaking order" refers to the order in which corresponding text information is arranged according to the sequence of different speaking actions. It does not have to be strictly equivalent to the absolute time order, but it reflects the rounds of speaking.
[0262] "Meeting minutes information" refers to a collection of text data that is formed by integrating multiple text information fragments in chronological or speaking order, and can reflect the overall process and content of the meeting.
[0263] "Agenda" refers to topics that are pre-set or formed during a meeting and require discussion or decision-making, including but not limited to project progress, schedule, resource allocation, and risk response.
[0264] "Decision-making matters" refer to specific issues or tasks that require discussion at meetings to arrive at clear solutions or conclusions.
[0265] A “fragment” refers to a portion of textual information extracted from meeting minutes according to predetermined rules, which is semantically relatively complete or related to a particular topic.
[0266] "Instruction text" refers to natural language text content used to explain control information such as task requirements, output format, and focus to generative artificial intelligence models.
[0267] "Prompt statements" refer to input content consisting of instruction text and optional contextual information, which are natural language expressions used to guide generative artificial intelligence models to process subsequent data according to the expected roles, tasks, and output structures.
[0268] "Generative artificial intelligence models" refer to artificial intelligence models that can generate new natural language text or other forms of output based on input text or other structured data through internal parameterized models, including but not limited to language generation models based on deep neural networks.
[0269] "Input information" refers to the complete input data that the server combines fragments of meeting record information with corresponding prompts, and submits to the generative artificial intelligence model for semantic parsing processing.
[0270] "Semantic parsing processing" refers to the internal computational process by which generative artificial intelligence models perform semantic understanding, information extraction, relational analysis, and reasoning on input information, in order to extract useful information from text and form a structured internal representation.
[0271] "Response information" refers to natural language text data output by a generative artificial intelligence model after semantic parsing processing, used to indicate meeting opinions, suggestions, summaries, or other relevant content.
[0272] "Structured processing" refers to the process in information processing devices that parses, segments, classifies, and reorganizes response information in order to map different types of content to predefined information categories or data fields.
[0273] "Meeting key information" refers to a concise collection of project information extracted and organized from the response information, which summarizes the core content and main conclusions of the meeting discussion.
[0274] "Time setting suggestion information" refers to the suggestive content in the response information related to time scheduling, deadlines, milestone settings, etc., used to set or adjust time parameters.
[0275] "Recommendation information related to risks or resources" refers to the recommendations in response information related to potential risk identification, risk mitigation measures, staffing, and allocation of physical or computing resources.
[0276] "Display information" refers to a collection of meeting key points, time setting suggestions, and risk or resource-related suggestions that have been structured and organized in a format suitable for presentation on a terminal display interface.
[0277] "Emotion estimation processing" refers to the process of analyzing the textual content of meeting minutes to infer the speaker's emotional tendency (such as positive, neutral, negative, etc.) and intensity.
[0278] "Emotional information" refers to quantitative or qualitative data obtained through emotion estimation processing, used to characterize the emotional state of meeting participants or the overall atmosphere of the meeting.
[0279] "Information for external use" refers to data content selected from meeting minutes and response information and organized according to a predefined format for providing meeting outcomes or intelligent analysis results to external entities.
[0280] "External information processing infrastructure" refers to a computing platform used to receive, store, display, or reprocess information provided to external parties outside of an organization or system, including remote servers, cloud computing platforms, or other network service systems.
[0281] In one embodiment of the present invention, a server, a terminal, and a user collaboratively constitute a computer system for conference support. The server includes a processor, memory, a network interface, and non-volatile storage media. The processor runs an operating system and multiple application modules, including an audio receiving module, a speech recognition interface module, a text processing module, a prompt generation module, a generative artificial intelligence model interface module, a sentiment analysis module, a result structuring module, and a display data generation module. The terminal includes a processor, memory, an audio acquisition device (microphone), a display device (liquid crystal display or organic light-emitting diode display), and a network communication module, and runs recording and communication applications as well as user interface applications thereon. Users operate the terminal at the conference venue to record audio and view analysis results.
[0282] At the hardware level, the server can utilize a general-purpose server hardware platform. Its central processing unit (CPU) can be a multi-core general-purpose processor, supplemented by a graphics processing unit (GPU) or a dedicated accelerator to accelerate the inference processing of generative artificial intelligence models. At the software level, the server can run a general-purpose operating system and backend applications built on the server framework, such as web services implemented using a general-purpose web framework, used to receive audio data sent from terminals and return analysis results. The server can also perform some computational tasks by calling external speech recognition services and external generative artificial intelligence model services. These external services can be accessed through standard HTTP interfaces or software development kits, such as general-purpose cloud speech recognition APIs and general-purpose generative language model APIs. In its implementation, the server can use software components such as audio processing libraries (e.g., FFmpeg, SoX), word segmentation libraries (e.g., Chinese word segmentation libraries), and deep learning inference frameworks (e.g., general-purpose tensor computation frameworks).
[0283] At the hardware level, the terminal can be a portable or desktop terminal. Its processor executes a terminal application to call the audio acquisition interface provided by the operating system, such as the audio recording interface provided in a general mobile operating system or the multimedia interface provided in a general desktop operating system. The terminal uses a built-in microphone to convert the voice signals of the user and other participants into digital audio data, and uploads the audio data to the server through a network communication module (such as a wireless LAN interface or a cellular communication module). The terminal also uses a graphical user interface framework to display the analysis results returned by the server on the display device in the form of lists, paragraphs, highlights, etc.
[0284] Before the meeting begins, users launch the meeting assistance application via their terminal, selecting or entering information such as the meeting identifier and key agenda items. Users can select parameters such as audio sampling rate and language type through the terminal interface, which then configures the audio acquisition module and network upload strategy accordingly. During the meeting, users do not need to manually record in segments or manually upload data; the terminal continuously and automatically acquires and uploads the meeting audio.
[0285] After receiving audio data uploaded by the terminal, the server first uses the audio receiving module to maintain an independent audio buffer in memory for each conference identifier. The server associates and stores each audio data unit with metadata such as timestamps and segment numbers, forming a time-ordered sequence of audio segments. The server can call the audio processing library as needed to perform unified format conversion on audio from different encoding formats, such as converting it to a linear pulse code modulation format with a fixed sampling rate and bit depth, to meet the input requirements of the subsequent speech recognition interface. Through this preprocessing, the server can ensure that the audio data input to the speech recognition module maintains consistency in sampling parameters, thereby reducing the recognition error rate.
[0286] The server calls external or local speech recognition services through a speech recognition interface module to convert continuous audio segments into text information. The server can use a streaming recognition mode, sending an audio window of a certain duration to the speech recognition service and gradually receiving the returned partial transcription results. When receiving each transcribed text segment, the server binds it to its corresponding time range and appends it chronologically to the text buffer corresponding to that meeting. The server can use a structured data structure for storage, such as a collection of records based on meeting identifier, start time, end time, and text content, facilitating subsequent querying and reorganization by time and topic.
[0287] After receiving a certain length of meeting minutes, the server's text processing module preprocesses the text, including sentence segmentation, paragraphing, and removal of meaningless filler words. When needed, the server can also invoke word segmentation algorithms and simple syntactic analysis algorithms to perform basic structural analysis of the text information to extract potential keywords. For example, the server can count the frequency of keywords such as "deadline," "time," "risk," and "resources" and their co-occurrence relationships in the meeting minutes, and based on this, identify the decision-making matters that may be involved in the current meeting segment.
[0288] After recognizing a text fragment related to a predetermined topic, the server constructs a prompt statement using the prompt statement generation module. The server pre-stores several prompt statement templates in the storage medium. These templates are represented in natural language text containing placeholders and are used to describe the target role, task requirements, output format, etc. The server selects an appropriate template based on the current topic and replaces the placeholders in the template with a fragment of the current meeting minutes, thereby generating the complete input text.
[0289] For example, when the server detects that the meeting minutes contain discussions related to "next project deadline," it can generate the following prompt: "You are a senior project management consultant. Please refer to the following meeting minutes:" 1. Use 3 Five key points summarizing the discussion related to 'next project deadline'; 2. After summarizing, give 2 Three specific and actionable deadline adjustment suggestions are provided, including: recommended deadlines, resources that need to be adjusted, and risks that need to be noted.
[0290] The following are the meeting minutes: Meeting minutes begin (Insert meeting minutes text here) Meeting minutes closed. For example, the server can generate prompts focusing on divergence points as needed: "You are now acting as a neutral project coordinator. Please read the meeting minutes below and complete the following tasks:" 1. Identify the main points of disagreement regarding the 'next project deadline' and list the different viewpoints in a list format; 2. Propose a compromise solution and explain why it achieves a good balance between time, quality, and cost.
[0291] The meeting minutes are as follows: (Insert meeting minutes text here) When constructing prompts, the server doesn't simply concatenate text. Instead, based on word segmentation and keyword statistics, it prioritizes sentences strongly related to the target topic and includes them in the prompts. It can also prune excessively long historical content based on the current meeting progress to control the input length of the generative AI model. This selection strategy, based on semantic relevance and length constraints, reduces the interference of irrelevant information on the model's reasoning, thereby improving the relevance of the generated results and computational efficiency.
[0292] The server sends the constructed input information to the generative AI model through the generative AI model interface module. The generative AI model can be a sequence-to-sequence neural network structure based on a multi-layer self-attention mechanism, including multi-layer encoders and multi-layer decoders, or a stacked decoder network using an autoregressive structure. When the server calls the model, it merges the prompt and meeting minutes into a single input sequence, converts it into a discrete labeled sequence through text tokenization, and then maps it to a high-dimensional vector space through the model's embedding layer. Internally, the model calculates attention weights between different labels using a multi-head attention mechanism to establish long-distance dependencies, and performs non-linear transformations on the intermediate representations through a feedforward network in each layer. When generating output, the model uses a conditional probability distribution to generate response text label by label in an autoregressive manner.
[0293] When deploying generative AI models, servers can employ a pre-training and fine-tuning approach. During pre-training, the model uses a large-scale, general corpus for language modeling training, employing a cross-entropy loss function to measure the difference between the output labeled sequence and the training reference sequence, and updating weight parameters through backpropagation. In its secondary training phase, the server can further fine-tune the model using corpora relevant to meeting recording scenarios, allowing it to better adapt to meeting summary and suggestion generation tasks. During the inference phase, the server can set temperature parameters, sampling strategies (such as beam search or temperature-controlled random sampling), and maximum output length to control the diversity and stability of the generated text.
[0294] In some implementations, the server uses a sentiment analysis module to perform sentiment estimation on meeting transcripts. The server can employ models such as convolutional neural networks or bidirectional long short-term memory networks, accepting textual representations or word vector sequences of meeting speeches as input and outputting corresponding sentiment category labels and sentiment intensity scores. When training the sentiment analysis model, the server can use sentiment-annotated corpora and adjust model parameters using multi-class cross-entropy or regression loss functions. During inference, the server adjusts the input or decoding strategies of the generative AI model based on sentiment information; for example, it reduces the probability of using overly aggressive wording in the model output when emotions are tense, thereby generating suggested text more suitable for the current meeting atmosphere.
[0295] After receiving the response information output by the generative artificial intelligence model, the server uses a result structuring module to parse and classify the text. The server can use a combination of rule-based and statistical methods to identify structural markers in the response text, such as keywords like "Summary:", "Recommendation:", and "Risk:", and divide the text into multiple parts based on these markers. The server can also utilize simple classification models or keyword-matching strategies to categorize sentences or paragraphs into predefined categories such as "Meeting Key Points Information", "Time Setting Proposal Information", and "Risk or Resource Related Recommendation Information", storing this information in structured data structures, such as object records or tabular records.
[0296] When generating display information, the server converts the structured results into a format suitable for display on the terminal. The server can organize multiple key items into an ordered list, present time-based suggestions in tabular form, and display risk warnings and resource suggestions separately in different sections. The server can also generate highlight markers or color codes to distinguish between high-priority risks and general alerts. When sending display information to the terminal, the server can use a lightweight data format for transmission over the network, thereby reducing transmission load and terminal parsing overhead.
[0297] After receiving display information from the server, the terminal parses it into internal data structures and displays them using user interface components. The terminal can display different types of information in a split-screen view, such as a list of meeting highlights on the left and scheduling suggestions and risk warnings on the right. Users can scroll through, expand, or collapse specific items on the terminal and mark or share certain suggestions as needed. The terminal can also cache structured information locally for offline viewing after the meeting.
[0298] The system of this invention does not merely automate manual meeting recording; rather, it achieves collaborative optimization of audio data, text data, emotional information, and generated results by establishing an end-to-end data processing pipeline within the server. In the audio access stage, the server reduces latency and improves speech recognition accuracy through unified format conversion and streaming upload. In the text processing and prompt generation stage, segment selection and template filling based on keywords, word segmentation, and sentiment estimation make the input of the generative AI model more focused on the current decision-making issue, thereby reducing noise and computational redundancy from irrelevant information. In the model inference stage, the stability and readability of the output content are improved by appropriately setting temperature, sampling strategies, and length control. In the result structuring and display stage, the efficiency of terminal presentation and user comprehension speed are improved by mapping the model output into multi-dimensional information units. These processing steps form specific data structures and algorithmic flows within the computer, making the overall system superior to simple recording and full-text generation solutions in terms of processing speed, resource consumption, and output quality.
[0299] This invention can be extended in multiple embodiments. For example, the server can rely only partially on external speech recognition services while deploying a lightweight acoustic model locally for local recognition during network instability and for result comparison and correction after network recovery. The server can load different prompt statement template sets and topic keyword sets according to different meeting types to support various scenarios such as project management meetings, technical review meetings, and human resources meetings. The server can distribute the processing tasks of different meetings across multiple physical servers or container instances based on the number of terminals and load conditions, achieving horizontal scaling through task queues and load balancers. Terminals in different embodiments can also support multiple audio inputs to distinguish different speaker channels, and the server can utilize this channel information to further improve the accuracy of text segmentation and speaker identification.
[0300] Through the above structure and processing flow, this invention achieves efficient acquisition of conference voice information, structured management of conference text information, precise control of the input of generative artificial intelligence models, and multi-dimensional organization and display of model output results. This improves the speed, accuracy, and usability of conference data processing at the computer technology level and provides a high-quality data foundation for subsequent automatic archiving, retrieval, and statistical analysis.
[0301] use Figure 13 The processing flow is explained.
[0302] Step 1: Users launch the meeting application and configure meeting parameters on their terminals.
[0303] Users input or select parameters such as meeting identifier, meeting agenda, and target language through the terminal's graphical interface, and select configurations such as audio sampling rate and upload method as needed. The input consists of the meeting information and system parameters set by the user on the interface, and the output is a meeting session configuration data structure (containing meeting ID, agenda, audio parameters, etc.) generated in the terminal's memory. The terminal initializes the recording module and network communication module based on this configuration data structure, preparing for subsequent audio acquisition and data upload.
[0304] Step 2: The terminal uses its built-in microphone to capture conference voice and generate audio data.
[0305] The terminal invokes the operating system's audio acquisition interface to initiate recording at the set sampling rate and bit depth, converting the analog sound signal in the environment into digital audio samples. The input is the analog speech signal from the meeting venue, and the output is raw audio data frames (such as PCM format) stored in a buffer in chronological order. Internally, the terminal manages these audio frames through a circular buffer, providing a data source for subsequent segmentation and uploading.
[0306] Step 3: The terminal segments the audio data and attaches metadata.
[0307] The terminal periodically reads fixed-length audio data from the buffer (e.g., every 500 milliseconds) in the recording thread and appends metadata such as meeting ID, timestamp, and segment number to each audio segment. The input is continuous raw audio data frames, and the output is a series of audio data segments with metadata tags. During this process, the terminal performs simple slicing operations on the raw sampled data without changing the sampled content, thereby forming audio block units that are easy to transmit over the network and process by the server while maintaining the temporal order.
[0308] Step 4: The terminal uploads audio clips with metadata to the server.
[0309] The terminal establishes or reuses a connection with the server (such as an HTTP long connection or a WebSocket session) through the network communication module and encapsulates each audio segment into a network message for transmission. The input is the audio segment and its metadata generated in step 3, and the output is the data packet encapsulated by the network protocol and transmitted over the communication link. Before transmission, the terminal checks the segment length; if a network anomaly is detected, the segment is temporarily stored in a local queue and transmitted only after the network is restored, thus mitigating data loss.
[0310] Step 5: The server receives audio segments and performs format verification and caching.
[0311] The server receives audio data packets from terminals via a network interface and parses them to extract the meeting ID, timestamp, segment number, and audio content. The input is encapsulated data packets received from the network layer, and the output is a list of audio segments indexed by meeting ID in the server's memory. After parsing, the server checks if the audio format conforms to preset parameters (sampling rate, number of channels, bit depth). For segments that do not meet the requirements, the server can perform format conversion or log errors. The server inserts audio segments into a buffer in the order of arrival using a simple queue or list structure, providing the basic data for subsequent speech recognition window stitching.
[0312] Step 6: The server concatenates the cached audio and uses speech recognition to convert the audio into text.
[0313] The server extracts unrecognized continuous audio segments (e.g., 10 or 30 seconds) from the audio buffer in chronological order. If necessary, it uses an audio processing library for resampling or encoding conversion, concatenating the multiple segments into a continuous audio stream. The input is multiple audio segments arranged in chronological order, and the output is a single audio buffer block that meets the requirements of the speech recognition service. The server then submits this buffer block to the speech recognition interface module, which calls the recognition service via API and receives the returned text transcription results. The input is a uniformly formatted audio buffer block, and the output is text information with time range markers (e.g., a sentence and its start and end times). The server appends the recognized text to the corresponding meeting's text recording buffer and marks the audio range as "recognized" to avoid duplicate processing.
[0314] Step 7: The server organizes the text records and extracts text fragments related to the topic.
[0315] The server reads the accumulated text information of the current meeting from the text record buffer, and performs sentence segmentation and simple paragraphing (e.g., segmentation by punctuation and pause times). The input is a raw transcribed text sequence arranged in chronological order, and the output is a series of sentence or paragraph units with timestamps and sentence boundary markers. The server performs keyword statistics and word segmentation on these units, calculating their relevance to preset topic keywords (such as "deadline," "progress," and "risk"). Based on this, the server selects several sentences or paragraphs highly relevant to the target topic, concatenates them in chronological order, and generates the topic-related text fragments to be analyzed. The input is a complete set of text record units, and the output is a filtered set of topic-related text fragments used in the construction of subsequent prompts.
[0316] Step 8: The server generates prompts and constructs input information for the generative artificial intelligence model.
[0317] The server reads a prompt template matching the current topic from the storage medium and inserts the topic-related text fragments obtained in step 7 into the template. The input consists of the prompt template and the topic-related text fragments; the output is the complete prompt text, which is a combination of text including role settings, task descriptions, output requirements, and actual meeting minutes. During construction, the server truncates or summarizes the text based on fragment length and model input constraints to control the input length. For example, the server can generate the following prompt: "You are a senior project management consultant. Please refer to the following meeting minutes:" 1. Use 3 Five key points summarizing the discussion related to 'next project deadline'; 2. After summarizing, give 2 Three specific and actionable deadline adjustment suggestions are provided, including: recommended deadlines, resources that need to be adjusted, and risks that need to be noted.
[0318] The following are the meeting minutes: Meeting minutes begin (Insert relevant text snippets) Meeting minutes closed. The server can also generate prompts emphasizing points of divergence, such as: "You are now acting as a neutral project coordinator. Please read the meeting minutes below and complete the following tasks:" 1. Identify the main points of disagreement regarding the 'next project deadline' and list the different viewpoints in a list format; 2. Propose a compromise solution and explain why it achieves a good balance between time, quality, and cost.
[0319] The meeting minutes are as follows: (Insert relevant text snippets) These prompts guide generative AI models to generate outputs with specific structures and focuses through clear task descriptions and structural requirements.
[0320] Step 9: The server performs sentiment estimation as needed and adjusts the model inputs or parameters accordingly.
[0321] The server utilizes a sentiment analysis module to classify and estimate the sentiment intensity of the text representations of each sentence in the meeting transcript. The input consists of text units with sentence boundaries, and the output is a sentiment label (e.g., positive, neutral, negative) and sentiment intensity value for each unit. Based on the overall sentiment trend and the sentiment peaks of intensely debated sections, the server adjusts the wording of prompts or sets control parameters for the generative AI model (e.g., requiring more moderate suggestion expressions). In this way, the text information input to the model not only contains objective content but also implicitly considers emotional states, helping to reduce overly aggressive or inappropriate outputs.
[0322] Step 10: The server invokes a generative artificial intelligence model to perform semantic parsing on the input information and generate response information.
[0323] The server sends the complete prompt statement constructed in step 8, along with any accompanying control parameters (such as temperature and maximum output length), to the generative AI model interface. The input is a text string containing the prompt statement and meeting minutes, along with the control parameters; the output is the natural language response text generated by the model. Internally, the input text is first segmented and mapped into embedding vectors, then feature extraction is performed through a multi-layer self-attention network. The server receives the final decoded output token sequence at the interface layer and restores it to continuous text. The server uses this text as the original response information and stores it in memory as associated with the corresponding meeting ID.
[0324] Step 11: The server performs structured processing on the response information and generates information for display.
[0325] The server parses the response text output by the generative AI model, using rule matching and segmentation algorithms to identify paragraph start markers such as "summary," "recommendation," and "risk," dividing the text into several logical parts. The input is a continuous string of response text, and the output is a structured data object containing a list of meeting key points, a list of time-setting suggestions, and a list of risk or resource-related suggestions. The server then reorganizes these lists into a display structure, for example, generating an entry object for each key point and adding a recommended deadline and a brief reason field for each time suggestion. In this process, the server completes the transformation from unstructured text to a visually appealing data structure, making it easy for the terminal to present and for the user to understand.
[0326] Step 12: The server sends display information to the terminal, which then presents it for the user to view and utilize.
[0327] The server invokes the network sending module to encode the structured data generated in step 11 into a transmission format and send it to the terminal. The input is a structured display information object, and the output is a response message transmitted over the network. Upon receiving the message, the terminal parses the data structure and invokes user interface components to generate the corresponding list view, title area, and detailed description area. The terminal displays blocks such as "Meeting Highlights," "Deadline Suggestions," and "Risk and Resource Suggestions" on the screen. Users can browse, expand details, and record key points using touch or keyboard and mouse operations. Based on this information, users conduct discussions and make decisions during or after the meeting, thus combining the data processing completed internally by the server with the real-world meeting workflow to provide technical support for meeting decisions.
[0328] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0329] In the fields of meeting support technology and operation management / business management technology, existing computer systems typically only save meeting records as static text, or only perform statistical analysis and display of meeting content based on simple rules. Existing technologies have the following problems: First, existing systems struggle to perform multimodal fusion processing of meeting audio data, text data, image data, and environmental data (such as mobile device operation information and business operation information) under a unified architecture, resulting in an inability to achieve a comprehensive and real-time computerized understanding of the meeting process. Second, while existing systems can call upon artificial intelligence models to generate opinions, this is often based solely on text content, failing to fully consider the emotional state and changes of meeting participants. This leads to opinions that do not accurately reflect the actual communication atmosphere, affecting decision-making efficiency and acceptance. Third, in scenarios involving mobile device operation management or complex business management, existing technologies typically require manual reinterpretation and input of the natural language results output by artificial intelligence into scheduling or business systems. This lack of an integrated computer processing flow from "meeting understanding → opinion generation → structured operation / business instructions → automatic control" results in information loss, increased latency, and increased error susceptibility.
[0330] Furthermore, existing systems generally lack a dynamic optimization mechanism for input prompts for generative AI models. In other words, they cannot automatically adjust the prompt generation strategy and model operating parameters based on historical output results and user feedback. This results in the model output quality remaining at a fixed level for a long time, making it difficult to continuously adapt and optimize in specific enterprise environments, meeting cultures, and operating scenarios.
[0331] In summary, from a computer technology perspective, how to provide a system that can: acquire and fuse meeting data in a multimodal manner, generate structured decision information by combining sentiment analysis, automatically drive operation management or business management devices, and continuously optimize prompts and generative artificial intelligence model behavior based on user feedback is a technical issue that urgently needs to be addressed by those skilled in the art.
[0332] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0333] In this invention, the server includes: a device for generating prompt statements based on meeting topic information and related environmental information to instruct a generative artificial intelligence model to generate opinions and proposals; a device for acquiring meeting audio information and converting the audio information into time-sequential meeting record data using speech recognition technology; a device for estimating the emotional state of participants and generating emotional time change information based on the meeting record data and image information and voice features of meeting participants using sentiment analysis technology; and a device for embedding the meeting record data and emotional state information into the prompt statements and inputting them into the generative artificial intelligence model to obtain output information containing opinions on the topics and responses considering emotional states. This system comprises a device for extracting path information, planning information, or business policy information related to the operation or business from output information and converting it into structured data that can be directly used by the operation or business plan; a device for presenting the structured data and output information on a display device or mobile device and providing a user interface for receiving user selection or confirmation operations; a device for outputting control commands or setting change information to a control or management device based on user operations to automatically execute operation management or business management; and a device for storing output information and evaluation or feedback information from users and updating the prompt statement generation conditions or generative artificial intelligence model operation conditions based on the information. This creates a closed-loop data processing chain within the computer, from multimodal meeting data acquisition, emotional state inference, automatic prompt statement construction, generative artificial intelligence model reasoning, result structure conversion to operation / business control execution and continuous feedback learning. This substantially improves the data processing efficiency, decision automation, and intelligence level of the meeting support system and the operation / business management system, achieving technical optimization of the input / output process of the generative artificial intelligence model and the overall system performance.
[0334] A "system" refers to an overall technical solution consisting of one or more information processing devices, storage devices, network devices, and input / output devices, which coordinate data acquisition, data processing, information output, and control command generation among the devices through a pre-set program.
[0335] "Meeting topic information" refers to structured or unstructured information related to the topics, objectives, and decision-making matters discussed in the meeting, including topic names, background information, problems to be solved, and constraints, which is used to guide generative artificial intelligence models in generating and processing opinions.
[0336] "Environmental information" refers to external condition information related to the meeting venue or business scenario, including but not limited to the operational status information of mobile entities, geographical location information, traffic conditions information, business status information, time information, etc., which is used to provide context for opinion generation and decision-making.
[0337] "Generative artificial intelligence models" refer to artificial intelligence models trained based on machine learning algorithms that can automatically generate new natural language text or other forms of output based on input text, speech or other features, including natural language generation models trained using deep learning and large-scale corpora.
[0338] "Prompt statements" refer to the instructional or explanatory text used as input to generative artificial intelligence models. Their content includes task descriptions, role settings, data summaries, and generation requirements, which are used to constrain and guide generative artificial intelligence models to output results in the desired form and content.
[0339] "Voice information" refers to human voices or other audio signals, represented in analog or digital form, obtained from the meeting venue through a sound pickup device, used to reflect the content, intonation, and speaking time of the participants in the meeting.
[0340] "Speech recognition technology" refers to the computational processing technology that converts collected speech information into corresponding text information, including matching acoustic features with language models to output a character sequence representing the speech content.
[0341] "Text information" refers to text data obtained through speech recognition technology or other input methods, used to record meeting speeches, system outputs, or other linguistic expressions for subsequent analysis and storage.
[0342] "Meeting minutes data" refers to a collection of textual data obtained by chronologically organizing and structurally storing various speeches and related events during a meeting. It typically includes speaker identifiers, timestamps, speech content, and related metadata for subsequent analysis and retrieval.
[0343] "Meeting participants" refers to individuals or groups who speak, listen, or participate in decision-making during a meeting, including human participants and, in some embodiments, artificial intelligence entities that participate in the meeting process.
[0344] "Image information" refers to image data acquired by imaging devices and stored in digital form, used to reflect the visual characteristics of the participants in the meeting, such as their appearance, facial expressions, and posture.
[0345] "Speech features" refer to the set of parameters extracted from speech information for analyzing emotion or identifying speaker characteristics, including pitch, volume, speech rate, formants, spectral features, etc.
[0346] "Emotional analysis technology" refers to the technology that identifies, classifies, or quantifies the emotional state in speech, text, or image information based on machine learning, signal processing, or rule models, and is used to infer the type and intensity of emotions such as pleasure, anger, tension, and anxiety.
[0347] "Emotional state" refers to the result of judging the psychological or emotional tendencies of meeting participants within a specific time period through emotion analysis technology, including the emotion category and corresponding intensity.
[0348] "Time change information" refers to a sequence of values or symbols that represent how a certain quantity changes over time. In this invention, it is mainly used to represent the dynamic trajectory of emotional state changes during a meeting.
[0349] "Output information" refers to textual or structured data obtained from generative artificial intelligence models, which includes opinions, proposals, and responses that take into account the emotional state of the participants, and is used for subsequent decision-making and control processing.
[0350] "Run as an object" refers to a physical or virtual entity that serves as the target of operation management in the system, including but not limited to mobile bodies, transportation equipment, production line equipment, or other objects that require path planning and operation scheduling.
[0351] "Business as an object" refers to business processes, business units, or organizational activities that are the targets of business management and optimization in the system, including supply chain processes, marketing activity processes, production planning processes, etc.
[0352] "Path information" refers to spatial location and connection information related to the movement or task execution sequence of the object being run, including path nodes, path order, and related constraints, which is used for route planning and navigation control.
[0353] "Planning information" refers to the time schedule and resource allocation information formulated for operation objects or business objects, including start and end times, execution order, resource usage, priority, etc., which is used to achieve scheduling and control.
[0354] "Business policy information" refers to the objectives and strategic information used to guide decision-making in business management or enterprise operations, including development direction, risk appetite, market strategy, and resource allocation principles.
[0355] "Structured data" refers to data forms with clearly defined fields and data types that facilitate retrieval, computation, and association by computer programs, such as hierarchical data with tables, records, or annotation fields.
[0356] "Display device" refers to a hardware device that can output text, images or graphics information in a visual manner, including displays, projection devices, vehicle display terminals, etc., used to present system processing results to users.
[0357] "Mobile in-body devices" refers to terminal devices installed inside a mobile body for information display, interaction, or control, including vehicle-mounted information terminals, ship control console terminals, engine room display terminals, etc.
[0358] A "user interface" refers to the combination of hardware and software used for information interaction between people and systems, including graphical interfaces, touch interfaces, voice interaction interfaces, etc., which are used to receive user input and provide feedback on system status.
[0359] "User" refers to an individual or organization that directly operates or indirectly utilizes the system's functions, including meeting hosts, participants, drivers, operations managers, or other end users.
[0360] "Selection operation" refers to the input behavior of a user choosing from multiple candidate options through a user interface, including clicking, touching, voice command selection, etc.
[0361] "Confirmation operation" refers to the input behavior by which a user makes a final approval of a candidate solution, suggestion, or control instruction, including pressing the confirmation key, submitting the submission button, or confirming via voice.
[0362] "Control device" refers to a device used to perform specific control actions on a moving object, including vehicle control unit, production equipment control unit, robot control unit, etc., which can adjust the physical state according to control commands.
[0363] "Management device" refers to the device or system used to perform planning, resource allocation and status monitoring for business-related tasks, including business management servers, scheduling systems, enterprise resource planning systems, etc.
[0364] "Control commands" refer to instruction information generated by the system and sent to the control device to drive the object to perform specific actions or adjust operating parameters, including commands such as start, stop, speed adjustment, and path change.
[0365] "Configuration change information" refers to configuration information used to modify the internal parameter settings of control or management devices, including target values, thresholds, priorities, strategy parameters, etc., which can be used to change operation or business control strategies.
[0366] "Operations management" refers to the monitoring, scheduling, optimization, and control activities carried out on operations, including route planning, driving plan formulation, and real-time adjustments, in order to improve operational efficiency and safety.
[0367] "Business management" refers to the organization, planning, execution, and monitoring activities focused on business operations, including task allocation, resource scheduling, progress tracking, and performance evaluation, in order to improve business efficiency and quality.
[0368] "Evaluation information" refers to objective or subjective evaluation data given by users to the output information or control results provided by the system, including scores, labels, selection results, etc., which are used to reflect output quality and satisfaction.
[0369] "Feedback information" refers to the opinions, suggestions, or correction requests made by users based on the system output results, including natural language descriptions, question and answer results, revision instructions, etc., which are used to guide the subsequent optimization of the system.
[0370] "Conditions for generating prompt statements" refers to the set of parameters and rules used when generating prompt statements, including topic category, data summary length, tone requirements, output format requirements, and weight settings related to historical effects.
[0371] "The operating conditions of generative artificial intelligence models" refers to the set of parameters and strategies used when calling generative artificial intelligence models for inference, including temperature parameters, maximum output length, decoding strategies, context truncation strategies, etc. These conditions affect the output style and content quality of the model.
[0372] This invention will describe a conference support and operation / business management system that operates collaboratively among a server, a terminal, and a user. The following description uses "server," "terminal," and "user" as the main terms, focusing on the internal data structures, algorithm flows, model structures, and their technical effects, without involving program code.
[0373] (I) Overall Hardware and Software Composition A server includes a processor, main memory, non-volatile memory, a network interface, and an optional graphics processing unit. The server runs several software components on top of an operating system, including: - Voice acquisition and encoding module (e.g., based on a general audio library); - A speech recognition client module for calling external speech recognition services (such as a general cloud speech recognition API, corresponding to "Google Speech" in existing technologies). to Text, etc.); - Natural language processing libraries (such as general-purpose word segmentation and dependency syntax libraries, and Transformer-based text embedding libraries); - Generative AI model client module, used to invoke remote generative AI models (such as large-scale language models trained on the Transformer architecture, functionally similar to the existing GPT). 3 / GPT 4); - Emotion recognition module, including: - Image processing submodule (e.g., face detection and expression feature extraction based on OpenCV and similar image processing libraries); - Text sentiment analysis submodule (can call general sentiment analysis services, such as the "TextAnalytics" API). - Voice sentiment analysis submodule (using a voice feature extraction and classification model); - Path / Plan Generation and Optimization Module (used to generate and adjust mobile vehicle operation paths or business plans); - User interface service modules (e.g., based on web server frameworks such as Flask or other HTTP servers) - Data storage module (combined with relational database or document database).
[0374] The terminal includes hardware such as a display, touch panel or keypad input device, microphone, camera, and network interface, as well as a browser or dedicated application. In this invention, the terminal primarily performs interface display, user input collection, and audio / video reporting functions.
[0375] Users are human operators such as meeting participants, operations managers, and drivers, who interact with the server through terminals.
[0376] (II) Server-side data structure and module composition The server maintains several core data structures in main memory: - Meeting minutes data structure: - Fields include: speech ID, speaker ID, timestamp, original text content, confidence level, source channel ID, etc.; The server creates an ordered conference text stream by inserting speech recognition results into the structure in chronological order.
[0377] - Meeting information context structure: - Fields include: meeting agenda information, meeting objectives, current agenda paragraph summary, and relevant environmental information (mobile location, operating status, service status, etc.).
[0378] - Emotional timeline structure: - For each participant in the meeting, the server maintains a time series, with elements being (timestamp, sentiment vector). The sentiment vector can include numerical indicators such as pleasure, tension, and anger.
[0379] - Prompt statement template structure: - Store template text for various task types (such as route optimization, business analysis, and advertising creative), with placeholders, for injecting meeting minutes summaries, sentiment states, environmental data, etc.
[0380] - Output information structure: - Includes the raw text, structured fields (such as suggested routes and suggested strategies), confidence scores, and post-processing labels returned by the generative AI model.
[0381] - Control command structure: - Control commands for mobile devices or business systems, such as waypoint lists, parameter settings, and time schedules, can be easily issued to external control or management devices.
[0382] The server divides its functions into multiple logical modules, which communicate with each other through clear interfaces and data structures, thereby improving processing efficiency and maintainability.
[0383] (III) Specific Implementation Forms of Speech and Image Processing The server uses a speech recognition client module to process the audio stream uploaded via the terminal. For each audio segment, the server performs the following calculations: The server uses audio preprocessing algorithms (such as short-time Fourier transform, noise estimation and suppression) to denoise and normalize the energy of the input waveform.
[0384] The server calls a cloud-based speech recognition service with the processed audio segment. The cloud service uses acoustic and language models internally. This invention does not limit the specific implementation, but the server-side is limited to returning a text sequence with a timestamp.
[0385] The server writes the text sequence into the meeting record data structure and records the timestamp and speaker identifier for each text (which can be distinguished using a voiceprint separation algorithm or a multi-channel microphone array).
[0386] The server uses an image processing submodule to analyze the video frames uploaded by the terminal. The server uses face detection algorithms (such as detectors based on Haar features or convolutional neural networks) to locate face regions in video frames.
[0387] The server extracts local feature vectors of the face region (e.g., coordinates of 68 key points, local texture features) and feeds them into a pre-trained expression classification neural network. This network can be a multi-layer convolutional + fully connected structure, outputting probability distributions for several emotion categories.
[0388] - The server records the emotional expression of each frame along with a timestamp in the emotional timeline structure.
[0389] The server uses a speech sentiment analysis submodule to extract acoustic features from the audio (such as Mel-frequency cepstral coefficients (MFCC), pitch contours, and energy envelopes), which are then input into a multilayer perceptron or a convolutional + recurrent hybrid network to output a numerical value for the sentiment dimension. The server then performs a weighted fusion of speech sentiment and image sentiment to obtain a more robust sentiment state estimate.
[0390] The server uses the text sentiment analysis submodule to analyze the meeting transcript data: The server uses a Transformer-based text embedding model to encode each message into a vector.
[0391] The server will embed a vector input sentiment classifier (which can use a finely tuned Transformer output header) to obtain positive / negative / neutral sentiment labels and intensity scores.
[0392] The server merges the results from multiple channels to form a comprehensive emotion vector for each time slice, which is then written into the emotion timeline structure in time series form. This significantly improves the accuracy and robustness of emotion recognition.
[0393] (iv) Specific implementation forms of generative artificial intelligence models and prompt statements The server accesses a large-scale language model through a generative artificial intelligence model client module. This model employs a multi-layered Transformer architecture, including self-attention layers, feedforward layers, and layer normalization components, and is trained through unsupervised language modeling and instruction fine-tuning. The server does not need to generate model weights locally, but it constructs precise prompts at the time of invocation to control the model's output content and style.
[0394] The server configures different prompt message templates for different tasks. For example: - Example of a prompt message from the server during a conference on autonomous vehicles: "The autonomous vehicle is currently traveling from the starting point to its destination. The following is a record of the conversation over the past 5 minutes:" ... Please provide three feasible routes, taking into account current traffic conditions and safety. For each route, please describe the estimated arrival time, main road sections, potential risks, and applicable scenarios. - Example of server-side prompts in enterprise business analysis: "You are the enterprise decision support system. Below are the quarterly sales data and key market trends for the past three years:" ... Assuming the macroeconomic environment remains stable, please forecast the sales range for the next quarter, and explain the main uncertainties and three actionable risk management strategies. - Example of a prompt message from the server in an ad creative scenario: "Design five creative advertising prompts for the following new product that will generate buzz and social media impact."
[0395] Product Information: - Product Type: Low-Sugar Energy Drink - Key selling points: 0 sugar, guilt-free, designed specifically for urban dwellers aged 20-30 - Brand Tone: Young, Courageous, and Accelerating Life Output format: Each prompt should be on a separate line, using concise and vivid language. When constructing the prompt statement, the server follows these technical rules: - The server limits the input length (e.g., the text from the most recent N minutes or the most recent M messages) to control the computational load of the model, reduce invalid context, and lower communication and inference latency.
[0396] - The server adds structured tags to the prompt statements, such as "meeting minutes" and "sentiment analysis," to guide the model to output according to the expected structure, which is beneficial for subsequent automatic parsing.
[0397] - The server injects a summary of the emotional timeline into the prompts, such as "Participant A's tension increases" or "Participant B's anger increases," to prompt the model to consider soothing or softening language when outputting suggestions.
[0398] By using the aforementioned unconventional prompt structure and emotion injection method, the server can enable generative AI models to produce outputs that are easier for subsequent programs to process and more sensitive to emotions, thereby improving the overall system's processing accuracy and decision-making quality.
[0399] (v) Specific Implementation Forms of Output Information Structure and Equipment Control After receiving the natural language output from the generative artificial intelligence model, the server does not simply present it to the user, but performs structured parsing: - The server uses a natural language parsing submodule (which can be based on regular expressions, dependency parsing, and entity recognition) to extract key fields from the text, such as path name, estimated arrival time, risk description, priority label, etc.
[0400] - The server maps these fields to predefined control instruction structures, such as: - For moving objects: Convert place names into a list of latitude and longitude nodes, and verify connectivity using existing map data; - For business processes: convert them into task nodes, dependencies, and time windows.
[0401] - The server runs graph algorithms or scheduling algorithms (such as Dijkstra, A*, heuristic scheduling algorithms) in the path / plan generation and optimization module to further calculate the shortest path, minimum delay, or minimum cost of the generated candidate solutions, and obtain optimized structured data.
[0402] The server sends the final structured data, conforming to a predefined protocol, to the specific control or management device via a network interface. Upon receiving the structured data, the control device can directly modify the movement path of the mobile device or the scheduling strategy of the business system. This automatic conversion process from natural language output to executable control commands avoids secondary manual input, significantly reduces human error, and improves response speed.
[0403] (vi) Specific implementation forms of terminal-user interaction As the carrier of the user interface, the terminal mainly performs the following technical actions: - The terminal renders structured data and output information returned by the server through a browser or dedicated application, such as drawing multiple candidate paths on a map control, displaying sales forecast intervals in a chart component, and displaying generated suggestions and sentiment analysis results in a text area.
[0404] - When a user clicks on a specific route, plan, or suggestion, the terminal packages the selected identifier and additional parameters into a request message and sends it to the server via HTTP or WebSocket to achieve low-latency interaction.
[0405] The terminal provides a text input box and a voice input button. Users can directly enter new prompts or record natural speech through the microphone. The terminal uploads the recording data to the server, which then calls the speech recognition and generative artificial intelligence models.
[0406] Users browse system-generated candidate solutions, sentiment reports, and historical records on their terminals, and make final decisions based on their responsibilities. User selections, along with feedback data such as ratings and comments on output quality, are collected by the terminals and transmitted back to the server for subsequent optimization.
[0407] (vii) Adaptive optimization of prompt statements and model running conditions The server not only makes a one-time call to the generative artificial intelligence model, but also uses evaluation and feedback information to achieve continuous optimization: - The server records the mapping relationship between each prompt statement, model output, user selection, and rating in the database.
[0408] The server periodically uses these records to train a lightweight rating prediction model (e.g., a regression or classification model based on text embeddings and fully connected layers) to predict the expected performance of a specific type of prompt statement in the current context.
[0409] - The server dynamically adjusts the conditions for generating prompt statements based on the score prediction results, including: - Whether to increase or decrease the weighting of emotional information; - Should the output structure be changed (e.g., should the advantages and disadvantages be listed in bullet points)? - Whether to change model operating parameters (such as temperature, maximum output length) to balance creativity and stability.
[0410] Through these specific data-driven rules, the server can enable the generative AI model to continuously approach its optimal performance in specific enterprises and meeting cultures, thereby achieving technical improvements to the "model invocation method" and "prompt statement design strategy," rather than relying solely on one-time manual settings.
[0411] (viii) Explanation of technical effects and causal relationship The server achieves the following technical effects through the aforementioned specific data structure and unconventional processing flow: - By jointly modeling multimodal sentiment timelines and meeting minutes, the server makes the input of generative artificial intelligence models more comprehensive and the output more in line with the actual situation, thereby reducing the proportion of manual corrections required and improving the effectiveness of decision-making suggestions; - By automatically performing structured parsing and secondary optimization (path / plan algorithm) on the model output, the server significantly reduces the overhead of converting natural language into control commands, thereby improving overall processing speed and accuracy; - By limiting context length and simplifying the structure of prompt statements, the server reduces the transmission of irrelevant information and the amount of computation required for model inference, thereby reducing network communication load and model computation load; - The server trains a rating prediction model based on historical feedback and adjusts the conditions for generating prompt statements to achieve adaptive optimization of the prompt statement design and model parameters, so that the output quality of the system automatically improves over time without relying on repeated manual debugging.
[0412] These effects arise from the specific technical measures taken by the server within the computer to handle data flow, model invocation, and control command generation, rather than simply transferring human decisions to the machine for execution. Therefore, they represent an improvement to computer technology itself.
[0413] (ix) Other implementation forms and variations The server can be implemented in different ways: - In mobile scenarios, we focus on optimizing path information and operation plans to enable autonomous vehicles, drones and other mobile objects to achieve low-latency and highly robust route adjustment capabilities. - In enterprise operation scenarios, we focus on using a combination of time series forecasting models and generative artificial intelligence models to optimize sales forecasting, risk assessment, and strategy generation in an integrated manner; - In advertising creative scenarios, user feedback is used to iterate on the style and content of prompts, and this is combined with image generation models to form a cross-modal content generation system.
[0414] Terminals can take many forms, such as in-vehicle central control screens, industrial control panels, desktop terminals, or mobile smart terminals, as long as they can perform display, input, and data reporting functions.
[0415] Users are not limited to a single role; they can be drivers, meeting hosts, corporate managers, or front-line operators. The technical effect of this invention is reflected in the data processing and model optimization logic within the computer system it is used in, rather than in a specific business role.
[0416] Through the above-described embodiments and their variations, the system of the present invention can achieve meeting understanding, sentiment analysis, generative artificial intelligence model-driven suggestion generation and control execution in different application scenarios with a unified technical framework, thereby gaining comprehensive technical advantages in terms of processing speed, accuracy, communication load and system adaptability.
[0417] use Figure 14 The processing flow is explained.
[0418] Step 1: The server receives multimodal raw data related to the meeting from the terminal.
[0419] The server's input includes: audio and video data streams uploaded by the terminal, as well as meeting topic information and environmental information (such as the location of the mobile device, its operating status, or its business status) entered by the user through the terminal.
[0420] The server performs preprocessing operations on the input audio data, such as sampling rate unification, format conversion (e.g., conversion to PCM or FLAC), and segmentation, and performs frame extraction and resolution adjustment on the video data, and performs field validation and format standardization on the topic information and environmental information.
[0421] The server outputs standardized audio segment sequences, video frame sequences, and structured meeting agenda and environmental information records for subsequent processing.
[0422] Step 2: The server uses speech recognition technology to convert audio data into text information.
[0423] The server's input is the audio segment sequence output from step 1.
[0424] The server sends each audio segment to the speech recognition service interface via the speech recognition client module. It then parses and reassembles the returned character sequences and timestamps, and splices the recognition results of each segment according to time sequence. During this process, the server marks or re-recognizes segments with low confidence, thereby improving overall recognition accuracy.
[0425] The server output is a sequence of meeting text information with timestamps and optional speaker identifiers, i.e., the initial text data of the meeting transcript.
[0426] Step 3: The server organizes the meeting text information into a meeting record data structure and performs basic natural language preprocessing.
[0427] The server's input is the sequence of meeting text information output in step 2.
[0428] The server performs sentence segmentation, word segmentation, noise removal, and removal of common meaningless words from the input text, sorts the text according to timestamps, and writes the results into the meeting record data structure. The server also extracts key terms from each speech (using word frequency statistics or dependency parsing) and appends them to the corresponding record for subsequent summary and topic extraction.
[0429] The server outputs structured meeting minutes data containing time information, speech content, and key terms.
[0430] Step 4: The server extracts the emotional characteristics of the meeting participants from video and audio data.
[0431] The server's input consists of the video frame sequence and audio segment sequence output from step 1.
[0432] The server uses an image processing module to detect facial regions in video frames, extract facial key points and expression feature vectors, and inputs them into a pre-trained expression classification neural network to calculate the probability distribution of each emotion category. Simultaneously, the server extracts speech features (such as MFCC, pitch, energy, and speech rate) from audio segments and calculates corresponding emotion scores using a speech emotion classification model.
[0433] The server performs a weighted fusion operation on the image sentiment results and the voice sentiment results, for example, by setting weights based on data quality and model confidence, and calculates a comprehensive sentiment vector.
[0434] The server outputs: emotional timeline data for each meeting participant, sorted by time, including emotional category probability and emotional intensity value.
[0435] Step 5: The server combines meeting minutes data and sentiment timeline data to generate a session summary and sentiment profile.
[0436] The server's inputs are: the meeting minutes data output in step 3 and the emotional timeline data output in step 4.
[0437] The server first uses text summarization algorithms (such as extractive or generative summarization models based on Transformer, or key sentence selection based on TF-IDF) to filter out core statements representing the current topic from the meeting minutes. Then, based on the emotional timeline within the same time interval, the server analyzes the emotional changes of participants around the core statements (such as increased tension or dissatisfaction). The server combines this information into a structured conversation summary and emotional profile text.
[0438] The server output is a summary of the meeting context data, including a "core list of speeches" and a "corresponding sentiment summary".
[0439] Step 6: The server generates prompts based on meeting topic information, environment information, session summary, and sentiment summary.
[0440] The server's inputs are: the structured meeting agenda information and environment information output in step 1, and the meeting context summary data output in step 5.
[0441] The server selects a template from the prompt template structure that matches the current task type (e.g., path optimization, business decision-making, advertising creative), and fills the template placeholders with meeting topics, environmental information (e.g., current location, traffic conditions, or business status), core speech summaries, and sentiment summaries. The server controls the length and hierarchical format of the prompt statements according to system settings (e.g., using "Meeting Notes" and "Sentiment Analysis" tags for subsequent analysis by the generative artificial intelligence model.
[0442] The server output is one or more complete prompt texts generated for the current meeting context.
[0443] Step 7: The server will input the prompt statement into the generative artificial intelligence model and obtain the model's output.
[0444] The server's input is the prompt text output in step 6.
[0445] The server calls a generative artificial intelligence model interface (such as a large-scale language model service based on the Transformer architecture), sending the prompt statement as the input sequence, and setting model running parameters (such as temperature, maximum output length, decoding strategy, etc.). The server receives the natural language text output generated by the model, performs basic format validation (checking whether it contains agreed chapter titles, list symbols, etc.), and trims or completes any parts that do not conform to the format.
[0446] The server output is a model output text containing opinions, proposals, and responses that take into account the emotions of the participants regarding the meeting topics.
[0447] Step 8: The server performs structured parsing and secondary optimization on the output text of the generative artificial intelligence model.
[0448] The server's input is the model output text from step 7.
[0449] The server extracts key data from the output text using natural language parsing algorithms (including segmentation, entity recognition, pattern matching, and dependency analysis). This data includes fields such as suggested route name, locations along the route, estimated arrival time, risk description, business adjustment suggestions, and priority. The server then populates these fields into a predefined structured data structure and, when necessary, invokes path planning or scheduling algorithms to perform numerical calculations and optimizations on the suggested solutions (e.g., recalculating the shortest path, comparing the total cost or delay time of multiple solutions).
[0450] The server outputs structured decision data that can be directly used by control or management devices, including path information, planning information, or business strategy information.
[0451] Step 9: The server distributes structured decision data to the control or management device and generates visual information for display.
[0452] The server's input is the structured decision data output from step 8.
[0453] The server encodes structured data into control commands or setting change information according to the target device's communication protocol, and sends it to the corresponding control or management device through the network interface, thereby enabling adjustments to the operating path, modifications to equipment parameters, or updates to the business plan. Simultaneously, the server generates the data formats required for visualization on the terminal, such as converting path nodes into map coordinate sequences, forecast data into chart data series, and risk levels into color labels.
[0454] The server output includes: control commands / configuration change information sent to external devices, and visual data packets for terminal display.
[0455] Step 10: The terminal receives visual data packets and presents the decision results and sentiment analysis results to the user.
[0456] The terminal input is the visual data packet output in step 9.
[0457] The terminal uses a front-end rendering module to display maps, graphs, bar charts, and text descriptions on the screen, showing information such as candidate routes, estimated arrival times, risk point descriptions, and the emotional curves of meeting participants. The terminal provides interactive controls for each candidate solution, such as "Adopt," "Modify," and "Regenerate" buttons, and displays prompts and key sentences from the model's output in appropriate locations to help users understand the basis for the system's recommendations.
[0458] The terminal outputs a graphical user interface and the status of user-operable interactive controls.
[0459] Step 11: Users can make selections or confirm actions based on the content displayed on the terminal, and can also input evaluations and feedback.
[0460] The user's input includes: the candidate solutions displayed on the terminal, the sentiment analysis results, and explanatory text.
[0461] Users can click the "Accept" button for a candidate route or business plan on the terminal interface, or select different priority plans through a drop-down list, and provide a rating or text comment in the rating area. Users can also enter new natural language requests in the text input box, such as "Please provide a more conservative sales forecast only for the East China region."
[0462] The user's output includes: selection results, confirmation instructions, and evaluation or feedback information, which are collected by the terminal.
[0463] Step 12: The terminal uploads the user's selection results and feedback information to the server.
[0464] The terminal input consists of the selection data, evaluation data, and new natural language input generated by the user in step 11.
[0465] The terminal encapsulates this data into a request message and sends it to the server via HTTP or WebSocket, along with a session identifier and timestamp. Upon successful transmission, the terminal can update its local interface status, such as marking the solution as adopted or displaying "Feedback submitted."
[0466] The terminal output is an upload request containing the user's selection results and feedback information.
[0467] Step 13: The server records user feedback and updates the conditions for generating prompts and the operating conditions of generative artificial intelligence models.
[0468] The server's input includes: the user selection results, rating data, text comments, and possible new natural language requirements uploaded in step 12.
[0469] The server establishes a correlation between the feedback data and the corresponding prompts, model outputs, and structured decision data, and writes this information to the database. The server uses these historical records to train or update a rating prediction model or strategy selection model, calculating the expected performance metrics under different prompt structures and different combinations of model parameters. Based on this prediction result, the server adjusts the prompt generation conditions (e.g., increasing / decreasing sentiment information weight, adjusting summary length, modifying output format requirements) and the operating conditions of the generative AI model (e.g., temperature, maximum output length), thereby automatically adopting a more optimal configuration in subsequent calls.
[0470] The server output includes: updated prompt generation conditions, updated model running parameter configurations, and an experience database containing historical records and feedback information, which are used to support the optimization of the next round of processing.
[0471] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0472] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0473] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0474] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0475] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0476] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0477] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0478] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0479] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0480] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0481] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0482] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0483] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0484] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0485] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0486] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0487] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0488] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0489] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0490] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0491] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0492] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0493] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0494] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0495] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0496] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0497] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0498] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0499] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0500] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0501] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0502] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0503] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0504] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0505] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0506] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0507] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0508] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0509] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0510] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0511] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0512] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0513] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0514] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0515] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0516] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0517] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0518] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0519] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0520] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0521] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0522] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0523] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0524] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0525] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0526] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0527] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0528] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0529] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0530] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0531] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0532] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0533] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0534] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0535] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0536] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0537] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0538] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0539] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0540] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0541] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0542] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0543] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0544] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0545] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0546] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0547] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0548] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0549] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0550] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0551] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0552] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0553] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0554] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0555] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0556] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0557] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0558] In addition, the following notes are provided in response to the above explanation.
[0559] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring information related to meeting topics and meeting content, and recording the meeting information as text-based meeting information; A device for performing natural language processing algorithms on the text-based conference information to perform summary processing and key word extraction processing, thereby generating summary information of the conference information and key word information of the conference information; An apparatus for instructing a generative artificial intelligence model to generate prompt statements for meeting opinions, based on the summary information and the key word information, combined with meeting topic information; An apparatus for acquiring voice data from a voice acquisition device, converting the voice data into text data using a voice recognition algorithm, and integrating the text data into the meeting information as text-based meeting information; An apparatus for analyzing the emotional information of meeting participants based on information related to the meeting participants contained in the meeting information and the voice data, using an emotion inference algorithm, and reflecting the emotional information in the prompt statements and / or the generated meeting opinion content; An apparatus for inputting the prompt statement into the generative artificial intelligence model, enabling the generative artificial intelligence model to perform data processing and data operations to generate opinion information in the meeting, and organizing the generated opinion information into structured information; A device for sending the structured information to a terminal device and presenting the opinions in a form that facilitates multi-faceted and rapid decision-making by meeting participants.
[0560] (Note 2) According to the information processing system described in Appendix 1, the device for acquiring meeting information includes: an input device for inputting meeting content via an operation interface from an information management device, and / or an acquisition device for automatically acquiring meeting content from voice data during the meeting using a voice acquisition device and a voice recognition algorithm.
[0561] (Note 3) The information processing system according to Appendix 1 is characterized in that it further includes: an output device for generating meeting outcome information based on the generated opinion information and the meeting information, and outputting the meeting outcome information to an external information providing device, and providing the meeting outcome information to an external subject through an information publishing device, so as to generate topicality and promote business results.
[0562] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for generating prompts that instruct a generative artificial intelligence model to generate opinions on meeting topics and content about production plans, based on meeting topics and the operational status of production objects. A device for acquiring operational data from an information acquisition device installed on the production object, converting the operational data into standardized time series data, and using the time series data as input feature quantities for the generative artificial intelligence model; A device for acquiring voice or image data related to meeting participants, converting the voice or image data into text data and emotional information, and combining the emotional information with the prompts and input features and inputting it into the generative artificial intelligence model; A device for generating opinions on meeting topics and candidate production plans, including time allocation, target objects, target quantities, and resource allocation for each production line, based on the generative artificial intelligence model; An apparatus for sending the candidate production plan to an information notification device, receiving correction operations from the user, and determining the final production plan based on the correction operations; An apparatus for generating instruction information for a control device based on the final production plan, and for controlling the operation of the production object through the control device.
[0563] (Note 2) The information processing system according to Appendix 1 is characterized in that, It includes a device for automatically acquiring meeting content and the operational status of production objects through input means of information management departments, or through voice recognition technology and operational data acquisition technology, and automatically generating the prompt statements and input feature quantities based on the acquired information.
[0564] (Note 3) The information processing system according to Appendix 1 is characterized in that, It includes a device for outputting opinions on meeting topics, production plans, and evaluation information related to their execution results generated by the generative artificial intelligence model to an external information providing device, and providing the evaluation information in a form that can be used to disseminate information about the organization's operations and generate business results.
[0565] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring voice information related to a meeting through the voice input or recording function of a terminal, and sending the acquired voice information to an information processing device via a predetermined communication path; A device for converting speech information into text information using speech recognition in the information processing apparatus, and for integrating the text information in chronological order or speaking order to store it as meeting record information; In the information processing device, at least a portion of the meeting record information is extracted from the meeting record information, and the extracted meeting record information is combined with prompt statements composed of instruction text defined according to the meeting agenda or decision-making object to form an input information device for inputting to a generative artificial intelligence model. A device for inputting the input information into the generative artificial intelligence model in the information processing apparatus, causing the generative artificial intelligence model to perform semantic parsing processing, and generating response information for indicating opinions or proposals in a meeting based on the result of the semantic parsing processing; A device for performing structured processing on the response information in the information processing apparatus and organizing it into display information containing at least some key meeting points, time-setting related suggestion information, and risk or resource-related suggestion information; A device for sending the display information to the terminal and presenting the display information to the conference participants through the display function of the terminal.
[0566] (Note 2) The information processing system according to Appendix 1 is characterized in that, In the information processing device, sentiment estimation processing is performed on the speech content contained in the meeting record information to generate sentiment information that indicates the emotional state of the meeting participants, and the sentiment information is considered as a condition for generating the input information or the response information, thereby controlling the input to the generative artificial intelligence model or the structured processing of the response information.
[0567] (Note 3) The information processing system according to Appendix 1 is characterized in that, In the information processing device, at least a portion of the meeting record information and the response information are stored as information for external distribution. When predetermined disclosure or distribution conditions are met, the information for external distribution is distributed to external information processing infrastructure or user terminals to release the meeting results or the opinions generated.
[0568] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for generating prompts that instruct generative artificial intelligence models to generate opinions and proposals based on meeting topic information and relevant environmental information. An apparatus for acquiring audio information from a meeting, converting the audio information into text information using speech recognition technology, and organizing the text information in chronological order for storage as meeting record data. An apparatus for inferring the emotional state of a participant based on the meeting record data and the image and voice features of the participant, using sentiment analysis technology, and calculating the emotional state as time-varying information; A device for embedding the meeting record data and the emotional state-related information into the prompt statement and inputting it into the generative artificial intelligence model, and obtaining output information from the generative artificial intelligence model containing opinions on the meeting topics and responses that take into account the emotional state of the participants. A device for extracting path information, planning information or business policy information related to the operation or business from the output information and converting it into structured data that can be applied as an operation plan or business plan for the operation or business. A means for visually presenting the structured data and the output information on a display device or a mobile device, and providing a user interface for receiving selection or confirmation operations from a user. A device for outputting control commands or setting change information to a control device or management device that is an object of operation or a business based on the selection operation or confirmation operation, so as to perform operation management or business management; An apparatus for storing the output information and evaluation or feedback information from the user, and for updating the generation conditions of the prompt statement or the operating conditions of the generative artificial intelligence model based on the evaluation or feedback information.
[0569] (Note 2) According to the information processing system described in Appendix 1, the meeting environment includes a meeting space set within the mobile body, the meeting agenda information and the meeting record data are associated with the mobile body's operation information, location information or traffic information, and the structured data is generated as operation management information for optimizing at least one of the mobile body's operation path, arrival time or operation conditions.
[0570] (Note 3) According to the information processing system described in Appendix 1, the output information and the operation management information or business management information are notified to external subjects through an information providing device or communication device, and are provided in a way that enables external subjects to know and generate topics related to the fact that the generative artificial intelligence model participates in the meeting and decision-making.
Claims
1. An information processing system, characterized in that, include: processor, The processor is configured to: Based on the meeting's agenda, prompts are generated to instruct generative AI models to generate opinions. Acquire audio data and convert the audio data into text data; Analyze the emotions of the meeting participants and generate opinions that take into account those emotions.
2. The information processing system according to claim 1, characterized in that, The processor is configured to input meeting content through the information management department or to automatically acquire meeting content using speech recognition technology.
3. The information processing system according to claim 1, characterized in that, The processor is configured to release the generated opinions and meeting outcomes to generate buzz and achieve marketing results regarding the innovative initiative.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A