Information processing system
Patent Information
- Application Number
- CN202610319296.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-16
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]在现有技术中,体育赛事解说主要依赖人工解说员或预先设定的固定解说脚本,存在以下问题:其一,解说内容难以根据单个用户的兴趣、知识水平和情绪状态进行动态个性化调整,用户只能被动接收统一的解说信息,交互性和沉浸感不足;其二,现有虚拟角色或简单动画角色即使能够叠加在体育赛事画面上,其动作往往是预先编排的,与实际比赛画面和具体运动项目的细节关联度较低,难以及时对应不同体育项目及其复杂多变的比赛情景;其三,现有系统普遍缺乏能够结合生成式人工智能模型进行自然语言问答与连续对话的能力,用户在观看赛事过程中提出的问题和评论,往往无法得到即时、内容丰富且贴合上下文的回应;其四,即使某些系统能够记录部分用户数据,也缺乏有效的学习机制来基于用户的历史行为和偏好持续优化解说风格和互动内容,从而难以形成长期、个性化的陪伴式解说体验
[0005]为了解决上述课题,本发明提供了一种信息处理系统,该系统包括处理器,其中,所述处理器被配置为:向用户提供用于设置用户偏好的虚拟角色的用户界面,使用户能够在移动终端上选择和配置所希望的虚拟角色;当用户利用携带信息终端对准体育比赛现场或体育比赛的影像时,对采集到的影像数据进行解析,识别当前体育项目和比赛情景,并基于所述解析结果生成与当前情景相匹配的虚拟角色动作;使虚拟角色参照体育动作数据库,从包含多个不同体育项目的动作数据中,依据为各体育项目和动作类型预先设定的索引,自动选取与当前运动情景相对应的动作数据,从而使虚拟角色在不同体育项目和不同比赛事件中均能呈现合适的动作表现;对来自用户的输入进行解析,将用户提出的与体育相关的问题或评论抽取为语义结构,并据此生成用于指示生成式人工智能模型生成响应的提示文本,使生成式人工智能模型能够在考虑当前比赛情景、用户输入内容和既有上下文的基础上,生成自然、连贯且内容丰富的问答或对话响应;进一步地,处理器还被配置为识别用户在互动过程中的情绪状态,并根据所识别的情绪对虚拟角色的解说风格进行调整,包括但不限于语气、用词和解说的详略程度,然后利用生成式人工智能模型基于所述提示文本生成与当前情绪及偏好相一致的解说内容和回复文本;此外,处理器还被配置为使虚拟角色通过与用户的多轮交互进行学习,记录用户的过去行为和偏好信息,并通过学习算法对所记录的数据进行分析,从而在后续的比赛解说和问答互动中,自动选择更符合用户兴趣和理解水平的解说角度、解说深度及互动方式,实现对用户的个性化体育赛事解说与持续优化的互动体验。
Smart Images

Figure CN122799323A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] In existing technologies, sports commentary mainly relies on human commentators or pre-set fixed commentary scripts, which presents the following problems: First, the commentary content is difficult to dynamically and personally adjust according to the individual user's interests, knowledge level, and emotional state. Users can only passively receive uniform commentary information, resulting in insufficient interactivity and immersion. Second, even if existing virtual characters or simple animated characters can be superimposed on sports event footage, their movements are often pre-choreographed, with low correlation to the actual match footage and the details of specific sports, making it difficult to respond promptly to different sports and their complex and ever-changing match scenarios. Third, existing systems generally lack the ability to combine generative artificial intelligence models for natural language question answering and continuous dialogue. Questions and comments raised by users during the match often do not receive immediate, rich, and context-appropriate responses. Fourth, even if some systems can record some user data, they lack effective learning mechanisms to continuously optimize the commentary style and interactive content based on users' historical behavior and preferences, making it difficult to create a long-term, personalized, companion-style commentary experience.
[0004] In view of the above problems, it is necessary to provide an information processing system that can automatically analyze the images of sports competitions or videos captured on mobile terminals to generate virtual character actions, and link with a motion database containing multiple sports and a generative artificial intelligence model to achieve personalized sports commentary services with scene-matched action presentation, intelligent question answering, and style adaptation for individual users, thereby effectively improving the viewing experience and interactivity. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to: provide a user interface for setting user-preferred virtual characters, enabling the user to select and configure desired virtual characters on a mobile terminal; when the user uses a portable information terminal to view images of a sports event or a sports competition, analyze the acquired image data, identify the current sports event and competition scenario, and generate virtual character actions matching the current scenario based on the analysis results; automatically select action data corresponding to the current sports scenario from a sports action database containing action data of multiple different sports, according to pre-defined indexes for each sports event and action type, thereby enabling the virtual character to present appropriate action performance in different sports events and competition events; and analyze user input, extracting sports-related questions or comments from the user into semantic structures, and generating instructions based on these structures. Generative AI models generate prompt text for responses, enabling them to produce natural, coherent, and content-rich question-and-answer or dialogue responses based on the current match context, user input, and existing context. Furthermore, the processor is configured to recognize the user's emotional state during interaction and adjust the virtual character's commentary style accordingly, including but not limited to tone, word choice, and level of detail. The generative AI model then generates commentary content and response text consistent with the current emotion and preferences based on the prompt text. Additionally, the processor is configured to enable the virtual character to learn through multiple rounds of interaction with the user, recording the user's past behavior and preferences. The learned data is then analyzed using learning algorithms to automatically select commentary angles, depths, and interaction methods that better match the user's interests and comprehension levels in subsequent match commentary and Q&A interactions, achieving personalized sports commentary and a continuously optimized interactive experience.
[0006] "System" refers to the overall technical configuration including at least one processor and optional storage devices, communication interfaces and user terminals, used to perform the functions of virtual character setting, image analysis, motion generation, question-and-answer interaction and personalized narration described in this invention.
[0007] A processor is a hardware unit that can execute program instructions and perform calculations, control, and logical judgments on input data. It can be a single CPU, GPU, NPU, DSP, ASIC, FPGA, or any combination thereof, or a computing resource deployed on a local terminal or a remote server.
[0008] The "user interface" refers to the interface presented by the system to the user for information display and input interaction, including graphical user interfaces, touch interfaces, voice interfaces, menu interfaces, etc. Through this interface, users can set virtual character preferences, enter questions or comments, and view system feedback.
[0009] "Virtual character" refers to a computer-generated character image presented in a display interface or augmented reality scene, including but not limited to two-dimensional characters, three-dimensional characters, virtual anchors (Vtuber), etc. The character can have a visual appearance, actions, expressions, and the ability to interact with the user through language.
[0010] "Portable information terminal" refers to a terminal device that a user can carry with them and that has information processing and communication functions, including but not limited to smartphones, tablets, wearable devices, etc. This terminal can capture sports competition images and run the applications related to this invention.
[0011] "Sports competition venue" refers to the physical location where a sports event actually takes place, such as a football field, basketball court, or tennis court. Users can use their mobile devices to take pictures or view the event at the venue.
[0012] "Sports competition footage" refers to video footage of sports events acquired through cameras, television, online video, etc., and displayed on portable information terminals, including live footage and recorded footage.
[0013] "Image data analysis" refers to the process by which a processor processes and analyzes image and video data acquired from cameras or video sources to identify the type of sport, scene elements, competition events, and other information related to the sports context.
[0014] "Virtual character actions" refers to the dynamic performance of virtual characters in the display screen or AR scene, including limb movement, posture changes, facial expression changes, and position changes. Its data can be represented by skeletal animation, keyframe animation, or other motion description data.
[0015] The "Sports Action Database" refers to a database that stores action data sets corresponding to various sports and their related competition events. This database organizes action resources that virtual characters can call upon by dimensions such as sports and action types.
[0016] "Motion data" refers to the digital data used to drive virtual characters to perform specific actions, including skeletal animation data, pose sequences, keyframe parameters, facial animation data, and motion-related metadata.
[0017] An "index" refers to the identifier and structure used to retrieve, locate, and select motion data corresponding to a specified sport and motion type in a sports motion database, including but not limited to key values, labels, index tables, or multidimensional index structures.
[0018] "User input" refers to various types of information provided by users to the system through the user interface, including text input, voice input (recognized text), touch selection, button operation, and other interactive instructions that can be parsed by the system.
[0019] "Sports-related question and response" refers to the interactive process in which the system provides answers or feedback in natural language when users ask questions or express opinions about sports competitions, sports rules, tactical analysis, player information, etc.
[0020] "Generative AI models" refer to AI models that can automatically generate text, speech, or other content output based on input prompts, including but not limited to large-scale language models, text generation models, and multimodal generation models.
[0021] "Prompt text" refers to text instructions or descriptions generated by the processor based on user input, the current sports scenario, and contextual information, used to guide generative artificial intelligence models to produce the desired response content.
[0022] "User emotions" refers to the emotional state expressed by users during the interaction between users and the system, such as excitement, tension, disappointment, happiness, calmness, etc. These emotions can be identified through voice features, text tone, facial expressions, or other signals.
[0023] "Commentary style" refers to the expression and performance characteristics adopted by virtual characters when commentating on sports events, including the strength of tone, emotional color, level of humor, frequency of use of professional terms, depth of commentary, and rhythm.
[0024] "Learning algorithms" refer to algorithms used to model, analyze, and update users' historical behavior and preference data, including but not limited to machine learning algorithms, deep learning algorithms, reinforcement learning algorithms, and rule-based learning mechanisms.
[0025] "Users' past behavior and preferences" refers to information recorded during long-term interactions between users and the system, including users' historical operations, questions asked, evaluations, choices made, dwell time, and inferred interests, knowledge levels, and style preferences.
[0026] Personalized sports commentary refers to a sports commentary service that automatically adjusts the content, style, level of detail, and interaction methods of the commentary based on the user's past behavior and preferences as well as the current context, so that the commentary better meets the specific needs and interests of the user. Attached Figure Description
[0027] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0028] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0029] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0030] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0031] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0032] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.
[0033] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0034] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0035] Figure 9 This represents an emotion map that maps multiple emotions.
[0036] Figure 10 This represents an emotion map that maps multiple emotions.
[0037] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0038] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0039] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0040] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0041] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0042] First, let me explain the terminology used in the following instructions.
[0043] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0044] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0045] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0046] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0047] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0048] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0049] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0050] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0051] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0052] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0053] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0054] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0055] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0056] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0057] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0058] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0059] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0060] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0061] With the development of generative artificial intelligence models and image recognition technology, users' demands for interactive and personalized commentary while watching competitive events are constantly increasing. However, existing technologies suffer from the following problems: First, the recognition results of the game screen on the terminal side and the commentary generation on the server side are usually processed separately, lacking unified context modeling. This results in a lack of synchronization and consistency between the virtual character's actions and the language commentary, leading to a poor user experience. Second, existing systems mostly use fixed scripts or templates for commentary, failing to dynamically adjust the commentary style based on real-time competitive status, statistical data, and the user's current emotions, making it difficult to achieve truly personalized and adaptive interaction. Third, although generative artificial intelligence models can generate natural language, in actual systems, prompts are often statically and coarsely constructed, failing to comprehensively utilize game status information, virtual character attributes, user historical preferences, and multi-turn dialogue context, resulting in insufficient professionalism and poor coherence in the generated results. Fourth, the collaboration between the user terminal and the server is often limited to simple data uploading and result distribution, lacking a learning mechanism based on structured competitive status and user behavior data, making it impossible to gradually improve dialogue quality and system performance through long-term use. Therefore, how to improve the relevance, consistency, and personalization of the narration content by constructing structured prompts, utilizing generative artificial intelligence models, and combining user emotion and preference information within the collaborative framework of information processing devices and external devices, thereby achieving an overall improvement in computer technology (including human-computer interaction, content generation, state recognition, and learning algorithms), has become a pressing technical issue in this field.
[0062] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0063] In this invention, the server includes: a processing unit for receiving competition status information uploaded by an information processing device and natural language query information input by a user, and processing it in association with action information corresponding to various competition categories and competition scene categories stored in an external device; a retrieval unit for retrieving and specifying action information corresponding to the competition status information from the external device based on the competition status information, and returning response information containing the action information to the information processing device; a generation unit for generating prompt statements for inputting into a generative artificial intelligence model based on the query information, the competition status information, and attribute information of a virtual display object, and inputting the prompt statements along with statistical information into the generative artificial intelligence model to generate response text information; an emotion and preference analysis unit for inferring the user's emotional state based on the user's facial expression information, voice information, and operation information, and dynamically adjusting the speaker setting and expression style in the prompt statements based on the emotional state and the user's preference information; and a learning unit for recording and updating the user's past query information, viewing history information, and operation history information, and integrating the summarized historical information into the prompt statements, so that the generative artificial intelligence model can generate consistent response text information in multi-turn dialogues. This allows for the formation of a unified context construction and dynamic prompt generation mechanism on the server side for generative artificial intelligence models. It structurally integrates image recognition results from the terminal side, action information from external devices, real-time user emotions, and long-term preference data. This not only improves the professionalism and coherence of the response text but also enables the action control and language output of virtual display objects to be highly coordinated at the system level, thereby substantially improving the computer technology performance of content generation and scene understanding in human-computer interaction systems.
[0064] "Information processing device" refers to an electronic device with computing and storage functions, used to execute programs to perform image processing, data communication, interface display and data recording, etc., and may include terminal devices, server devices or any combination of these devices.
[0065] "External device" refers to a data processing device or storage device that is connected to an information processing device via communication, used to store motion information, statistical information or other data related to competitive activities, and to return response information according to requests from the information processing device.
[0066] "User" refers to a person who watches the game, asks questions, sets the attributes of virtual display objects, and interacts with the system by operating the user interface of the information processing device.
[0067] "Virtual display objects" refer to digital characters or graphic objects presented on a display device in the form of images or three-dimensional models. These objects can change their movements and expressions based on the action information, narration information, and expression styles generated by the system, and serve as the main body for interaction with the user.
[0068] "Attribute information" refers to a set of predefined parameters related to virtual display objects, including character type, tone style, professional level, performance tendency, etc., which are used to control the narration method and performance form of virtual display objects.
[0069] "Dialogue style" refers to the combination of linguistic expression characteristics, such as tone, wording, politeness, humor, and professionalism, used by virtual display objects when interacting with users.
[0070] "Imaging device" refers to image acquisition hardware used to acquire image information related to competitive activities, including cameras, video acquisition devices, or other devices that can generate digital image signals.
[0071] "Image information" refers to still images or dynamic video frame data acquired by an imaging device and represented in digital form, used to reflect visual content such as the arena, participants, and scoreboards.
[0072] "Image processing algorithm" refers to the program logic that performs operations such as preprocessing, feature extraction, target detection, and region segmentation on image information, in order to provide structured data for subsequent recognition processing.
[0073] "Recognition algorithm" refers to an algorithm that judges and classifies targets or scenes in an image based on the results of image processing. It includes target recognition, text recognition, scene classification, etc., and is used to infer state information related to competitive activities.
[0074] "Competitive status information" refers to structured data obtained by analyzing image information using image processing and recognition algorithms. It is used to represent elements related to the current competitive activity status, such as the category of the competitive event, score, time, scene type, and key participants.
[0075] "Motion information" refers to a set of parameters used to control the physical movements, posture changes, and facial expressions of virtual display objects. These parameters include skeletal animation parameters, facial expression control parameters, motion duration, and motion type identifiers.
[0076] "Information set" refers to a combined dataset stored in an external device that contains motion information and its index information associated with multiple competition categories and competition scenario categories.
[0077] "Response information" refers to the structured data generated and returned to the information processing device by the external device after receiving the competition status information from the information processing device, based on the retrieved action information, to guide the display control of the virtual display object.
[0078] "Natural language query information" refers to text or voice content in which users input questions in natural language, targeting competitive activities or related objects.
[0079] "Prompt statements" refer to text strings generated by information processing devices and used as input to generative artificial intelligence models. These include descriptions of inquiry information, competition status information, virtual display object attribute information, and historical information, as well as settings for the model's role, task, and response style.
[0080] "String information" refers to data represented in the form of a sequence of characters, used to form prompts or other text content, which can be parsed and processed by generative artificial intelligence models or other program modules.
[0081] "Generative artificial intelligence models" refer to data processing models trained through machine learning that can generate natural language text or other content based on input prompts and related context. These models typically include neural network models.
[0082] "Statistical information" refers to a set of quantitative data related to competitive activities, including numerical or categorical data such as the number of points scored, shots taken, assists made, fouls committed, and playing time, which can be used to analyze and generate commentary content.
[0083] "Response text information" refers to the natural language text content generated by generative artificial intelligence models based on prompts and input data, used to answer users' inquiries.
[0084] "Output data" refers to data obtained by formatting or multimodal fusion of response text information and action information, which can be directly used by display devices or audio devices, including text data, audio data, animation control data, etc.
[0085] "Face information" refers to data obtained through image acquisition or other sensing methods that reflects the facial expression characteristics of a user, including the location of facial key points and expression category labels.
[0086] "Voice information" refers to audio data acquired by audio acquisition devices such as microphones, which reflects the characteristics of a user's speech content, speaking speed, pitch, volume, etc.
[0087] "Operation information" refers to the data generated when a user performs input operations on an information processing device, including records of interactive behaviors such as touch operations, key input, swipe gestures, and click frequency.
[0088] "Emotional state" refers to the user's current emotional tendency inferred from facial expressions, voice information, and operational information, including emotional categories such as excitement, calmness, tension, and disappointment.
[0089] "Preference information" refers to characteristic information that reflects a user's long-term tendencies, inferred from the user's historical behavioral data, including preferences for competitive sports, preferred commentary styles, and types of objects of interest.
[0090] "Viewing history information" refers to recorded data about users' past viewing of competitive activities or related content, including viewing time, viewed items, dwell time, and interaction frequency.
[0091] "Operation history information" refers to the time sequence record of user operations such as input, selection, questioning, and feedback in the system, which is used to reflect user interaction habits and points of interest.
[0092] "Learning processing" refers to the process by which an information processing device updates preference information and dialogue strategies based on the user's past inquiry information, viewing history information, and operation history information, including steps such as feature extraction, model training, or parameter updating.
[0093] "Dialogue history information" refers to the time-series text or structured record of multiple rounds of question-and-answer content and corresponding system responses that have occurred between the system and the user, which is used to support the consistent generation of subsequent rounds of dialogue.
[0094] "Historical information" refers to summary data obtained by summarizing or compressing dialogue history information and preference information, which can be used to construct prompts to maintain the continuity and personalization of the dialogue.
[0095] In this embodiment of the invention, the system mainly consists of a server, a terminal, and a display and input device operated by the user. The components interact with each other via wired or wireless communication networks. The server performs centralized data processing and generative artificial intelligence model inference, while the terminal performs image acquisition, local image analysis, virtual object rendering, and user interaction.
[0096] The server runs on one or more computing devices with processors and memory, such as computer devices running a general-purpose operating system. The server may employ a hardware architecture based on a central processing unit and a graphics processing unit (GPU), with the GPU suitable for performing large-scale matrix operations to support inference in generative artificial intelligence models. At the application layer, the server can execute service programs developed using scripting languages or general-purpose programming languages. These service programs interact with the terminal via network communication protocols, displaying game status information, prompts, and response text.
[0097] The terminal is a portable information processing device, such as a smartphone or tablet running a mobile operating system. The terminal includes an imaging device, a display device, an audio output device, a touch input device, and a local processor. The terminal runs a dedicated application locally, calling an image acquisition interface to obtain image information related to the competition, and calling a graphics rendering interface to implement the animated display of virtual objects.
[0098] The server stores the generative AI model in storage. This generative AI model can be a neural network model based on the Transformer architecture. The server sets up a multi-layer self-attention encoder and a multi-layer decoder in the model structure, with each layer including a multi-head attention sublayer and a feedforward network sublayer. During the training phase, the server pre-trains using a corpus containing text from multiple domains and further fine-tunes it using datasets related to sports commentary. During training, the server uses the cross-entropy loss function as the error function and updates the model weights based on backpropagation and gradient descent optimization algorithms. The server can employ data augmentation techniques during training, such as paraphrasing and sentence transformation of the input text, to improve the model's robustness to diverse query information.
[0099] During runtime, the server encodes prompts, competitive status information, and statistical information into a sequence of input to the model. In the encoding phase, the server maps the text to discrete tokens, converts these tokens into vector representations through an embedding layer, and incorporates positional encoding. The server calculates attention weights between different tokens in a multi-layered self-attention module to capture the dependencies between the query information and the competitive status information. In the decoding phase, the server generates the tokens for the response text information sequentially, using decoding strategies such as bundle search to improve accuracy while maintaining fluency. After generation, the server reconstructs the output token sequence into natural language text.
[0100] The server also stores index information and action information sets of external devices in its storage device. These external devices can be independent data storage servers or logical partitions within the server. The server assigns a unique identifier to each competition category and competition scene category in the index information, which includes competition item tags, scene tags, and references to the corresponding action information records. Upon receiving competition status information uploaded by the terminal, the server searches the index structure based on the competition category and scene category information contained within the competition status information. A hash table or tree structure can be used to speed up the retrieval. After matching the corresponding index, the server reads records from the action information set, including skeletal animation parameters, facial expression control parameters, and action duration, and encapsulates these records into response information to return to the terminal.
[0101] The terminal stores the model and texture data of the virtual display object in its local storage. It loads the skeletal structure of the virtual display object using a 3D rendering engine (e.g., an engine based on a general graphics API). Upon receiving motion information from the server, the terminal maps the skeletal animation parameters in the motion information to the skeletal nodes of the virtual display object, and interpolates the rotation and displacement of each node according to a specified time sequence to generate smooth motion trajectories. For facial expression control, the terminal adjusts the deformation of features such as the eyes, mouth, and eyebrows by controlling the weight parameters on the virtual display object's facial mesh to achieve different emotional expressions.
[0102] The terminal continuously acquires image information of the arena through the camera interface provided by the operating system. The terminal inputs this image information to the local image processing module. The local image processing module can perform preprocessing operations using an image processing library, including resizing, color space conversion, and noise suppression. Based on this, the terminal calls the recognition algorithm module, which may include a lightweight convolutional neural network, to detect arena lines, target objects, and scoring information superimposed on the image. The terminal constructs a data structure for the arena's status information based on the detection results. This data structure includes fields such as arena category, score, time, scene type, and key participant identifiers.
[0103] The terminal sends its competition status information to the server via the communication unit. The server, at the receiving end, uses a message parsing module to parse and verify the information, ensuring field integrity. After parsing, the server inputs the competition status information into the retrieval module and the prompt generation module. By uniformly receiving competition status information on the server side, the server can utilize centralized algorithms to aggregate and compare status data from multiple terminals, thereby optimizing the global consistency of action information selection and commentary content generation.
[0104] Users configure the attributes and dialogue style of virtual display objects through the user interface on the terminal. The terminal displays multiple candidate roles and their corresponding attribute options in the user interface module. After the user selects a specific virtual display object on the touch screen, the terminal generates configuration data containing information such as role identifier, tone style, and level of expertise. The terminal saves the configuration data locally and synchronizes it to the server. When generating prompts, the server incorporates this attribute information as part of the prompt content, enabling the generative AI model to adopt the tone and role settings desired by the user when responding.
[0105] During the competition, users can ask questions to the virtual display object via a text input module. After receiving the natural language query information, the terminal combines this information with the current competition status and virtual display object attribute information, packages it, and sends it to the server. The server uses a templated approach to construct the text in the prompt generation module, including a description of the current competition background, the original text of the user's question, and instructions for the model character. For example, the server can generate the following Chinese prompt: "The user is watching a football match. Current match information: Home team vs. Away team, score 1-0, match time 32 minutes. Key player identifier: player_9."
[0106] The question asked by the user: How many assists has this player had this season? You are a football data analysis expert and a virtual coach selected by the user. You need to answer questions using professional but easy-to-understand language based on the statistical data (number of shots, number of goals, number of assists, etc.) provided in the database.
[0107] Please answer from the perspective of a virtual coach, first providing specific data, then briefly analyzing the level of that data in the league. When generating prompts, the server can also insert summaries of the dialogue history and user preference information into the text to maintain consistency across multiple turns of conversation. By structurally incorporating competitive status information, statistical information, and user preference information into the prompts, the server explicitly guides the generative AI model to focus on key features, reducing the model's search space in irrelevant directions and improving the relevance and accuracy of the generated content. This specific method of constructing prompts is an improvement over traditional simple natural language input, allowing the generative AI model to utilize more refined context during computation, which technically improves the efficiency of model inference and the quality of output.
[0108] In terms of user emotion analysis, the server receives facial expression, voice, and operational information from the terminal. The terminal can preprocess the raw sensor data using a local facial expression recognition model and voice analysis module, sending the results to the server in the form of emotion feature vectors. In the emotion and preference analysis unit, the server inputs these feature vectors into a multilayer perceptron-based classifier to determine the user's emotional state. The server adjusts the dialogue style of the prompts based on the emotional state; for example, when the user is tense, the technical density of the explanation is reduced and reassuring language is increased; when the user is excited, the passionate expression of the explanation is enhanced. Through this algorithm-driven dynamic adjustment, the system can adaptively control the output style of the generative artificial intelligence model without increasing the user's explicit input burden, improving the human-computer matching degree and response quality of the interactive system.
[0109] In terms of learning and processing, the server aggregates and models users' historical query information, viewing history information, and operation history information. The server can use sequence models or clustering algorithms to extract key patterns from user behavior sequences, such as preferred sports, preferred data types of queries, and preferred explanation styles. The server encodes these patterns into preference vectors and maps these preference vectors to natural language prompts when generating prompts. This approach allows generative AI models to automatically assign higher weights to statements representing user preferences within their internal attention mechanism when processing prompts, thereby achieving long-term personalized effects in dialogue. Furthermore, the server uses incremental learning when updating preference vectors, updating only the subset of parameters corresponding to new data, thus reducing computational burden and improving online update efficiency.
[0110] Through the above structure, this invention does not simply replace human narration with a computer, but rather establishes a collaborative working mechanism among data structure design, prompt generation, model architecture and training, terminal image parsing, and server retrieval algorithms. The server injects multi-source information into the input of the generative artificial intelligence model through structured prompts, making the model's internal representation learning more closely resemble the competitive scenario. The terminal, through local image recognition, data compression, and state extraction, converts high-dimensional visual information into compact competitive state information, reducing network transmission load. The server improves the response speed of virtual display object action control through index optimization and action information retrieval. These technical means, overall, improve the system's computational efficiency, content generation accuracy, and interaction coherence when processing multimodal data, thereby achieving specific technical effects at the computer technology level, such as increased processing speed, reduced error rate, and optimized data management.
[0111] use Figure 11 The processing flow is explained.
[0112] Step 1: The user launches a dedicated application on the terminal and completes the login process. The input consists of the account identifier and authentication information entered by the user on the login screen, and the output is a session token stored locally on the terminal. The terminal packages the account identifier and authentication information into an authentication request and sends it to the server over the network. Upon receiving the request, the server compares the account and authentication information in the user database, performs hash verification and permission checks, generates a session token, and returns it. The terminal receives the session token, writes it to local storage, and appends the token to the header fields of all subsequent requests.
[0113] Step 2: After obtaining a valid session token, the terminal requests a list of virtual display objects from the server. The input is the session token, and the output is a set of attributes for multiple virtual display objects built in the terminal's memory. Upon receiving the request, the server reads the identifier, basic attributes, and optional dialog styles of each virtual display object from the configuration data store, serializes these fields into structured data, and sends it to the terminal. The terminal parses this data and presents it in the user interface as a list of icons and text.
[0114] Step 3: Users select a target virtual display object and set the narration style in the terminal interface. Input consists of user clicks and option selections on the interface; output is attribute configuration data containing fields such as virtual display object identifier, tone type, and level of professionalism. The terminal receives user action events, maps selected options to enumerated values or numerical parameters, combines them into a configuration structure, writes it to local storage, and sends it to the server via the network. The server associates this configuration with the user identifier and stores it in the configuration database for later use when constructing prompts.
[0115] Step 4: The terminal activates the imaging device to acquire image information of the competition scene. The input is a series of image frames output by the imaging device, and the output is a series of preprocessed image matrices. The terminal calls the camera interface provided by the operating system to retrieve image data at a preset resolution and frame rate. It performs size scaling, color space conversion, and normalization operations on each frame, converting the original pixel array into a tensor format suitable for subsequent recognition models, and caches it in local memory.
[0116] Step 5: The terminal uses image processing and recognition algorithms to extract competition status information from image information. The input is a preprocessed sequence of image tensors, and the output is competition status information including competition category, score, time, scene type, and key participant identifiers. The terminal inputs the image tensors into a local convolutional neural network model, calculates feature maps, and outputs the field type and competition category through a classification head; the terminal combines edge detection and line segment detection algorithms to identify field lines to determine the direction of attack; simultaneously, the terminal calls the text recognition module to extract scoreboard numbers from the image and converts the obtained values into a score field; the terminal encapsulates the above results into a structured data object.
[0117] Step 6: The terminal sends its competition status information to the server to retrieve action information. The inputs are the competition status information and a session token; the output is the action information of a virtual display object available locally on the terminal. Upon receiving the request, the server parses the competition status information, reads the competition category and scene type fields, and performs a key-value lookup or tree search in the index structure based on these fields to locate the corresponding action record. The server combines the skeletal animation parameters, facial expression parameters, and action duration from the action record into a response data structure and returns it. The terminal receives this structure and caches it in the rendering module for subsequent use.
[0118] Step 7: The terminal drives the virtual display object to display based on the motion information returned by the server. The input is motion information and locally cached virtual display object model data, and the output is the animated image presented on the display device. The terminal loads the skeleton and mesh of the virtual display object in the 3D rendering engine, inserts the bone rotation and displacement sequences from the motion information into the animation trajectory, performs interpolation calculations on each joint node to generate continuous frame poses; at the same time, the terminal adjusts the weights of facial control points according to expression parameters to generate corresponding expression deformations, and outputs the rendering results to the screen.
[0119] Step 8: During viewing, users input natural language queries through the terminal. The input is a natural language string typed by the user in the text input box, and the output is a question-and-answer request data containing the query information and the current competition status information. The terminal listens for the user's click to send the request, reads the string from the text input area, encodes and performs simple cleaning (such as removing leading and trailing whitespace), merges the string with the most recently updated competition status information and the virtual display object attribute configuration into a structured object, and attaches a session token to form a request packet.
[0120] Step 9: After receiving a question-and-answer request, the server constructs prompts for the generative AI model. The input includes the question information, the competition status information, and the attribute information of the virtual display object; the output is a prompt text for model inference. Internally, the server uses template concatenation to insert the competition background, user question, and role settings into a pre-designed text template. For example, the server generates the following prompt: "The user is watching a football match. Current match information: Home team vs. Away team, score 1-0, match time 32 minutes. Key player identifier: player_9."
[0121] The question asked by the user: How many assists has this player had this season? You are a football data analysis expert and a virtual coach selected by the user. You need to answer questions using professional but easy-to-understand language based on the statistical data (number of shots, number of goals, number of assists, etc.) provided in the database.
[0122] Please answer from the perspective of a virtual coach, first providing specific data, then briefly analyzing the level of that data in the league. When generating the prompt, the server combines the fields into continuous text by string concatenation and placeholder replacement.
[0123] Step 10: The server feeds the prompts and statistical information as input to the generative AI model and generates response text. The input consists of the prompt text and statistical data related to the key object; the output is a response text string in natural language. First, the server tokenizes the prompts, mapping each token to an integer index using a vocabulary, and then converts it into a vector sequence via an embedding layer. In the Transformer encoder, the server calculates self-attention weights, synthesizing the semantic relationships between different parts of the prompts, and encodes the statistical data as additional features, concatenating them into the vector sequence. In the decoder, the server generates output tokens step-by-step, calculating a conditional probability distribution based on the previous generation and encoding results at each step, selecting the highest-scoring token, until a final token is generated. Finally, the server reconstructs the output token sequence into natural language text as the response text.
[0124] Step 11: After generating the response text, the server performs consistency checks and style adjustments on the text content. The input consists of the original response text, actual statistical information from the database, and user preference information. The output is the final response text after correction and stylization. The server parses the numerical segments in the response text and compares them with the statistical fields in the database. If discrepancies are found, they are replaced with accurate values. The server then determines whether tactical analysis or simplified explanations are needed based on user preference information, adjusting the text style by inserting or replacing specific sentence templates to create output that matches the set character's tone.
[0125] Step 12: The server sends the final response text to the terminal. The inputs are the final response text and the session token, and the output is the text data cached locally by the terminal. The server encapsulates the text into a response message and transmits it to the terminal over the network; the terminal parses the message body in its receiving module, stores the response text in the chat history data structure, and triggers the display and speech synthesis process.
[0126] Step 13: The terminal uses response text information to drive virtual display objects to perform language output and motion coordination. The input consists of response text information and virtual display object attribute information; the output is a combined multimedia output presented on the display and audio devices. The terminal displays text in the form of speech bubbles on the interface, while simultaneously calling a text-to-speech engine to convert the text into speech signals. Speech rate and intonation parameters are set according to the virtual display object attributes. The terminal divides the speech duration into several time slices, synchronizes them with the skeletal animation system, generates corresponding lip-sync animation parameters, and overlays them with existing motion trajectories to achieve synchronized speech and lip-sync animation.
[0127] Step 14: During interaction, the terminal continuously collects the user's facial expressions, voice information, and operation information, and uploads them to the server for emotion analysis. The inputs are facial images captured by the camera, voice segments collected by the microphone, and user interaction event logs; the output is a feature vector for emotion estimation. The terminal performs facial detection and keypoint localization on the images, normalizing the keypoint coordinates as facial expression features; it extracts features such as pitch, speech rate, and energy from the voice signal; and it statistically analyzes operations such as click frequency and swipe speed as behavioral features. These features are combined into a vector and sent to the server.
[0128] Step 15: The server estimates the user's emotional state and updates preference information based on the uploaded feature vectors. The inputs are the user's emotional feature vectors and historical behavioral data; the outputs are emotional state labels and updated preference vectors. In the emotion classifier, the server inputs the feature vectors into a multilayer perceptron network, calculates the probability distribution of each emotion category, and selects the label corresponding to the maximum value. Simultaneously, the server updates the preference vectors associated with the user, employing a weighted average or incremental update strategy to incorporate new behavioral patterns into historical preferences. When constructing subsequent prompts, the server maps the emotional state and preference vectors into text descriptions and embeds them into the prompts, guiding the generative AI model to generate responses with a style more closely aligned with the user's state.
[0129] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0130] In existing sports event viewing systems, computer devices are mostly used only as tools for transmitting and playing video content. They lack fine-grained structured processing of live event data and efficient collaboration with generative artificial intelligence models, resulting in the following shortcomings in computer technology.
[0131] First, the event video data collected at the terminal side is usually in the form of continuous video streams. Servers often only perform simple forwarding or coarse-grained labeling, without fine-grained association of video data with time information and event type, nor generating "event extraction information" that can be efficiently utilized by downstream algorithms through a unified data structure. Therefore, servers cannot perform indexing optimization and fast retrieval of sports action databases at a systematic event data level, resulting in inefficient use of storage and computing resources.
[0132] Secondly, existing systems, when utilizing generative AI models, often directly concatenate user input with a small amount of context before sending it to the model. They lack a mechanism for generating prompts that is integrated with the domain data structure, resulting in a lack of rigorous mapping between prompts and event information, rules, and user preferences. This not only increases the burden of processing irrelevant information during model inference but also makes it difficult for the generated results to meet the requirements of real-time interactive scenarios in terms of stability, consistency, and controllability regarding response latency and quality.
[0133] Furthermore, most commentary systems remain at the level of simple parameter configuration or template switching in terms of user interaction and personalization. They lack mechanisms for machine-processable modeling of multimodal data such as user behavior history and emotional state, and cannot use the computer's internal data structures and learning algorithms to persist and abstract user preferences and feed them back into subsequent prompt generation and model invocation processes. This "stateless" or "weakly stateful" design makes it difficult for the server to make dynamic and nuanced adjustments to the commentary style, content depth, and question-and-answer granularity at the system level.
[0134] Furthermore, in existing technologies, image overlay processing and generative AI processing are mostly completed by relatively independent modules with loose coupling, lacking an integrated pipeline centered on "narrative text" that connects event extraction, model inference, and video frame overlay rendering. This not only causes data to be copied multiple times between different subsystems, undergo format conversions, and involve unnecessary network transmissions, but also increases processing latency, affecting the overall performance and resource utilization of the system in mobile terminal and network bandwidth-constrained environments.
[0135] Therefore, how to design a data processing flow centered on event extraction information, action information storage structure, prompt statement generation logic, and user preference model under the overall architecture of server and terminal collaboration, so that generative artificial intelligence models can run efficiently under structured context constraints and be deeply integrated with video frame overlay rendering, thereby reducing computational and communication load while improving response speed, narration quality, and personalization, has become a technical issue that urgently needs to be solved in this field.
[0136] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0137] In this invention, the server includes a module for receiving sports competition video data with time information uploaded by a portable terminal device and generating event extraction information representing competition types and events; a module for performing an index retrieval on an action information storage unit storing multiple competition types and multiple competition events based on the event extraction information and user-associated setting information to obtain corresponding action information and rule information; a module for combining virtual character attributes, event extraction information, action information, and rule information to generate prompt statements input to a generative artificial intelligence model and instructing the generative artificial intelligence model to generate explanatory responses or question-and-answer responses based on the prompt statements; a module for associating text responses from the generative artificial intelligence model with the original video data and outputting structured explanatory data for overlay rendering on the terminal side; and a module for storing user setting information, question text, response results, and sentiment inference results as historical information and updating the user preference model based on statistical processing or machine learning processing to provide feedback and adjust subsequent prompt statement generation strategies. This allows for the formation of an integrated processing chain within the server, encompassing everything from extracting event footage, indexing domain data, automatically constructing prompts, controlling the invocation of generative AI models, to outputting directly overlaid and rendered data to the terminal. This enables collaborative optimization of computing and storage resources, reduces redundant information processing in prompt construction and model inference, improves the real-time performance and stability of response generation, and allows for refined and personalized adjustments to the commentary style and Q&A content of virtual characters at the computer system level through continuous updates to the user preference model. Ultimately, this improves the overall human-computer interaction experience and the technical performance of the sports commentary system.
[0138] "Information processing device" refers to an electronic device, including a processor, memory, and communication interface, used to execute programs and comprehensively process input information from terminals, data from the server side, and the output of generative artificial intelligence models.
[0139] "Portable terminal device" refers to a portable electronic device with a display unit, camera unit, microphone and communication functions, used to capture sports competition images, obtain user input and send data to the server.
[0140] "Display device" refers to a display component used to present a graphical interface, text information, and video images with overlaid explanatory text to a user, including but not limited to liquid crystal display panels and organic light-emitting display panels.
[0141] "User interface" refers to a graphical user interface displayed on a display device that can receive user touch, click, swipe and other operations, and is used by users to select virtual character attributes, explain policies and input questions, etc.
[0142] "Virtual characters" refer to computer-generated digital characters with personalized settings. Their commentary style, language characteristics, and response methods can be adjusted according to settings and user preferences to perform sports commentary and Q&A.
[0143] "Virtual character attributes" refer to the parameterized settings related to the virtual character, such as personality traits, tone style, language type, and level of detail in the explanation, which are used to control the presentation of the explanation text and the question and answer text.
[0144] "Commentary guidelines" refer to strategic information used to guide virtual characters in sports commentary, including setting commentary directions such as emphasizing emotional expression or tactical analysis, highlighting basic rule explanations or advanced statistical interpretations.
[0145] "Setup information" refers to the configuration information generated and stored by associating the user's selected virtual character attributes, explanation guidelines, and user identification information with the user, which is used in subsequent prompt statement generation and response control.
[0146] "Image data" refers to digital signal data of sports competition footage captured by camera or video acquisition units, existing in the form of video streams or image frame sequences.
[0147] "Time information" refers to the time stamps corresponding to each image frame or event in the video data, including timestamps, relative playback time, or game minutes, which are used to establish time associations in event extraction and data retrieval.
[0148] "Event extraction information" refers to structured information generated based on the analysis of image data and time information, used to represent the type of sports competition, the type of competition event that occurred, and its time location.
[0149] "Competition type" refers to the category information of sports, used to distinguish different sports such as football and basketball, so that the corresponding data set can be selected in the motion information storage department.
[0150] "Competition events" refer to specific actions or situations in a sports competition, such as goals, fouls, shots, and penalty kicks, which are used to drive commentary and rule explanations.
[0151] "Motion Information Storage Unit" refers to a storage resource unit used to store motion information and related rule information corresponding to multiple competition types and multiple competition events in a data structure manner.
[0152] "Action information" refers to action description data related to a specific competition event, including participating roles, action process, location area, etc., which is used to provide a structured semantic basis for commentary and Q&A.
[0153] "Rule information" refers to the competition rules, judging criteria, and their interpretations related to a sport, stored in a form accessible to the program for use in generating commentary and Q&A.
[0154] "Index retrieval" refers to the process of using a pre-defined index data structure (such as key value, time index, event type index) to query the action information storage unit in order to efficiently obtain target action information or rule information.
[0155] "Generative artificial intelligence models" refer to artificial intelligence models built on deep learning technology that can automatically generate natural language output based on input text or structured information, including but not limited to language models based on transformer structures.
[0156] "Prompt statements" refer to text content constructed for input into generative artificial intelligence models. This text explicitly describes the current context, role settings, event information, and output requirements, and is used to constrain and guide the model to generate the expected explanations or responses.
[0157] "Commentary responses" refers to natural language text generated by generative artificial intelligence models based on commentary prompts, used to describe and analyze current sports competition events.
[0158] "Question-answering responses" refers to natural language text generated by generative artificial intelligence models based on user questions and related contextual prompts, used to answer user questions.
[0159] "Structured explanation data" refers to a data set that organizes explanation responses or question-and-answer responses together with time information, event markers, etc., and is suitable for overlay rendering and synchronized playback on the terminal.
[0160] "Image processing software" refers to software components or library functions that run on information processing devices or terminals and are used to perform operations such as drawing text, overlaying layers, and calculating coordinates on images or video frames.
[0161] "Display image data" refers to video frame data that is output to users by the terminal display device after overlaying and drawing explanatory text, marking information, etc.
[0162] "Voice input" refers to the audio signals provided by the user to the system through a microphone, which are acoustic data used to represent the content of questions or instructions.
[0163] "Speech recognition processing" refers to the computational process of extracting features, acoustic modeling, and decoding language from speech input to generate a corresponding text representation.
[0164] "Question text" refers to data obtained through speech recognition processing or user input, which describes the content of the user's question in text form.
[0165] "Question content" refers to a structured set of information, which is a combination of question text, event-extracted information, and action information, used to explain the context and background of a user's question to a generative artificial intelligence model.
[0166] "Face information" refers to the feature data acquired through cameras or other sensors that represents the state of a user's facial expressions and can be used to infer the user's emotions.
[0167] "Voice feature information" refers to acoustic parameters such as pitch, speech rate, and volume extracted from user voice input, which are used to help infer the user's emotions or state.
[0168] "Operation history information" refers to the log of user interaction behavior such as selection records, click behavior, and question frequency on the operation interface, which is used for user preference analysis.
[0169] "Emotion inference processing" refers to the computational processing that infers a user's current emotional state based on facial expression information, voice feature information, or operation history information, using rules or learning models.
[0170] "Emotional presumption results" refer to the representation of a user's emotional category or intensity obtained from emotional presumption processing, which is used to adjust the explanation style or response method.
[0171] "Historical information" refers to data such as settings, question texts, response results, and sentiment inference results that the system saves for the same user over a period of time, which are used to build and update user preference models.
[0172] A "user preference model" is a data model built based on historical information through statistical analysis or machine learning methods to represent users' preferences in terms of narration style, content depth, and topic interest.
[0173] "Personalized live commentary" refers to adjusting the content, style, and level of detail of commentary responses based on user preference models and the current competition context, thereby providing differentiated commentary for different users.
[0174] Personalized Q&A refers to the process of generating Q&A responses by incorporating user preference models and historical behavior information to adjust the professionalism, explanation methods, and language style of the answers, thereby tailoring the Q&A process to the individual user's needs.
[0175] In one embodiment of the invention, the system comprises a collaborative working structure consisting of a server, a terminal, and a user. The server operates in a data center or cloud computing environment, and the terminal is a portable electronic device with a camera and a display. The user interacts with the server through the terminal. The server can be deployed on a computer system equipped with a general-purpose processor (e.g., a multi-core CPU), a graphics processor (e.g., a general-purpose graphics accelerator card), main memory, and large-capacity storage devices, running an operating system (e.g., a general-purpose server operating system), application server software, and a generative artificial intelligence model framework for inference (e.g., a deep learning framework based on a transformer architecture). The terminal can run a mobile operating system and use a graphical user interface framework, a camera access interface, an audio acquisition interface, and a network communication module. The user performs operations such as selecting a virtual character, watching overlaid narration videos, and asking questions through the terminal.
[0176] During system initialization, the server prepares multiple functional modules in the storage device, including: an event extraction module, an action information storage module access interface, a prompt statement generation module, a generative artificial intelligence model inference module, a commentary data packaging module, a user preference modeling module, and a log recording module. The server loads sports action information and rule information into the main memory. This information is stored in a hierarchical data structure: the upper layer is the competition category identifier, the middle layer is the competition event type, and the lower layer contains fields such as action description, participating entities, location area, and rule description text. The server maintains an index structure for this data, such as an inverted index with competition category and event type as keys, and a time index with time range as keys, enabling the server to quickly query relevant action and rule information upon receiving event extraction information.
[0177] The terminal uses its built-in camera to capture image data while the user watches sports events. The terminal converts the raw image signal into digital image frames via the system's multimedia interface and compresses the frames into a video stream format or samples keyframe images at preset intervals using a local hardware encoding module. The terminal appends timestamp information to each frame or keyframe, organizing "image frame data + timestamp + competition identifier (e.g., selected by the user or pre-configured)" into a structured data packet, which is then sent to the server via a network communication module (e.g., based on TCP / IP and secure transmission protocols). Users can select the virtual character's personality tags, tone style, language type, and level of commentary detail in the terminal's interface. The terminal collects these options as configuration data and sends it to the server along with the user's identification information.
[0178] After receiving image data uploaded by the terminal, the server analyzes the image data using an event extraction module. The server can preprocess the image frames using image processing libraries (such as general-purpose open-source image processing libraries), performing tasks such as resolution scaling, color space conversion, and denoising. Then, it uses pre-trained visual classification or detection models (such as recognition models based on convolutional neural networks or video converter structures) to extract features. From these features, the server identifies the competition type, possible competition event types (such as goals, fouls, shots, and penalty kicks), and relevant time intervals, thereby generating event extraction information. This event extraction information is represented in a structured form, including fields such as "Competition Type Identifier," "Event Type," "Start Time," "End Time," and "Confidence Level." Because the server represents events from the video using a unified data structure, this abstraction layer can be reused in subsequent index retrieval and prompt generation, thereby reducing redundant analysis and improving computational efficiency.
[0179] After generating event extraction information, the server accesses the action information storage module. Based on the competition type and event type in the event extraction information, the server uses an index structure to retrieve the action information table in the storage device. For example, when the event extraction information indicates "the competition type is football, and the event type is a goal," the server finds the action templates, common tactical descriptions, and rule entries related to penalties corresponding to this type of event in the action information storage. The server further filters for more common tactical patterns in the current stage of the match based on time information to compress the candidate set. This retrieval method based on event extraction information and index structure allows the server to quickly obtain highly relevant action and rule information from a large amount of sports data, thereby reducing the dependence of generative artificial intelligence models on long background information, reducing the model input length, and improving inference speed.
[0180] After acquiring action and rule information, the server uses a prompt generation module to construct prompts for input into the generative AI model. Based on the virtual character attributes (e.g., "personality: excited, humorous," "style: professional and analytical") and event descriptions extracted from the event data, the server organizes these elements into natural language text according to a predefined template. The prompts not only contain an objective description of the current event but also control constraints on the explanatory output, such as word count, tone, and whether rule explanations are required. This templated and parameterized approach to prompt generation ensures a stable and predictable structure, facilitating subsequent rule checks and output filtering outside the generative AI model, thereby improving the overall system's controllability and stability.
[0181] For example, when the server detects a goal, and the user's virtual persona is set to "excited commentator," it can generate the following message: "You are a virtual football commentator character. Personality: very excited and humorous. Current scene: The home team scores in the 67th minute through a shot from striker number 10 inside the penalty area. Please describe the goal in detail using a live commentary style, between 80 and 150 words, with rhythm and engaging language." For example, when the server receives a user's question about the foul rules, it can generate the following prompt: User question: 'Why was that goal ruled offside?' Current match: Football. Recent event: In the 54th minute, the away team's striker received a through ball from his teammate and scored a one-on-one goal, but was ruled offside. According to the action database, at the moment of the pass, the striker was closer to the goal than the second-to-last defender. Please explain the reason for this offside decision to a viewer with limited football knowledge in simple terms, and also briefly introduce what the offside rule is. Your answer should be limited to 120 words. The server inputs the aforementioned prompts into the generative artificial intelligence model's inference module. In this module, the server uses a transformer-based neural network model, consisting of multiple encoders and decoders, employing a self-attention mechanism to model the input prompts. The server first calls the word segmentation and encoding components to convert the Chinese prompts into a sequence of symbols at the sub-word level, then maps them to a high-dimensional vector representation. The server performs matrix multiplication, attention weight calculation, and feedforward network operations on the graphics processor, progressively generating the output token sequence through multi-layer stacking. During inference, the server can employ beam search or temperature-controlled sampling strategies to balance the diversity and stability of the generated data.
[0182] To improve the model's inference performance, the server uses a large-scale corpus of sports commentary and rule explanations during the training phase. During learning, the server employs cross-entropy loss as the basic error metric, calculates gradients using backpropagation, and updates model weights based on an adaptive learning rate optimization algorithm (e.g., adaptive moment estimation). The server can also incorporate data augmentation into the training data, such as synonym replacement, random masking, and text recombination for different match scenarios, to enhance the model's generalization ability across diverse contexts. After training, the server distills or quantizes the model to reduce the number of parameters and improve inference speed, making it more suitable for real-time commentary scenarios.
[0183] After receiving the labeled sequence output by the generative AI model, the server decodes it into Chinese text, i.e., explanatory responses or question-and-answer responses. The server then performs post-processing on this text, including removing redundant and repetitive segments, checking word count ranges, and filtering inappropriate language. The server can also compare event-extracted information with action information to trim or correct content in the model output that is clearly inconsistent with reality, such as detecting incorrect mentions of non-existent team names or incorrect score descriptions, thereby improving the overall output accuracy.
[0184] After receiving the commentary data from the server, the terminal performs image overlay rendering locally. First, the terminal acquires the image frame to be displayed via a video player or a custom drawing pipeline. Then, it uses image processing software to draw text in specific areas of the image frame. The terminal can calculate the text's drawing position and line width based on screen resolution and font size to avoid obscuring key areas. When drawing text, the terminal can first draw a semi-transparent rectangle in the text background area, then draw white text on the rectangle to improve readability against complex backgrounds. The terminal sends the processed image frame to the graphics rendering pipeline, and the final output is displayed to the user. Because the server pre-outputs the structured interpretation data, the terminal can perform only lightweight image overlay operations with limited computing resources, thereby reducing the CPU and battery load on the terminal side.
[0185] During viewing, users can ask questions to virtual characters via the terminal. When users ask questions using voice, the terminal calls the audio acquisition interface to obtain audio signals from the microphone in real time, and converts the audio into text locally or through a connected speech recognition service. The terminal sends the obtained question text along with extracted event information to the server. After receiving the question text, the server uses relevant records and rule information from the action information storage to supplement the context of the question, such as the most recent ruling, the current score, and the roles involved in the event. The server combines this background information with the question text in the prompt generation module to generate more targeted and informative prompts for question-and-answer interaction. This structured prompt design allows generative AI models to obtain the key context needed to generate questions and answers without having to backtrack through all information in a long dialogue history, thus reducing the burden of input length and attention computation, thereby improving inference speed and reducing memory usage.
[0186] Throughout the session, the server records historical information such as user settings, question text, generated responses, and sentiment inferences. During non-real-time phases or periods of low load, the server can perform statistical or machine learning processing on this historical information. For example, the server can analyze the distribution of user questions across different competitions, the depth of preferred explanations, the level of attention paid to rule descriptions, and the dwell time under different commentary styles. Based on these statistical features, the server trains a user preference model. This model can employ structures such as shallow neural networks or tree models, taking a vector representation of the user's historical behavior as input and outputting multiple preference parameters (e.g., "preference for tactical analysis weight," "preference for sentimental explanation weight," "preference for rule explanation weight," etc.). When generating subsequent prompts, the server references these preference parameters to automatically adjust the output requirements of the prompts, such as increasing the proportion of tactical analysis, reducing lengthy background information, and appropriately adding or removing emotional vocabulary. This feedback mechanism allows the system to dynamically optimize the prompt construction process at a technical level, achieving personalized control over the output style of the generative artificial intelligence model.
[0187] Through modular collaboration, the server and terminal achieve technological improvements in several aspects. First, the server structures the raw video using event extraction and index retrieval, avoiding repeated video content parsing each time the generative AI model is invoked, thus significantly reducing redundant computation and communication overhead. Second, the server explicitly injects event data, rule information, and user preferences into the model input through structured prompts, enabling the model to generate scene-appropriate results with shorter prompts, thereby shortening the model input length, reducing attention computation, and improving inference latency. Third, the server automatically adjusts the prompts using a user preference model. This computer-based approach controls model behavior at the prompt level, allowing for multi-dimensional parameter adjustments automatically within the system compared to traditional methods where human operators manually modify prompts. This represents a technical optimization of the model invocation strategy. Furthermore, the terminal performs lightweight image overlay and audio playback processing locally, enabling real-time fusion of server-generated text with video frames, providing a smooth real-time overlay narration experience with limited hardware resources.
[0188] In another implementation, the server can employ different generative AI model structures, such as a decoder-only transformer model, or adding a dedicated structured information encoding layer before the text model to directly receive vector encodings of event extraction and action information, rather than relying solely on natural language prompts. The server can also train the model using multi-task learning, enabling it to simultaneously learn commentary generation and rule-based question answering tasks. While sharing parameters, it can generate different types of text using different output heads, further enhancing the model's adaptability to the sports domain. The server can introduce various error function combinations during training, such as combining the fluency loss of the commentary text with the factual consistency loss, to reduce factual errors in the generated text. The server can also explicitly label event types in the training data, allowing the model to employ different generation strategies for different events during decoding; these are all variations of this invention.
[0189] In another implementation, the terminal can partially handle event extraction. For example, the terminal can run a simplified motion detection model locally to perform coarse-grained event recognition on video frames, sending only event tags and thumbnail frames to the server. The server then performs more refined inference and text generation based on these tags. This distributed processing approach can further reduce the server's video analysis burden, making it particularly suitable for scenarios with a large number of users. The terminal can also offload some user preference calculations locally, such as using local storage to record simple preference parameters and adjusting text display or voice playback methods offline, thereby enhancing the system's robustness.
[0190] In summary, in the embodiments of this invention, the server, terminal, and user form a tightly coupled data processing link around the event extraction information, action information storage structure, and prompt statement generation mechanism. This enables the generative artificial intelligence model to be deeply integrated with video rendering and user interaction in specific hardware and software environments, thereby achieving comprehensive technical improvements in terms of narration generation speed, text quality, personalization, and system resource utilization.
[0191] use Figure 12 The processing flow is explained.
[0192] Step 1: Users select a virtual character and set their narration preferences on the terminal.
[0193] Users can click on virtual character cards in the character list interface via the terminal's touchscreen, select personality (such as excited, calm), style (such as humorous, professional) and language type from the drop-down menu, and set the level of detail in the explanation using the slider.
[0194] Input: Touch coordinates, click events, drop-down selection results, and slider values received by the terminal interface components.
[0195] Based on these input events, the terminal encapsulates the character identifier, personality tag, style tag, language type, and level of detail into structured setting data, and adds user identification information.
[0196] Output: The terminal generates a set of configuration data including the user ID and virtual character setting parameters, and sends it to the server via the network.
[0197] Step 2: The server receives and stores user settings.
[0198] The server receives configuration data from the terminal through the network interface, and the application parses the message body to extract fields such as user ID, role identifier, style, and level of detail.
[0199] Input: A JSON data packet containing user-defined fields.
[0200] The server uses a data parser to map JSON to an internal data structure and calls the database access module to write these fields into a session table or user preference table, establishing a link between user IDs and virtual role settings.
[0201] Output: The server generates or updates a session record in the database and caches the corresponding settings information in memory for use in generating subsequent prompt statements.
[0202] Step 3: The terminal collects video data of sports events and adds time information.
[0203] The terminal calls the camera interface to periodically read raw frame data from the image sensor, uses a hardware encoding module to encode the raw data into a video stream or sample it into keyframe images at intervals, and reads the current timestamp from the system clock.
[0204] Input: Raw image signal captured by the camera and system clock time.
[0205] The terminal performs format conversion (such as YUV to compressed format) and resolution scaling on the raw data, and adds a timestamp and a preset competition category identifier to each frame or keyframe, organizing it into a data structure of "image / video frame + timestamp + competition identifier".
[0206] Output: The terminal generates an image data packet with time information and sends it to the server over the network.
[0207] Step 4: The server extracts events from the image data and generates event extraction information.
[0208] The server receives timestamped image data from the network interface, performs preprocessing on the image frames using an image processing library, such as scaling, denoising, and color space conversion, and then inputs the preprocessing results into a pre-trained visual recognition model.
[0209] Input: An image data stream containing image frames and timestamps.
[0210] The server extracts feature vectors from the visual recognition model and matches these features with learned patterns to determine whether there are competitive events such as goals, fouls, or shots in the current frame, recording the corresponding time intervals and confidence levels. The server combines the competition type, event type, start and end times, and confidence levels into structured event extraction information.
[0211] Output: The server generates one or more event extraction information records for use in subsequent action information retrieval and prompt statement construction.
[0212] Step 5: The server retrieves action and rule information based on event extraction information.
[0213] The server reads the competition type and event type from the event extraction information, calls the retrieval interface of the action information storage department, and performs a query on the database using the index with the competition type and event type as the key.
[0214] Input: Event extraction information and hierarchical motion data and rule data stored in the motion information storage unit.
[0215] The server filters the corresponding action template (such as "shooting action in the penalty area"), the description of the participating roles, and the relevant rule entries based on the event type, and transforms the original table scan into a small number of index lookups and primary key accesses through indexes.
[0216] Output: The server receives a set of action information and related rule information that match the current event, which serves as the context input for generating the prompt statement.
[0217] Step 6: The server constructs prompts for explanation.
[0218] The server reads user settings information from the session cache, reads event description fields from event extraction information and action information, and populates these text fields into a predefined prompt template.
[0219] Input: User settings (character personality, narration style, language, etc.), event extraction information, action information, and rule information.
[0220] The server concatenates text fragments in a fixed order, such as first describing the character's identity and personality, then describing the time, place, and actions of the current event, and finally adding output constraints (such as word count and tone) to complete the overall construction of the prompt statement.
[0221] Output: The server generates a prompt in natural language to be used as input to the generative artificial intelligence model, for example: "You are a virtual football commentator character. Personality: very excited and humorous. Current scene: The home team scores in the 67th minute through a shot from striker number 10 inside the penalty area. Please describe the goal in detail using a live commentary style, between 80 and 150 words, with rhythm and engaging language." Step 7: The server calls a generative artificial intelligence model to generate explanatory responses.
[0222] The server inputs the prompt statement into the word segmentation and encoding module, which converts the text into sub-word units and maps them into vector sequences. Then, forward inference of the transformer network is performed on the graphics processor.
[0223] Input: A structured prompt text.
[0224] The server computes self-attention weights and feedforward network outputs within the model and iteratively generates output tokens, using beam search or temperature sampling to control the generation process until a termination token is generated or the character limit is reached. Then, the server decodes the output token sequence into Chinese sentences and performs length pruning and inappropriate word filtering on the generated results.
[0225] Output: The server receives a commentary response text corresponding to the current event.
[0226] Step 8: The server packages and interprets the data and sends it to the terminal.
[0227] The server encapsulates the narration response text along with the corresponding time information and event type into a narration data structure, so that the terminal can display it synchronously according to the timeline.
[0228] Input: Explanation of the time field in the response text and event extraction information.
[0229] The server adds metadata such as "start timestamp", "end timestamp" and "display location parameters" to the commentary data, and sends this commentary data back to the terminal via network response.
[0230] Output: The server outputs a structured parsed data packet containing text and control information for aligning the video timeline.
[0231] Step 9: The terminal overlays the narration text onto the video screen.
[0232] The terminal parses the narration data from the server response, determines the narration text items to be displayed based on the current playback time, and retrieves the current image frame through the video playback module.
[0233] Input: The narration data packet returned by the server and the image frame being played locally.
[0234] The terminal uses an image processing library to draw text at specified positions within an image frame. It calculates coordinates based on font size and screen resolution, first drawing a semi-transparent background bar, then drawing the explanatory text character by character. The terminal then outputs the image frame with the overlaid text to the screen via the graphics rendering pipeline.
[0235] Output: The terminal generates and displays a game screen overlaid with commentary text, allowing users to see a screen overlay commentary synchronized with the event.
[0236] Step 10: Users can submit questions related to the competition through the terminal.
[0237] If users have questions about the rules, tactics, or a particular event while watching, they can click the "Ask a Question" button on the terminal and choose between voice or text input.
[0238] Input: User actions (button clicks) and text input via voice or keyboard.
[0239] Based on the user's selection, the terminal turns on the microphone to record or activates the text input box, collects complete audio segments or text questions, and prepares to send them to the server.
[0240] Output: The terminal generates a question request containing the source data of the question (voice or text) and proceeds to subsequent processing.
[0241] Step 11: The terminal converts the voice problem into a text file and uploads it.
[0242] When a user asks a question in voice form, the terminal sends the audio signal collected by the microphone to the speech recognition module, performs feature extraction and acoustic decoding on the audio, and generates the corresponding text result.
[0243] Input: User's voice audio stream or directly entered text questions.
[0244] The terminal performs framing and feature extraction (such as Mel-frequency cepstral coefficient calculation) on the audio, and maps the acoustic features into a text sequence through built-in or connected speech recognition services to obtain the question text. The terminal packages the question text together with the current time and current event extraction information identifier, and sends it to the server via the network.
[0245] Output: The terminal outputs structured problem request data, which includes the problem text and contextual identification information.
[0246] Step 12: The server generates prompts for answering questions based on the question and context.
[0247] After receiving a problem request, the server reads action and rule information related to recent events from the action information storage department and reconstructs the problem background by combining it with the current game context.
[0248] Input: Question text, current or recent event extraction information, and action and rule records from the database.
[0249] The server concatenates "User Questions," "Recent Events Overview," and "Rule Highlights" according to a predefined pattern, adds output requirements (such as word limits and explanation levels), and generates question-and-answer prompts. For example: User question: 'Why was that goal ruled offside?' Current match: Football. Recent event: In the 54th minute, the away team's striker received a through ball and scored a one-on-one goal, but the referee ruled him offside, and the goal was disallowed. According to the motion database, at the moment of the pass, the striker was closer to the goal than the second-to-last defender. Please, as a virtual commentator, explain the reason for this offside decision to novice viewers in simple and easy-to-understand language, and briefly explain what the offside rule is, within 150 words. Output: The server receives a structured question-and-answer prompt, preparing to invoke the generative artificial intelligence model.
[0250] Step 13: The server calls a generative artificial intelligence model to generate responses for the questions.
[0251] The server inputs the question and answer prompts into the generative artificial intelligence model, which then generates the response text through the same word segmentation, encoding, and attention calculation process.
[0252] Input: Question-and-answer prompts constructed in response to the user's question.
[0253] The server generates text tags progressively in the model output until the stopping condition is met. Then, it performs factual consistency checks and length control on the output text, deletes irrelevant content, and retains explanations and rule descriptions that are directly helpful to the problem.
[0254] Output: The server receives a question-and-answer text that clearly answers the user's question about the rules or the cause of the event.
[0255] Step 14: The server returns the Q&A results, which are displayed on the terminal and can optionally be played as audio.
[0256] The server packages the question and answer text into JSON or a similar structure and sends it to the terminal over the network.
[0257] Input: The generated question-and-answer text.
[0258] After receiving the message, the terminal inserts the question and answer text into the message list of the question and answer interface and displays it. When the user enables voice broadcasting, the local text-to-speech module is called to convert the text into audio and play it through the speaker so that the user can see and hear the virtual character's answer.
[0259] Output: The terminal generates an updated dialogue interface display and optional voice output, providing users with complete question-and-answer feedback.
[0260] Step 15: The server updates the user preference model based on this interaction.
[0261] After completing a presentation or Q&A session, the server writes the settings, question text, generated response, and emotion-related parameters (such as user interaction frequency) used in this session into the database as a new historical record.
[0262] Input: Configuration information, event extraction information, prompt statements, response text, and related metadata for this session.
[0263] The server performs batch statistical or incremental learning processing on the accumulated historical records, and adjusts the parameter weights in the user preference model through feature extraction and model update algorithms, so that the model gradually reflects the user's preferences for narration style, explanation depth and topic type.
[0264] Output: The server obtains the updated user preference model and saves it in memory or storage, which is used to influence the parameter settings and generation strategy of subsequent prompt statements.
[0265] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0266] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0267] In computer-based sports viewing technology, with the development of virtual avatar presentation technology, data-driven motion generation technology, and generative artificial intelligence models, users' demands for personalized, interactive, and deeply understandable viewing experiences are constantly increasing. However, existing technologies typically suffer from the following problems: First, on the terminal side, virtual characters are simply overlaid onto the video screen, lacking fine-grained structured management and intelligent retrieval of sports motion data. This makes it difficult to dynamically select appropriate motion data for display based on the user's fine-grained needs during the viewing process. Second, when processing prompts input by users in natural language, existing systems mostly use rule matching or simple keyword retrieval methods. They cannot combine multi-source contextual information (such as sports type, scene type, level of detail requirements, etc.) to perform high-precision semantic parsing of prompts, making it difficult to accurately map the user's subjective intent to the corresponding behavioral information set in the motion database. Third, generative AI models are often used solely to generate text responses, lacking a collaborative processing mechanism with structured action data. This prevents them from performing a "selection-optimization-ranking" computational process based on candidate behavior information, hindering intelligent filtering and priority control of candidate actions. Consequently, dynamic, fine-grained control of virtual character behavior sequences cannot be achieved on the terminal side. Fourth, existing systems generally lack a systematic learning and storage mechanism for the relationship between users' historical prompts and actual selected behavior information. The system struggles to automatically adjust subsequent candidate behavior information extraction conditions and prompt generation strategies during multiple rounds of interaction, causing the system to remain at a static response level for extended periods, failing to demonstrate the computational performance improvement of "adaptive optimization with use." Fifth, the control of action reproduction and explanation presentation on the terminal side is mostly based on fixed logic, lacking a dynamic scheduling mechanism for the reproduction speed, reproduction order, and explanation granularity based on the evaluation information output by generative AI models. This prevents adaptive optimization of the virtual character's presentation for actions of varying complexity or different user preferences, thus limiting the improvement of human-computer interaction capabilities and computational resource utilization efficiency on the terminal side.
[0268] Therefore, it is necessary to provide a new system based on the collaborative work of information processing devices and generative artificial intelligence models. This system achieves the following by dividing the work between the server and the portable information processing device: structured indexing and retrieval of sports behavior information, semantic parsing and generation of natural language prompts, intelligent selection and sorting of candidate behavior information, and dynamic presentation control of the behavior sequences and commentary information of virtual display objects. This will enable overall improvements at the computer technology level, including "computational structure," "data organization method," "model calling process," and "interactive control mechanism," thereby enhancing the ability of virtual character-driven action demonstrations and commentary in spectator applications.
[0269] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0270] In this invention, the server includes a behavior information storage area for managing behavior information corresponding to various competitive activities, and attaching index information, including competition category information, scene category information, and detail information, to each behavior information. This allows the server to extract candidate behavior information related to the current context from the behavior information storage area using the index information, based on the parsing results of viewing video information from the terminal and the user's natural language prompts. The server also includes a unit for generating prompts for input into a generative artificial intelligence model based on the user's input natural language prompts and the candidate behavior information, and inputting the generated prompts along with the candidate behavior information into the generative artificial intelligence model. The generative artificial intelligence model then performs relevance evaluation and selection processing on the candidate behavior information to obtain target behavior information optimized for user needs. Furthermore, the server includes a unit for accumulating the correspondence between past user prompts and the target behavior information in a learning storage area, and dynamically adjusting the extraction conditions for subsequent candidate behavior information and the prompt generation conditions based on this accumulation result to perform adaptive learning processing. This allows for a combined computational process on the server side, encompassing a data retrieval mechanism based on structured indexes, a behavioral information selection mechanism based on generative artificial intelligence models, and an adaptive learning mechanism based on interaction history. This systematically optimizes the organization and retrieval of behavioral data, the generation of prompt statements, and the model invocation path, thereby improving the parsing accuracy and behavioral information matching accuracy for complex natural language requirements.
[0271] In this invention, the server includes a unit for outputting target behavior information selected by a generative artificial intelligence model and associated response information to a portable information processing device via a communication interface. This enables the terminal to obtain action data and semantic information that are suitable for the current user's needs and have been preliminarily optimized and filtered on the server side, providing high-quality input for behavior reproduction and interpretation control on the terminal side, reducing the retrieval burden on the terminal side and improving the data path efficiency of the overall system.
[0272] In this invention, the server includes a unit for encapsulating the target behavior information into behavior sequence description data that can be directly driven by the terminal to display virtual objects according to a predetermined data format. This allows for friendly adaptation of the terminal's rendering engine at the data level, enabling the terminal to reduce additional parsing and conversion calculation steps after receiving the behavior sequence description data, thereby reducing the consumption of computing resources on the terminal side and shortening the interaction response time.
[0273] "System" refers to an overall collection of devices consisting of one or more information processing devices, portable information processing devices, storage devices, and communication devices, used to execute the processing steps described in this invention.
[0274] "Computing device" refers to an electronic processing unit that has a processor and memory and is capable of executing program instructions to perform data processing and control flow, including but not limited to server devices, general-purpose computers or special-purpose processing devices.
[0275] "Information processing device" refers to an electronic device that processes data, executes applications, and works in conjunction with storage and communication devices to achieve functions such as behavioral information management, index retrieval, and generative artificial intelligence model invocation.
[0276] "Portable information processing device" refers to a mobile terminal device that has a display and an input unit, is able to interact with a server via wireless or wired communication, and performs some data processing and rendering locally, including but not limited to smartphones, tablets, or wearable terminals.
[0277] "User" refers to an entity that uses the system to watch the game, input prompts, and interact with virtual display objects. It can be a single individual or a collection of multiple individuals.
[0278] "Virtual display objects" refer to digital images that are presented graphically on a display and can demonstrate actions and provide explanations based on behavioral information, including but not limited to virtual characters, virtual figures, or other virtual images.
[0279] "Behavioral information" refers to the data set used to drive virtual display objects to perform specific actions or action sequences. It typically includes skeletal keypoint sequences, posture change parameters, time series information, and action-related descriptive information.
[0280] "Behavioral information storage area" refers to a logical or physical area pre-divided in a storage device for storing behavioral information related to multiple competitive activities, used for centralized management of behavioral information and its index information.
[0281] "Candidate behavioral information" refers to a subset of behavioral information selected from the behavioral information storage area based on user prompts, viewing video information, and index information, which serves as the input object for generative artificial intelligence models.
[0282] "Target behavioral information" refers to the behavioral information or combination of behavioral information that best meets the user's current needs after the candidate behavioral information has been evaluated and selected through a generative artificial intelligence model.
[0283] "Index information" refers to structured labeled data attached to various behavioral information to facilitate retrieval and filtering, including but not limited to competition category information, scene category information, detail information, and other identifying information used to quickly locate behavioral information.
[0284] "Competitive activities" refers to all kinds of sports or competitions involved in sports viewing scenarios, including but not limited to ball games, track and field events, or contact sports.
[0285] "Competition category information" refers to data items used to indicate the specific sports category to which a certain behavior belongs, such as classification information used to distinguish different sports like football and basketball.
[0286] "Scenario category information" refers to data items used to indicate the context category to which behavioral information belongs in competitive activities, such as offensive scenarios, defensive scenarios, set-piece scenarios, or fast break scenarios.
[0287] "Detail information" refers to data items used to represent the level of detail in behavioral information in terms of temporal resolution, degree of action breakdown, or level of detail display.
[0288] "Prompt statements" refer to text information input by users in natural language to express their viewing needs or need for action explanations, and serve as one of the main natural language inputs for generative artificial intelligence models.
[0289] "Generative artificial intelligence models" refer to models built based on artificial intelligence technologies such as deep learning, which can generate text responses, evaluate the relevance of behavioral information, or output action selection results based on input prompts and candidate behavioral information, including but not limited to large-scale language models or multimodal generative models.
[0290] "Response information" refers to the semantic content output by a generative artificial intelligence model based on prompts and candidate behavior information, used to drive explanations or descriptions, including but not limited to textual explanations, step-by-step instructions, or question-and-answer content.
[0291] "Audio information for narration" refers to audio data generated based on response information and used for audio narration by virtual display objects. It can be obtained by converting text responses using text-to-speech technology.
[0292] "Explanatory text information" refers to text data generated based on response information and used to present the narration content in text form on the display.
[0293] "Behavior sequence" refers to a sequence of actions composed of multiple frames or stages of behavioral information arranged in chronological order, used to drive virtual display objects to be displayed on the display screen in the form of continuous animation.
[0294] "Replay speed" refers to the time scaling or playback rate when playing a sequence of actions on a display, used to control the speed of the action demonstration.
[0295] "Reproduction order" refers to the temporal or logical arrangement order used when displaying multiple behavioral information or multiple behavioral fragments, and is used to determine the order in which virtual display objects perform each action.
[0296] "Explanation granularity" refers to the level of detail when explaining or describing a sequence of behaviors, including different levels of explanation such as overall overview, stage level, or action key points level.
[0297] "Evaluation information" refers to the quantitative results related to relevance, importance, or fit of candidate behavior information output by the generative artificial intelligence model, which are used by the terminal to calculate the priority of candidate behavior information.
[0298] "Priority" refers to a ranking index calculated based on evaluation information, used to characterize the order in which candidate behavioral information is selected and the weight of each candidate during the presentation and explanation process.
[0299] The "learning storage area" refers to the storage space used to accumulate the correspondence between the user's historical prompts and the corresponding target behavior information, so as to facilitate subsequent learning processing and condition adjustment.
[0300] "Learning processing" refers to the computational process of updating parameters or adjusting strategies for the extraction conditions of candidate behavior information and the generation conditions of prompt statements based on the historical correspondence stored in the learning storage area.
[0301] In one embodiment of the present invention, the server serves as the central information processing device, the terminal serves as the portable information processing device, and the user serves as the interactive subject. The three work together through a communication network to realize the virtual display object-driven demonstration and commentary of sports actions.
[0302] In terms of hardware, servers can employ rack-mounted computing devices based on general-purpose processors, such as multi-core central processing units, main memory, solid-state storage devices, and network interface controllers. In terms of software, servers can run general-purpose operating systems, such as Unix-like operating systems, and deploy scripting environments (such as Python runtime environments) and relational database management systems (such as MySQL). Servers can also deploy deep learning inference frameworks, such as tensor-based inference libraries, and load pre-trained generative artificial intelligence models.
[0303] The terminal can be a mobile computing device in terms of hardware, such as a mobile terminal with a touch screen, central processing unit, graphics processing unit, audio output unit, and wireless communication module. In terms of software, the terminal can run a mobile operating system, such as an open-source mobile operating system or a commercial mobile operating system, and run a viewing application on it. The terminal can integrate a local inference engine, such as a lightweight deep learning inference framework (e.g., a mobile tensor computation library, a lightweight inference engine, or a mobile device-specific inference framework), for performing some inference computations of generative artificial intelligence models locally.
[0304] After a user launches the viewing application on their terminal, the terminal displays a virtual display object settings interface. The terminal receives the user's touch input, text input, and selections, and integrates the user's chosen appearance parameters, voice type, and commentary style into a virtual display object configuration. The terminal sends this configuration to the server via an encrypted communication protocol, and the server creates a corresponding configuration record for each user in its database. Because the server centrally manages the virtual display object configuration, it can embed these configuration parameters into the prompt generation process and model input features in subsequent processing stages. This maintains stylistic consistency during generation and reduces the burden of repeated configuration data transmission by the terminal, achieving a technical effect of communication load reduction.
[0305] The server sets up a behavior information storage area in the database. Using a database management system, the server records the behavior identifier, competition category information, scene category information, detail information, and corresponding skeletal keypoint sequence, timestamp sequence, and semantic tags for each piece of behavior information within this storage area. The server establishes logical partitions or index tables for each competitive activity (e.g., "football," "basketball," etc.) within this behavior information storage area. The server encodes the skeletal data of the behavior information into a unified format, for example, representing the skeletal keypoints of each frame with joint numbers and three-dimensional coordinate arrays, and attaching a time step index. Because the behavior information is stored in structured data form and includes multi-dimensional index information, the server can quickly filter candidate behavior information that meets user needs through combined indexes and range queries. Compared to the traditional method of searching only with text tags, the server improves both query complexity and retrieval accuracy.
[0306] When a user enters natural language prompts on the terminal during the viewing process, the terminal collects the text data along with contextual information such as the current match type, time, and location. Examples of prompts that users can enter on the terminal include: "I want to see a detailed breakdown of the movements involved in that goal." "Please use a virtual character to demonstrate the standard penalty kick shooting motion step by step." "What's the difference between this defensive move and the standard stance?" "Please slow down the leg swing in step three and explain it in more detail." "I want to see a detailed video of a shot on goal in a football match." "Please use a virtual character to explain to me the positioning details in this counter-attack tactic." The terminal sends the prompt and context information to the server. Upon receiving the request, the server uses its natural language processing module to perform lexical analysis, syntactic analysis, and intent recognition on the prompt. The server can utilize a language representation model based on a deep bidirectional encoding structure or a context representation model based on a multi-layer transformation structure to convert the prompt into a dense vector representation. During the feature extraction stage, the server uses an attention mechanism to assign higher weights to action-type words (e.g., "shoot," "defend," "run"), scene-related words (e.g., "goal," "counterattack"), and detail-oriented words (e.g., "detailed," "step by step," "slow down"). Because the server uses learnable attention weights instead of fixed rules during the feature extraction stage, it can reliably identify key semantics even when faced with prompts that vary in word choice or expression, thereby improving the robustness and accuracy of prompt parsing.
[0307] The server constructs database query conditions based on the parsed results. For example, when the server parses a prompt containing "football," "goal," and "details," it combines the sports category information field, scene category information field, and detail information field in the behavior information storage area to form a multi-condition index query. The server can use query statements with range constraints, limiting the detail information to be greater than or equal to a certain threshold, and sort the results according to the similarity score field. This combination of data structure and query method allows the server to efficiently lock a small amount of highly relevant behavior information from a massive behavior dataset, reducing the input scale of subsequent generative artificial intelligence models and improving overall computational efficiency.
[0308] After receiving a set of candidate behavior information, the server combines the labels, text descriptions, and partial skeletal feature vectors of these candidate behavior information with the prompt statement vector to form the input of the generative artificial intelligence model. In one implementation, the generative artificial intelligence model can employ a sequence-to-sequence neural network based on a multi-layer self-attention structure. Its encoder encodes the prompt statement sequence and the candidate behavior information label sequence into semantic vectors, while the decoder generates new prompt statements and relevance scores for the candidate behavior information based on the encoding results. Internally, the model employs a multi-head attention mechanism, a multi-layer feedforward network, and a normalization layer to achieve high-dimensional feature modeling of long text prompt statements and multiple candidate information. The model uses a cross-entropy loss function to optimize the language quality of the generated prompt statements, and simultaneously uses ranking loss or contrastive loss to optimize the relevance ranking of the candidate behavior information. During the training phase, the model updates the gradients of each layer weight through backpropagation and can introduce data augmentation techniques into the training data, such as synonym substitution, word order perturbation, or partial masking of the prompt statements, to improve the model's generalization ability.
[0309] When generating new prompts, the server not only considers the user's original prompts but also encodes this information into additional conditional vectors based on the user's virtual display object configuration, the current game type, and the scene type, and then connects these vectors to the model input. The server-generated prompts thus maintain a consistent tone and style with the user-defined commentary style and more explicitly indicate the target behavioral information features, making it easier for the model to converge on the action segments that the user truly cares about when selecting candidate behavioral information. This method of reconstructing prompts on the server side essentially changes the distribution of the model input, making the subsequent generation and selection process more computationally stable and efficient.
[0310] After obtaining the relevance scores of each candidate behavior information, the server can execute a ranking algorithm locally, such as descending order based on the scores, or re-rank the candidate behavior information using a learning-based ranking algorithm. The server can select several target behavior information based on a threshold or a fixed number, and encapsulate their behavior sequence data along with supplementary descriptive text into a unified format before sending it to the terminal. In some implementations, the server can also send only the ranking results and behavior identifiers, allowing the terminal to retrieve specific behavior data from its local cache or edge nodes based on the behavior identifiers, further reducing the traffic pressure on the central server.
[0311] After receiving the target behavior information and response information from the server, the terminal uses a lightweight generative AI model locally to further refine the display of the behavior sequence. The model on the terminal can employ a transformation structure with small parameter sizes, or it can be distilled from a larger model on the server side using a teacher-student distillation method. Based on the evaluation information provided by the server and the output of the local model, the terminal calculates the priority of each behavior sequence and dynamically selects the playback speed, playback order, and explanatory granularity accordingly. For example, when the terminal determines that a certain action is a critical technical action, it can reduce the playback speed (e.g., by 0.5 times or 0.25 times) while increasing the explanatory granularity, breaking the action down into multiple keyframes and adding detailed text descriptions; for auxiliary actions, the terminal can quickly gloss over them at normal speed with minimal explanation. Through this priority-driven control on the terminal side, the terminal achieves fine-grained scheduling of the display of virtual object behavior, thereby presenting users with animations and explanations with more reasonable information density under limited screen and computing resources.
[0312] When rendering virtual display objects, the terminal calls a graphics rendering interface (such as a mobile graphics interface) and maps the skeletal keypoint sequence from the target behavior information onto the skeletal structure of the virtual display object. The terminal processes the skeletal data using interpolation and pose smoothing algorithms to reduce motion jitter and discontinuities. The terminal utilizes its graphics processing unit to perform parallel computations of vertex transformations and pixel shading, thereby reducing the load on the central processing unit while maintaining high rendering quality. Since the behavior sequence has already been filtered and compressed on the server side according to demand, the amount of data the terminal needs to process is reduced, thus achieving the technical effects of increased processing speed and reduced power consumption on the terminal side.
[0313] When the terminal outputs narration audio information, it sends the narration text from the response information to the text-to-speech module. The text-to-speech module can employ neural network-based acoustic models and vocoder models, such as sequence-to-sequence acoustic models and vocoders based on convolutional or autoregressive structures. The terminal's central processing unit or digital signal processing unit generates a speech waveform. The terminal aligns the generated speech with the action sequence timeline to ensure that the narration content is presented synchronously with the actions of the virtual displayed objects, thereby improving the user's comprehension.
[0314] During interaction, the server continuously records the correspondence between users' historical prompts and corresponding target behavior information. The server maintains a set of user profile parameters and preference vectors in the learning storage area. The server can periodically, or when a certain amount of data is reached, use this historical data to incrementally update some parameters of the generative artificial intelligence model. For example, it can fine-tune the weights of the last few layers of the model through a fine-tuning mechanism, making the model more inclined to select behavioral information consistent with the user's habits and interests when processing future prompts. The server can also group the preference vectors of multiple users based on clustering algorithms to form various typical user categories, and adopt different prompt generation strategies and candidate behavior information filtering rules for different categories. Through this historical information-driven adaptive adjustment, the server fundamentally changes the parameter space of behavior information retrieval and model invocation, enabling the system to continuously optimize over long-term operation, resulting in improvements in technical indicators such as improved prompt parsing accuracy, reduced candidate behavior information selection error, and shortened overall response time.
[0315] In another implementation, the server can avoid performing all computations of the generative AI model on its own. Instead, it can handle the parsing of prompts and the filtering of candidate behavior information on the server side, sending the simplified candidate set and structured features to the terminal, allowing the terminal to perform the computation of the generative part locally. Since the model parameters on the terminal are smaller, the server only transmits summary features rather than the complete original data over the network. This significantly reduces the amount of data transmitted across networks, alleviates network congestion, and improves interactive response performance in real-time viewing scenarios. Furthermore, when the terminal performs model inference locally, it can adjust the inference depth or precision based on the current device status (such as battery level and processor utilization), achieving adaptive allocation of computing power and further improving the overall computational efficiency of the system.
[0316] In another implementation, the server can divide the behavioral information storage area into multiple tiers of cache. For example, frequently accessed behavioral information can be stored in high-speed storage media, while low-frequency behavioral information can be stored in high-capacity storage media. Based on the high-frequency action types and scenario types appearing in the prompt statements, the server pre-loads potentially accessed behavioral information into the high-speed cache using a prediction model. In this way, when the terminal issues a prompt statement request, the server can directly return candidate behavioral information from the cache, significantly reducing disk access latency and improving the throughput of behavioral information retrieval.
[0317] As can be seen from the above embodiments, this invention, by introducing structured behavioral information management, multi-source feature fusion-based prompt statement parsing, candidate behavioral information selection and ranking based on generative artificial intelligence models, and priority-based dynamic display control of behavioral sequences, not only achieves personalized generation of virtual display object action demonstrations and narration content, but also introduces new technical solutions in terms of data structure design, algorithm flow, and resource scheduling within the computer. The server improves the accuracy and speed of behavioral data retrieval through a multi-dimensional indexed behavioral information storage structure and a deep learning-driven prompt statement parsing and selection mechanism; the terminal achieves efficient rendering and synchronous narration output under limited computing power through a lightweight generative artificial intelligence model and priority control strategy. These technical features work together to make this invention not merely a simple automation of human narration or human action demonstration, but a computer technology solution that substantially improves processing efficiency and result quality in terms of data management methods, model calling paths, and terminal display control.
[0318] use Figure 13 The processing flow is explained.
[0319] Step 1: Users launch the viewing application on their devices and set up virtual display objects.
[0320] Users tap the application icon on the terminal's touchscreen to open the virtual character settings interface, where they can select the appearance style, clothing color, voice type, and narration style.
[0321] Input: User touch operations, text selections (such as "sporty style" or "male voice narration") and other interactive information.
[0322] The terminal constructs a virtual display object configuration data in local memory based on the input, encodes each option into key-value pairs (such as style identifiers, color numbers, and voice type codes), and updates the preview image on the screen in real time.
[0323] Output: Configuration data structure of the virtual display object generated internally by the terminal.
[0324] Step 2: The terminal sends the virtual display object configuration to the server and receives confirmation.
[0325] The terminal uses an encrypted communication protocol through its communication module to package the configuration data generated in step 1 and the user identifier into a request message and send it to the server's configuration interface.
[0326] Input: Virtual display object configuration data, user ID.
[0327] After receiving the data, the server uses a parser to extract fields from the request message, uses the database interface to write the configuration to the user configuration table, and generates a write result (success / failure and error code).
[0328] Output: The server returns a configuration confirmation response, which the terminal receives and displays "Configuration saved" on the interface.
[0329] Step 3: Users can input prompts via their terminals during the viewing process.
[0330] Users type natural language prompts in the text input box on the terminal viewing interface, such as "I want to see a detailed breakdown of the goal I just scored." or "Please demonstrate the standard penalty kick step by step using a virtual character." and then click the send button.
[0331] Input: The text of the prompt entered by the user, the time and position of the current viewing screen, and the event type information.
[0332] The terminal temporarily stores this data in local memory, clears the input box on the interface, and displays the status of the prompt statement that has been sent.
[0333] Output: The terminal's internal prompt requesting data.
[0334] Step 4: The terminal sends a prompt request to the server.
[0335] The terminal uses a communication module to encapsulate the prompts, event types, timestamps, etc. prepared in step 3 into a request message, and sends it to the server's action query interface via a network protocol.
[0336] Input: Prompt text (prompt statement), event context information (sports event, current time, scene markers).
[0337] The server receives the message, performs integrity verification and decoding on it in the input buffer, and prepares it for subsequent parsing modules.
[0338] Output: The raw request data object available for further processing on the server side.
[0339] Step 5: The server parses the prompt statement and extracts its semantic features.
[0340] The server calls the natural language processing module to perform word segmentation, part-of-speech tagging, and important word extraction on the prompt statement. The server converts the prompt statement into a vector representation through an embedding layer and an encoding network (such as a bidirectional encoder or a multi-layer transform network), and assigns higher weights to keywords such as "goal", "details", and "step by step" under an attention mechanism.
[0341] Input: Prompt text, event context information.
[0342] During the encoding and computation process, the server performs matrix multiplication, weighted summation, and nonlinear activation operations on the text to obtain a high-dimensional semantic vector, from which it parses out the target movement type, action type (such as "shoot" or "defense"), scene category (such as "goal scene"), and detailed requirements.
[0343] Output: The semantic feature vector of the prompt statement and the parsed structured semantic tags (motion type, action type, scene category, and level of detail).
[0344] Step 6: The server retrieves candidate behavior information based on semantic tags.
[0345] The server uses the tags obtained in step 5, such as sports type, scene category, and level of detail, to construct database query conditions and execute multi-condition index queries in the behavior information storage area.
[0346] Input: Structured semantic tags, behavioral information database index structure.
[0347] The server generates and executes query commands through the database engine, filters behavior records stored on disk or in memory, uses competition category information fields, scene category information fields, and detail information fields for filtering, and sorts them according to the pre-stored relevance score field, and finally selects a number of candidate behavior information records.
[0348] Output: A set of candidate behavior information, including behavior identifier, skeletal keypoint sequence, time series, label and descriptive text for each record.
[0349] Step 7: The server generates prompts for generative AI models and prepares inputs for the models.
[0350] The server constructs new structured prompts based on the user's original prompts, parsed semantic tags, and tags and descriptions of candidate behavior information, to more accurately indicate the processing targets of the generative artificial intelligence model.
[0351] Input: User's original prompt statement, virtual display object configuration parameters, and labels and descriptions of candidate behavior information.
[0352] The server concatenates and formats these texts, embedding the style of the virtual display object (such as "sports commentary") and scene descriptions (such as "right-footed shot into the penalty area") into the prompt statements. At the same time, it encodes the candidate behavior information label sequence into a vector and merges it into the model input features.
[0353] Output: Input data for the generative artificial intelligence model, including a feature list of reconstructed prompt text and candidate behavior information.
[0354] Step 8: The server invokes a generative artificial intelligence model to evaluate and select candidate behavioral information based on its relevance.
[0355] The server feeds the input data constructed in step 7 into the generative artificial intelligence model. The model's encoding layer performs embedding mapping on the text and labels, and the attention layer calculates the semantic matching degree between the prompt and each candidate behavior information, outputting a relevance score for each candidate behavior information.
[0356] Input: Reconstructed prompt statement and candidate behavior information feature vector.
[0357] The server performs multiple matrix multiplications, summations, normalizations, and activation function operations within the model to calculate the corresponding score vector. It then uses sorting or threshold selection strategies to determine a set of target behavior information and generates auxiliary natural language response content (such as action explanation text).
[0358] Output: A set of target behavior information and their priority order, as well as the corresponding response text information and evaluation information.
[0359] Step 9: The server encapsulates the target behavior information and response information and sends them to the terminal.
[0360] The server packages the target behavior information selected in step 8, compresses and encodes the skeletal key point sequence, organizes the explanatory text and evaluation information into a response message structure, and sends it to the terminal through the network interface.
[0361] Input: Target behavior information data (skeletal sequence, time information), response text, evaluation information.
[0362] During the packaging process, the server performs format conversion and compression on the data to reduce the message size, and finally writes it into the sending buffer and transmits it over the network.
[0363] Output: A network response message that arrives at the terminal, containing target behavior information and response information.
[0364] Step 10: The terminal receives target behavior information and performs local priority calculation.
[0365] The terminal receives the server's response from the network layer, parses out the list of target behavior information, its evaluation information, and explanatory text. The terminal then calls its local lightweight generative artificial intelligence model, inputting the evaluation information along with local context (such as screen size, battery status, and user preferences) into the model to refine the calculation of the display priority of each segment of behavior information.
[0366] Input: Target behavior information, evaluation information, and local environment parameters returned by the server.
[0367] The terminal performs vector operations in the model to obtain the weights of each behavioral information segment in three dimensions: reproduction speed, reproduction order, and description granularity, and generates a set of display control parameters accordingly.
[0368] Output: Display control parameters corresponding to each segment of target behavior information, including playback speed, playback order, and level of detail in the description.
[0369] Step 11: The terminal drives the virtual display object to execute a sequence of behaviors based on the display control parameters.
[0370] The terminal uses the sequence of skeletal key points in the target behavior information to map it onto the skeletal model of the virtual display object, and sets the animation playback speed and order according to the display control parameters.
[0371] Input: Behavioral sequence data, display control parameters, and virtual display object skeletal structure.
[0372] The terminal uses an interpolation algorithm to smooth each keyframe, uses a graphics processing unit to perform vertex calculation and rasterization, draws the dynamic pose of the virtual display object onto the display screen, and plays multiple behavior segments in a predetermined order to meet priority requirements.
[0373] Output: A series of virtual display object motion animations presented on the terminal display screen.
[0374] Step 12: The terminal generates narration using both voice and text information based on the response text and displays it synchronously.
[0375] The terminal inputs the response text from step 9 into the local text-to-speech module, which performs text encoding, acoustic feature generation, and waveform synthesis to convert the text into an audio signal. Simultaneously, the terminal renders synchronized subtitles at the bottom of the screen, displaying the response text line by line as it is played along with the audio, segmented by time.
[0376] Input: Response text, display timeline information.
[0377] The terminal aligns the start and end times of audio playback with the animation playback time based on the timeline of the behavior sequence, and controls the audio module and rendering module to run synchronously through the clock module to ensure that the narration content is consistent with the action display.
[0378] Output: Narrative voice and subtitles synchronized with the movements of the virtual display object.
[0379] Step 13: The server records the correspondence between user prompts and target behavior information for subsequent learning.
[0380] When the server receives feedback from the terminal that the interaction has been completed or when the session ends, the server writes the original user prompt, the reconstructed prompt, the target behavior information identifier, and the display effect feedback (such as whether the user requests to slow down again or explain again) to the learning storage area.
[0381] Input: User prompts, target behavior information identifiers, and interactive feedback information.
[0382] Before writing the data, the server standardizes it, such as by unifying the encoding format, aligning timestamps, and filtering out duplicate or noisy data, and then stores it as training samples. In subsequent offline or online training, the server uses these samples to fine-tune the parameters of the generative artificial intelligence model, updating some weights through gradient descent, so that the model can more accurately select behavioral information in response to similar prompts.
[0383] Output: Updated learning storage records and optimized model parameters after the training cycle ends.
[0384] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0385] With the development of multimedia content distribution and human-computer interaction technologies, users' demands for personalization, interactivity, and a sense of presence are constantly increasing when watching sports events or other live content. However, existing computer-implemented commentary and question-and-answer systems mainly suffer from the following technical problems: First, in solutions where only simple overlay display or fixed script playback is performed on the terminal side, the server lacks the ability to deeply analyze the captured data and it is difficult to accurately infer the motion state in the scene from the image data uploaded by the terminal. Therefore, it is impossible to dynamically generate virtual object actions and commentary content that are highly matched with the actual situation, resulting in a relatively mechanical human-computer interaction experience and low utilization of multimodal input data by the computer system.
[0386] Second, traditional question-answering systems typically only perform keyword matching or rule-based processing on user text, failing to build a unified semantic parsing, emotion recognition, and generative artificial intelligence model call chain on the server side. They cannot automatically generate prompts that adapt to the context and user state, thus failing to fully leverage the generative artificial intelligence model's capabilities in natural language generation. The computer system lacks flexibility and scalability when processing natural language output with complex contexts and multi-factor constraints.
[0387] Third, existing systems are often fragmented in terms of emotion recognition and output style control: on the one hand, emotion-related information such as user facial expressions and voice features is not tightly coupled with the process of generating narration content; on the other hand, the server side lacks an output control mechanism that can uniformly adjust multi-dimensional parameters such as speaker characteristics, speech rate, tone, and level of detail in explanation. This makes it difficult to drive the virtual object's actions and language style to change in a structured way on the server side, even if the user's emotions are recognized. As a result, the computer system has insufficient control granularity in multimodal output coordination.
[0388] Fourth, many systems lack systematic modeling of long-term user interaction data. The server side does not treat historical input, response content, and emotional state as a whole data resource for statistical or machine learning processing, nor does it feed the learning results back to the prompt generation and interpretation style control stages. As a result, computer systems find it difficult to form adaptive strategies based on large-scale interaction data and cannot continuously improve the quality and efficiency of personalized output at the algorithm level.
[0389] Therefore, how to provide a unified computer implementation architecture on the server side, which can collaboratively process the shooting data, user input data and emotion data uploaded by the terminal, automatically generate prompts that are adapted to the generative artificial intelligence model, and dynamically adjust the actions and narration style of virtual objects based on learned user preferences, thereby substantially improving the computer system's processing capabilities in multimodal input understanding, natural language generation and personalized output control, has become an urgent technical issue to be solved in this field.
[0390] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0391] In this invention, the server includes: a processing unit for running on an information processing device and performing processing using computing resources; a unit for generating a display screen and providing a user interface through the display screen, the user interface enabling the user to configure selectable display objects in a computer system; an image analysis unit for receiving captured data from a portable information processing terminal with a shooting function and estimating the motion state of an object based on the captured data; an action generation unit for determining the display action of the display object based on the motion state, time information, and scene information; an action information acquisition unit for reading action information corresponding to multiple motion categories and managed according to index information from a storage device, selecting and applying the action information to the display object based on the motion state and scene information; an input parsing unit for receiving voice or text information from the user and extracting the semantic content of the user input through voice recognition processing and natural language processing; and an input parsing unit for automatically generating input for a generative artificial intelligence model based on semantic content, scene information, and user-related attribute information. The system comprises: a prompt statement generation unit that sends a request to generate a prompt statement to a generative artificial intelligence model; a response acquisition unit that receives response information from the generative artificial intelligence model and outputs it as explanatory or question-and-answer information as a display object; an emotion estimation unit that infers the user's emotional state based on facial and voice information obtained from a portable information processing terminal with shooting and recording functions; an output control unit that determines the explanation style, including speaker characteristics, speech rate, tone, and level of detail, based on the emotional state and the user's past usage history, and combines the explanation style with response and action information to adjust the display object output method; a sending unit that associates the display object, shooting data, and response information to generate display data and sends it to the portable information processing terminal; and a learning unit that associates and records the user's past input history, emotional state, and response information, extracts user preference information through statistical processing or machine learning processing, and updates the prompt statement and explanation style based on the user preference information. This allows for the construction of an integrated computer processing flow on the server side, oriented towards multimodal input and output. The computer system can automatically infer scene and user states from images, speech, and text, generate prompts adapted to generative artificial intelligence models, and finely adjust the actions and language style of virtual objects through a learnable output control mechanism. This improves the relevance and coherence of natural language generation, enhances the consistency and personalization of multimodal output, and achieves substantial improvements to existing computer technology in the field of intelligent explanation and interaction.
[0392] "Information processing device" refers to an electronic device with computing and storage resources, used to execute program code and process input data, including but not limited to server devices, terminal devices or other computing devices.
[0393] "Processing unit" refers to a hardware circuit or software module that operates in an information processing device and is used to call upon computing and storage resources to perform a predetermined function, including but not limited to a central processing unit, a graphics processing unit, or a control program executed by it.
[0394] "Display object" refers to a digital virtual entity generated by an information processing device and presented on a display screen, used to output explanatory or question-and-answer information to the user, including but not limited to virtual characters, virtual images or other graphic elements.
[0395] "User interface" refers to an interactive interface consisting of display screens and input controls, used to receive user operation commands and present system status to the user, including but not limited to menu interfaces, settings interfaces, and interactive windows.
[0396] "Portable information processing terminal" refers to an electronic terminal device that can be carried by a user and has computing, display, shooting and / or radio functions, including but not limited to mobile phones, tablet devices or wearable devices.
[0397] "Shooting data" refers to image or video data obtained by a terminal with shooting capabilities through a camera component, used to reflect the state of objects, environment, or motion in a real-world scene.
[0398] The "image analysis unit" refers to a functional module used to receive captured data and analyze and process the captured data in order to infer the motion state or other feature information of objects in the scene.
[0399] "State of motion" refers to the characteristic information that describes the changes of an object in space and time, including but not limited to changes in position, velocity, acceleration, attitude, or event type.
[0400] "Time information" refers to a time identifier associated with the captured data or the moment an event occurs, including but not limited to timestamps, match times, or system clock information.
[0401] "Scene information" refers to environmental description data related to the current environment, content category, or context state, including but not limited to sports type, score status, event type, or content tag.
[0402] The "action generation unit" refers to a functional module that determines and outputs the display actions that the display object should perform based on motion state, time information, and scene information.
[0403] "Storage device" means a hardware component used to store data or program code in a non-transitory manner, including but not limited to semiconductor memory, magnetic storage medium or optical storage medium.
[0404] "Motion information" refers to structured data used to define the motion patterns performed by display objects or other objects on the display screen, including but not limited to skeletal animation data, keyframe data, facial expression data, or motion parameters.
[0405] "Sports Category" refers to the classification information that divides the scene according to different sports or activity types, including but not limited to football, basketball, tennis or other sports.
[0406] "Index information" refers to identification information or retrieval keys used to retrieve and locate target data in a storage device, including but not limited to numbers, tags, key-value pairs, or combinations of conditions.
[0407] The "motion information acquisition unit" refers to a functional module that accesses the storage device and selects and reads motion information suitable for displaying the object based on the motion state, scene information, and index information.
[0408] "Voice information" refers to user voice data acquired by a microphone and represented in the form of an audio signal, used to express the user's questions, instructions, or emotions.
[0409] “Text information” refers to user input data represented in the form of a character sequence, including but not limited to keyboard input, touch input, or text converted from speech recognition.
[0410] The "input parsing unit" refers to the functional module that performs speech recognition and natural language processing on the user's voice or text information to extract semantic content and intent information.
[0411] "Semantic content" refers to structured or semi-structured information extracted from user input by the input parsing unit to express the user's intent or the meaning of the question.
[0412] "Attribute information" refers to characteristic data related to users, including but not limited to user identity information, historical behavioral characteristics, preference settings, or device environment information.
[0413] "Generative artificial intelligence models" refer to artificial intelligence models trained using machine learning methods that can automatically generate text responses or other output content based on input prompts.
[0414] "Prompt statements" refer to input text constructed to guide generative artificial intelligence models in generating output that conforms to the desired content and style, and include descriptions of the scenario, user state, and output requirements.
[0415] The "prompt statement generation unit" refers to a functional module that automatically constructs prompt statements based on semantic content, scene information, and attribute information, and forms a generation request to the generative artificial intelligence model.
[0416] "Response information" refers to text or other forms of information generated and output by a generative artificial intelligence model based on prompts, used to answer user input or provide explanations.
[0417] The “response acquisition unit” refers to the functional module that receives response information from the generative artificial intelligence model and provides the response information to subsequent output processing.
[0418] "Face information" refers to images or feature data acquired by the camera and used to reflect the user's facial muscle shape, facial expression changes, and other emotion-related characteristics.
[0419] "Emotional state" refers to the user's emotional tendency inferred based on facial expressions and / or voice information, including but not limited to emotional categories such as joy, sadness, anger, calmness, or tension.
[0420] The “emotion estimation unit” refers to a functional module that analyzes a user’s facial expressions and voice information to estimate and output the user’s emotional state.
[0421] "Commentary style" refers to a set of parameters used to describe how a displayed object expresses narration information, including but not limited to speaker characteristics, speaking speed, tone of voice, and level of detail in the explanation.
[0422] The "output control unit" refers to a functional module that adjusts the output method of the displayed object and controls the narration style and action presentation based on emotional state, user history, response information, and action information.
[0423] The “sending unit” refers to the functional module that sends the generated data for display to the portable information processing terminal via a communication link.
[0424] "Data used for display" refers to comprehensive display data that includes the image of the displayed object, overlay information of the captured data, response information and related control parameters, and is used to present the final image on the terminal.
[0425] "Input history" refers to the collection of voice or text information submitted by users to the system and their parsing results within a certain time frame.
[0426] "Preference information" refers to feature information that reflects a user's long-term preferences or behavioral patterns, obtained through statistical analysis or machine learning of user input history, emotional state, and response information.
[0427] A “learning unit” refers to a functional module that records and processes user-related data, extracts preference information, and updates prompts and explanation styles based on the preference information to achieve personalized output.
[0428] In this invention, the server operates as an information processing device, and the terminal operates as a portable information processing terminal. The user interacts with the server through the terminal. The server and terminal work together to execute multiple software modules, thereby realizing a multimodal explanation and question-and-answer system based on a generative artificial intelligence model and prompt statements.
[0429] In one embodiment, the server employs a general-purpose server hardware architecture, including a multi-core central processing unit, a graphics processing unit, high-speed main memory, and non-volatile storage devices. The server runs application services, database management programs, and deep learning inference services on an operating system (such as a Unix-like operating system). In another embodiment, the terminal is a mobile terminal device equipped with a camera, microphone, display screen, and communication module, and runs dedicated applications on a mobile operating system. Users select display objects, input voice or text, and view enhanced display content through a user interface on the terminal.
[0430] In one implementation, the server utilizes multiple functional modules, including an image parsing module, an action information management module, a speech and text parsing module, a prompt statement generation module, a generative artificial intelligence model inference module, an emotion inference module, an output control module, and a learning module. The server stores action information, user attribute information, user preference information, prompt statement templates, and model parameters in structured data format within its storage device.
[0431] In one implementation, the server utilizes an image processing software library (e.g., a general-purpose image processing library) to perform multi-level feature extraction on the captured data uploaded by the terminal. At the frame level, the server performs preprocessing calculations such as color space conversion, edge detection, and motion vector estimation. Then, it inputs the feature sequence within the time window into an event recognition model composed of a convolutional neural network and a temporal network. In this model, the server uses a multi-layer convolutional network to extract spatial features and a long short-term memory network or gated recurrent units to extract temporal relationship features, thereby inferring whether a specific motion event (e.g., a goal, a foul, a pause) has occurred in the scene. At each estimation, the server outputs an event category label and confidence score, and simultaneously generates a motion state vector containing multi-dimensional values such as object position, motion direction, and velocity estimation.
[0432] In one implementation, the server manages motion information through a database management system. The server groups motion information according to motion category, event type, and emotional state, assigning an index key to each group. The server records motion identifiers, animation resource paths, parameter vectors (e.g., motion amplitude, duration), and joint trajectories corresponding to the skeletal structure of the displayed object in a storage table. Upon receiving motion state vectors and scene information, the server generates index information based on motion category, event type, and current emotional state, and selects the motion information with the highest matching degree through a database query. During selection, the server can sort candidates based on a distance function between the motion state vector and the candidate motion template, thereby improving the accuracy of motion matching. In this process, the server reduces runtime computation through pre-computed indexes and vectorized retrieval, thus reducing retrieval latency.
[0433] In one embodiment, the server parses user input via voice and text information uploaded from the terminal. In this embodiment, the terminal invokes a speech recognition service to convert the audio signal into a text sequence, which is then sent to the server. The server uses a natural language processing library to perform word segmentation, part-of-speech tagging, dependency parsing, and intent recognition calculations on the text sequence. During this process, the server constructs a semantic graph structure, representing entities (e.g., "offside," "this player," "this season"), predicates (e.g., "what is it," "how many goals were scored"), and modifiers in the user input as graph nodes and edges, thereby extracting the user's question type and core query objective. After obtaining the semantic structure, the server generates a semantic content vector, which is used as input when constructing subsequent prompt statements.
[0434] In one implementation, the server calculates the user's emotional state by combining facial expression information and speech features through an emotion inference module. After the terminal uploads an image sequence containing the user's facial region and speech signal features, the server executes convolutional neural network inference in the facial emotion recognition model, outputting an emotion category probability distribution. Simultaneously, the server uses acoustic features such as Mel-frequency cepstral coefficients, fundamental frequency changes, and energy changes in the acoustic model, via a multilayer perceptron network, to perform emotion discrimination. The server weights and fuses the emotion distributions output by the visual and acoustic models to obtain a unified emotion state label and confidence level. The server uses weight parameters during fusion to adjust the proportion of visual and audio influence on emotion determination according to different application environments. Through this multimodal emotion inference, the server can estimate the user's emotional state with high accuracy, thus providing reliable input for subsequent narration style control.
[0435] In one implementation, the server automatically constructs prompt statements based on semantic content, scene information, and attribute information through a prompt statement generation module. The server first selects a suitable base template from a set of prompt statement templates, such as a rule explanation template, a data query template, or a tactical analysis template. The server reserves variable positions in the template (e.g., "current event," "score," "user sentiment," "description length," etc.), and then fills the template with motion state vectors, scene information, user attribute information (including preference information), and sentiment state to form the specific prompt statement text. During construction, the server can specify the output language, tone requirements, and length range, thereby constraining the output space of the generative AI model. For example, in a rule explanation scenario, the server can generate the following prompt statement: "A user is watching a football match and asks you: 'What is the offside rule?' Please explain the offside rule in simple Chinese, within two to three sentences, and give a simple example." The server can generate the following prompts during goal commentary: "You are a passionate football commentator. Current information is as follows:" - Match time: 67th minute Event: Home team's number 10 player scores with a long-range shot from outside the penalty area. Current score: Home team 2 – 1 Away team - User sentiment: Joy Please generate one to two sentences of narration in Simplified Chinese. The tone should be enthusiastic and engaging, but avoid using vulgar language. In some implementations, the server can also construct prompts for scenarios with a comforting tone, for example: "A user is watching a football match, and their team is currently trailing 0-2. The user is feeling disappointed. The user asks: 'How many goals has this striker scored this season?' The query result is: 10 league goals this season. Please answer the number of goals in Simplified Chinese, adding a touch of comfort and encouragement in 2-3 sentences." In one implementation, the server uses a generative AI model based on a self-attention mechanism as the core of text generation. The server can deploy a multi-layered transformer-structured language model containing multiple self-attention and feedforward sub-layers, modeling long-distance dependencies through a multi-head attention mechanism. During the inference phase, the server receives prompts as input, converts them into a sequence of subwords, processes them through embedding layers and positional encoding, and then inputs them into the encoder and decoder structure. The decoder uses a self-attention module at each layer to focus on the generated parts, while simultaneously utilizing a cross-attention module to obtain prompt information from the encoder output, thus generating multiple rounds of output. At each time step, the server calculates the probability distribution of the next word through a softmax layer, selecting the word with the highest probability or using a beam search algorithm to generate multiple candidate outputs to improve the coherence and diversity of the generation. During sampling, the server can control the randomness and repetition of the output based on parameters such as temperature and repetition penalty, thereby further refining the discourse style.
[0436] The server trains this generative AI model using a large-scale text corpus, incorporating a cross-entropy loss function as the error function, and updating model parameters through stochastic gradient descent or adaptive learning rate optimization algorithms. In some implementations, the server can use user-interactive response data for fine-tuning training, further optimizing the model's adaptability to specific domains through supervised learning or reinforcement learning methods. In this approach, the server calculates the loss between the generated output and the target output during training and updates the weights based on the gradient of the error function, thereby gradually reducing generation error and improving the accuracy of explanations and question answers.
[0437] In one implementation, the server achieves multi-dimensional adjustment of the narration style through an output control module. After receiving emotional state, user preference information, the current prompt, and the generated response information, the server calculates a target narration style parameter vector. This parameter vector includes speech rate coefficients, pitch adjustment values, volume range, formality of language, and level of explanation detail. The server uses this parameter vector to perform style post-processing on the response text, such as adding or removing explanatory sentences, adjusting the frequency of interjections, and selecting expressions that better align with user preferences. The server then sends the adjusted text to the terminal for speech synthesis. Through this structured parameter control method, the server can achieve fine-grained scheduling of the output style, achieving a balance between consistency and personalization in multimodal output compared to traditional fixed-template output.
[0438] In one embodiment, the terminal runs a rendering software environment, such as a graphics rendering engine. After receiving action information and response text from the server, the terminal loads the corresponding model and animation resources for the display object. In this embodiment, the terminal creates a virtual character object in the graphics engine scene and maps the server-specified motion trajectory to the character's skeletal system. Simultaneously, the terminal invokes a speech synthesis module to convert the server's response text into an audio signal. While playing the audio, the terminal synchronously drives the character's lip movements based on speech energy and phoneme boundary information, thus achieving visual consistency between speech and lip movements. During the image synthesis stage, the terminal uses the current frame from the camera as the background, renders the virtual character at a specified foreground position, and controls the character's facial expressions, movement amplitude, and camera position according to parameters provided by the server. This method couples the server's abstract action information with real-world footage, achieving augmented reality presentation.
[0439] In one implementation, the user first selects the display object and sets preferences through the terminal's user interface, such as choosing the commentary style as "enthusiastic" or "calm," and the description length as "brief" or "detailed." The user then holds the terminal up to the sports event screen and watches a virtual character overlaid on the screen providing commentary in real time. When the user asks a question, they simply speak natural language into the terminal, such as "What is the offside rule?" or "How many goals has this player scored this season?" In this implementation, the terminal automatically performs voice capture and uploading, while the server executes the aforementioned semantic parsing and generative artificial intelligence model inference process in the background, ultimately presenting the response to the user through the display object.
[0440] In one implementation, the server uses a learning module to continuously record and analyze user input history, emotional states, and response information. The server establishes user session tables and interaction record tables in the database, recording fields such as timestamps, question categories, answer categories, user emotion tags, and response lengths. The server periodically extracts features from these records and uses clustering algorithms, collaborative filtering algorithms, or deep representation learning methods to group users and model their preferences. After obtaining user preference information, the server writes it into a user attribute information structure and calls this information during the prompt generation stage. For example, when the server detects that a user has repeatedly requested detailed explanations of rules, it can add a requirement to "appropriately increase the explanation length and examples" when constructing the prompt. Through this closed-loop learning mechanism, the server gradually adjusts the prompts and explanation style, making the output of the generative artificial intelligence model more suitable for specific user groups, thereby improving the relevance and satisfaction of responses at the system level.
[0441] In this invention, the server achieves multimodal fusion processing of image, voice, and text data through the aforementioned modular structure and data flow design. Compared to traditional systems that rely solely on manually written narration scripts or fixed rules, this invention improves the accuracy of event recognition through the multi-layer neural network structure of the image analysis unit, enhances the robustness of emotion recognition through multimodal fusion in the emotion inference module, significantly enriches the diversity of natural language output through the collaborative work of the prompt generation unit and the generative artificial intelligence model, and automatically optimizes the output strategy within the system through continuous updates of preference information by the learning unit. Because the server uses an index-managed action information database and hierarchical caching technology in its processing flow, it can reduce retrieval latency and communication load while ensuring result quality, thereby improving the overall system response speed.
[0442] In this invention, the terminal is responsible not only for data acquisition and display but also for performing some preprocessing and rendering calculations locally. This distributed computing structure eliminates the need to upload most of the raw image and audio data to the server, reducing network bandwidth consumption. Even with only the characteristic captured data and compressed speech features, the server can still perform high-precision calculations using a deep learning model, thus achieving a balance between communication efficiency and inference accuracy. This data flow design and module division demonstrate the invention's technological improvement in the coordinated utilization of computing and communication resources.
[0443] In various implementations, the server can employ different generative AI model structures, such as transformer networks with different numbers of layers, attention heads, or hidden dimensions. The server can also be replaced with other types of autoregressive language models or encoder-decoder models. Regarding loss functions, the server can not only use standard cross-entropy loss but also introduce dialogue coherence loss, style consistency constraints, or reward-signal-based policy gradient methods to further improve generation quality. For data augmentation, the server can utilize methods such as synonym replacement, sentence rewriting, and noise injection to expand the training corpus, thereby enhancing the model's adaptability to diverse inputs. Through these variations, this invention can be flexibly configured for different application scenarios while maintaining the overall structure.
[0444] The terminal can employ different rendering architectures in various implementation forms. When the terminal has high graphics processing capabilities, it can perform complex 3D character rendering and physical animation locally; when graphics processing capabilities are limited, the terminal can perform only 2D texture animation or play pre-rendered sequences. Regardless of the rendering scheme used, the terminal drives the displayed objects based on the motion information and output control parameters provided by the server, thereby maintaining the temporal consistency between the virtual narration and the real-world scene.
[0445] In summary, through the specific hardware and software configurations described above, combined with generative artificial intelligence models, prompt statement construction, emotion inference, and learning mechanisms, the server not only achieves automated imitation of human interpretation behavior, but also proposes a new technical solution in terms of computer internal structure, data management, and multimodal processing flow. As a result, it has achieved substantial technical results compared to traditional manual operations and simple rule systems in terms of processing speed, recognition accuracy, output diversity, and personalization.
[0446] use Figure 14 The processing flow is explained.
[0447] Step 1: Users launch the application on the terminal and set the display objects.
[0448] Input: User's touch operation, and a list of preset roles stored in the terminal's local storage.
[0449] Users click the application icon on the terminal's main screen, and after the application loads, a role selection interface is displayed. Users can swipe to browse thumbnails of multiple displayed objects, click on an object, and select a style option (such as "Enthusiastic Explanation" or "Calm Analysis").
[0450] Based on user click events, the terminal reads the identifier and default parameters of the corresponding display object from the local configuration file or cache, and generates a role configuration data structure.
[0451] Output: Character configuration data containing display object identifiers and style parameters, ready to be sent to the server.
[0452] Step 2: The terminal sends the role configuration data to the server, and the server stores the user configuration.
[0453] Input: Character configuration data generated in step 1.
[0454] The terminal encapsulates the role configuration data into a request message through the network communication module and sends it to the server interface using an encrypted transmission protocol.
[0455] After receiving the request, the server parses the message body, extracts the user identifier, display object identifier, and style parameters, calls the database interface, and writes this data into the user configuration table (by inserting or updating records).
[0456] The server performs data validation during write operations, such as checking whether the display object identifier exists and whether the style parameters are within the allowed range.
[0457] Output: User configuration records on the server side, and confirmation responses received on the terminal side.
[0458] Step 3: Users point their devices at actual scenes or event footage, and the devices collect and preprocess the captured data.
[0459] Input: Raw video stream captured by the camera sensor.
[0460] The user raises the terminal and points it at the sports event venue or display device, and the terminal calls the camera interface to continuously acquire image frames.
[0461] The terminal performs scaling, compression, and basic filtering on each frame to remove noise and reduce resolution to decrease bandwidth. The terminal can extract regions containing significant motion using simple motion detection algorithms (such as inter-frame differencing) to generate a cropped key image.
[0462] Output: A preprocessed sequence of video frames or keyframe images, along with corresponding timestamps and basic scene information, ready to be sent to the server.
[0463] Step 4: The terminal sends the pre-processed shooting data to the server, which then performs image analysis and motion state estimation.
[0464] Input: The preprocessed image sequence and timestamps output from step 3.
[0465] The terminal packages the image data and timestamp and transmits it to the server over the network.
[0466] After receiving the data, the server uses an image processing library to further process the image, extracting edges, contours, and feature points. The server inputs the features of consecutive frames into the event recognition model (including convolutional networks and temporal networks), and calculates the event category probability and motion state vector through matrix multiplication and nonlinear activation operations.
[0467] The motion state vector contains the object's position coordinates, velocity estimate, direction information, and event labels (such as "goal", "shot", "foul").
[0468] Output: Motion state vectors and scene event descriptions generated on the server side, stored in memory or cache.
[0469] Step 5: The server selects suitable motion information from the motion information database based on the motion status and scene information.
[0470] Input: Motion state vector and event label obtained in step 4, and action information records in the database.
[0471] The server first forms query conditions based on the motion category and event tag, and then performs an index query in the database to retrieve candidate action information.
[0472] The server calculates the similarity between the parameter vector and the motion state vector of each candidate action (e.g., through Euclidean distance or cosine similarity), and selects the one or more records with the highest similarity as the final action information.
[0473] Output: Action information matching the current event, including animation identifiers, bone trajectories, and parameter vectors.
[0474] Step 6: Users input questions or commands into the terminal via voice or text, and the terminal performs voice recognition and generates text input.
[0475] Input: User's voice signal or keyboard text input.
[0476] Users speak questions into the terminal, such as "What is the offside rule?" or "What is the score now?".
[0477] The terminal uses a microphone to collect audio signals and transmits the audio segments to the speech recognition service. After feature extraction and acoustic model inference, the audio is converted into a text sequence. If the user directly inputs text, the terminal directly obtains the text string.
[0478] Output: The text form of the user's input question or instruction, along with a timestamp and session identifier.
[0479] Step 7: The terminal sends the user's text input to the server, which then performs input parsing and generates semantic content.
[0480] Input: User text input generated in step 6.
[0481] After receiving the text, the server uses a natural language processing module to segment and tag the text with words, then constructs a dependency syntax structure to identify key entities (such as "offside", "this player", "this season") and predicates (such as "what is it", "how many goals were scored").
[0482] The server encodes these structured results into semantic content vectors, which include question type labels (rule query, data query, etc.) and relevant parameters (target entity, time range, etc.).
[0483] Output: A semantic content vector with question type and parameters, used for constructing subsequent prompt statements.
[0484] Step 8: The terminal collects the user's facial expressions and voice features, and the server infers the user's emotional state.
[0485] Input: Sequence of user facial images captured by the terminal, and acoustic features of user speech.
[0486] Users naturally express facial expressions and tone of voice while watching, and the terminal acquires facial images through the front-facing camera and continuously collects voice signals through the microphone.
[0487] The terminal crops the image, keeping only the facial area, and sends it to the server, while also extracting and uploading basic acoustic features.
[0488] The server inputs facial images into the emotion recognition network, extracts facial expression features through convolution and pooling operations, and outputs the probability of each emotion category; the server inputs acoustic features into the emotion classification model to obtain the audio emotion probability distribution.
[0489] The server performs a weighted average or other fusion operation on the visual and audio emotion distributions to obtain the final emotion label and confidence score.
[0490] Output: The user's current emotional state (such as joy, disappointment, calmness, etc.) and its confidence level.
[0491] Step 9: The server combines semantic content, contextual information, and emotional state to generate prompts.
[0492] Inputs: Scene event information from step 5, semantic content vector from step 7, emotional state from step 8, and user attributes and preference information.
[0493] The server selects the corresponding template from the set of prompt statement templates based on the question type, and fills the placeholders with entity and event information from the semantic content; at the same time, the server reads the user's preferences (such as whether they prefer detailed explanations or brief answers) and sets the description length and tone requirements in the template.
[0494] The server then adjusts the tone of the description based on the emotional state. For example, it requires a comforting tone for the emotion of "disappointment" and a more enthusiastic expression for the emotion of "joy".
[0495] The server ultimately generates a prompt text in natural language, for example: "A user is watching a football match and asks you: 'What is the offside rule?' Please explain the offside rule in simple Chinese, within two to three sentences, and give a simple example." Output: A complete prompt text that drives the generative AI model to generate a response.
[0496] Step 10: The server invokes a generative artificial intelligence model to generate response information based on the prompts.
[0497] Input: The prompt text generated in step 9.
[0498] The server converts the prompt statement into a sequence of subwords, which are then embedded and positionally encoded before being input into a generative artificial intelligence model with a transformer structure.
[0499] The model performs multi-head self-attention computation on the prompts in the encoder to form a contextual representation, and generates the output word sequence step by step in the decoder. At each decoding time step, the model calculates the word distribution through the attention weight matrix based on the output of the previous time step and the encoder representation, and selects the next word through sampling or bundle search.
[0500] The server limits the maximum output length during generation and controls randomness based on temperature parameters to prevent excessively long or repetitive outputs. After generation, the server restores the labeled sequence output by the model into a natural language sentence.
[0501] Output: A natural language response text, such as an explanation of the rules or an answer to a data question.
[0502] Step 11: The server adjusts the style of response text and action information based on emotional state and preferences through the output control module.
[0503] Input: Response text generated in step 10, action information selected in step 5, emotional state in step 8, and user preference information.
[0504] The server first calculates a vector of explanatory style parameters based on emotional state and preference information, such as setting the level of detail in the explanation, whether to add comforting words, and the strength of the tone.
[0505] The server performs post-processing on the response text, such as adding encouraging sentences when the emotion is "disappointed" and adding exciting words when the emotion is "joyful"; the server expands or compresses the description content as necessary based on the explanation detail parameter.
[0506] The server also adjusts the action information based on the emotional state, such as increasing the amplitude of the action or changing the facial expression parameters.
[0507] Output: Styled response text and updated action information parameters, ready to be sent to the terminal.
[0508] Step 12: The server packages the action information and response text and sends them to the terminal, which then performs rendering and speech synthesis.
[0509] Input: Action information and response text output from step 11.
[0510] The server encapsulates the above data into a response message and sends it to the terminal over the network.
[0511] After receiving the data, the terminal uses a graphics rendering engine to load the display object model and specified animation, and sets the motion trajectory and facial expression. The terminal calls the text-to-speech module to convert the response text into an audio signal, and configures lip-sync animation for the display object according to the phoneme boundaries.
[0512] During the compositing stage, the terminal uses the current frame from the camera as the background, overlays and renders the display object with actions and lip movements onto the specified position, and plays audio at the same time.
[0513] Output: Virtual narration displayed on the terminal screen and voice narration output through the speaker; users see and hear dynamic responses.
[0514] Step 13: The server records interaction data and updates user preferences for adaptive adjustments to subsequent prompts and styles.
[0515] Input: Metadata such as user input text, emotional state, generated response text, narration style parameters, and timestamps from this round of interaction.
[0516] The server writes this data into an interaction log table, and after long-term accumulation, periodically performs statistical analysis or machine learning processing to calculate features such as the distribution of frequently asked user questions, common emotional change patterns, and responses to different styles.
[0517] The server updates user preference information based on these characteristics, for example, marking a user as "preferring detailed explanations and a calm tone", and saving the preference information back to the user attribute database.
[0518] The server reads the updated preference information when generating prompts and selecting narration styles, making the output of the generative artificial intelligence model more closely resemble the user's, thereby achieving long-term adaptive optimization of natural language generation behavior in a technical sense.
[0519] Output: Updated user preference information and interaction statistics, providing more accurate input conditions for the next interaction.
[0520] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0521] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0522] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0523] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0524] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0525] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0526] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0527] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0528] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0529] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0530] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0531] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0532] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0533] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0534] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0535] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0536] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0537] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0538] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0539] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0540] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0541] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0542] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0543] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0544] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0545] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0546] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0547] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0548] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0549] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0550] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0551] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0552] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0553] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0554] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0555] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0556] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0557] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0558] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0559] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0560] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0561] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0562] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0563] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0564] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0565] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0566] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0567] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0568] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0569] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0570] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0571] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0572] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0573] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0574] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0575] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0576] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0577] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0578] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0579] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0580] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0581] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0582] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0583] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0584] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0585] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0586] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0587] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0588] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0589] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0590] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0591] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0592] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0593] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0594] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0595] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0596] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0597] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0598] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0599] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0600] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0601] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0602] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0603] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0604] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0605] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0606] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0607] In addition, the following notes are provided in response to the above explanation.
[0608] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for providing a user interface to a user in an information processing apparatus, enabling the user to select a virtual display object and set the attributes and dialog style of the virtual display object; An apparatus for acquiring image information related to a competitive activity obtained by an imaging device in an information processing device, applying image processing algorithms and recognition algorithms to the image information to extract competitive state information, and generating motion information of the virtual display object based on the competitive state information. A device for transmitting the competitive status information to an external device via a communication device in an information processing device, receiving response information that specifically corresponds to the competitive status information from an information set containing action information corresponding to multiple competitive categories stored in the external device, and reflecting the response information in the display control of the virtual display object; An apparatus for acquiring natural language query information input by a user in an information processing device, and generating a string of prompt statements for input to a generative artificial intelligence model based on the query information, the competition status information, and the attribute information of the virtual display object. A device for inputting the prompt statement, the competition status information, and statistical information into a generative artificial intelligence model in an information processing device, generating response text information for the inquiry information by the generative artificial intelligence model, and converting the response text information into output data for explanation through the virtual display object; An apparatus for analyzing a user's facial expressions, voice information, and operational information in an information processing device to infer the user's emotional state, and for changing the speaker settings and expression styles in the prompt statements based on the emotional state and the attribute information of the virtual display object, and for adjusting the action information and voice output style of the virtual display object. An apparatus for recording a user's past inquiry information, viewing history information, and operation history information in an information processing device, and for inferring the user's preference information based on the recorded information, and for performing learning processing to generate the prompt statements and update the explanation style of the virtual display object according to the preference information.
[0609] (Note 2) According to the information processing system described in Appendix 1, the information processing device retrieves information based on index information associated with the action information stored in the external device, which corresponds to multiple competition categories and competition scene categories, and specifies the action information based on the competition category information and scene category information included in the competition status information.
[0610] (Note 3) According to the information processing system described in Appendix 1, the information processing device, when generating a prompt statement input to the generative artificial intelligence model, includes the current inquiry information, dialogue history information, and historical information obtained by summarizing the user preference information inferred by the learning process in the prompt statement, so that the generative artificial intelligence model generates response text information consistent with multi-turn dialogue.
[0611] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: This is a means of providing users with the selection of virtual character attributes and explanation guidelines through an operating interface on a display device in an information processing device, and associating the user's selection results with user identification information to generate and store them as setting information in the storage unit. Means for acquiring image data of a sports competition from the camera unit or image acquisition unit of a portable terminal device, attaching time information to the image data and receiving the image data, and generating event extraction information for determining the type of competition and competition events based on the image data. A means for obtaining action information and rule information corresponding to a competition event by performing a retrieval process using index information based on the event extraction information and the setting information, for an action information storage unit containing action information stored corresponding to multiple competition types and multiple competition events; The method is used to generate prompt statements for inputting into a generative artificial intelligence model based on the virtual character attributes, event extraction information, action information and rule information contained in the setting information, and to instruct the generative artificial intelligence model on the means of generating explanatory responses based on the prompt statements. A means for obtaining the narration response output from the generative artificial intelligence model as text data, and overlaying the text data onto each image frame of the image data using image processing software, thereby generating display image data with additional live narration information. This is a means of obtaining voice or text input from a user, performing speech recognition processing on the voice input to generate question text, and generating a prompt statement for question answering input to the generative artificial intelligence model based on the question text combined with the event extraction information and the action information, and generating a response for question answering based on the prompt statement. This is a means of performing sentiment inference processing based on the user's facial expression information, voice feature information, or operation history information, and adjusting the style, output length, and narration tendency of the prompt statement according to the sentiment inference result and the user's past behavior information, thereby dynamically changing the virtual character's live narration style and response style.
[0612] (Note 2) According to the information processing system described in Appendix 1, the action information storage unit is configured to: store multiple types of competition events and their corresponding action information in a hierarchical manner according to multiple different competition types, and switch the scope of retrieval using the index information based on the competition type information and time information contained in the event extraction information, thereby selecting action information corresponding to a predetermined competition event.
[0613] (Note 3) According to the information processing system described in Appendix 1, the information processing device is configured to: store the setting information, the question text, the response for question-and-answer, and the sentiment inference result as historical information in a storage unit according to the user; perform statistical processing or machine learning processing on the historical information to update the user preference model; and automatically adjust the explanation content, professionalism, and expression method contained in the prompt statement based on the updated preference model, thereby implementing personalized live explanation and question-and-answer for each user.
[0614] Example 2 (Note 1) An information processing system, characterized in that it comprises: A unit for providing users with an information input / output interface via a computing device for setting preferred virtual display objects; A unit for analyzing video information for watching sports or shooting information of sports space acquired by a portable information processing device through a computing device, and generating behavioral information of the virtual display object based on the analysis results; A unit for managing a behavior information storage area containing behavior information corresponding to various competitive activities through an information processing device, and extracting candidate behavior information associated with the video information or the shooting information and the user's instructions from the behavior information storage area; A unit for generating prompts for inputting into a generative artificial intelligence model based on natural language prompts obtained from the user and the candidate behavior information through an information processing device, and inputting the generated prompts together with the candidate behavior information into the generative artificial intelligence model so that the generative artificial intelligence model selects behavior information optimized for user needs from the candidate behavior information. A unit for determining a specific sequence of behaviors of a virtual display object based on optimized behavior information output from the generative artificial intelligence model using a portable information processing device, and applying the sequence of behaviors to the virtual display object for reproduction on the display unit of the portable information processing device; A unit for generating, via a portable information processing device, explanatory voice information or explanatory text information by the virtual display object based on response information output from the generative artificial intelligence model, and providing time-corresponding prompts to the reproduction of the behavior sequence; A unit for performing learning processing by accumulating the correspondence between past prompts and the optimized behavioral information of a user in a learning storage area through an information processing device, and adjusting the extraction conditions of subsequent candidate behavioral information or the generation conditions of the prompts based on the accumulation results.
[0615] (Note 2) According to the information processing system described in Appendix 1, the behavior information storage area stores index information including competition category information, scene category information, and detail information for behavior information corresponding to various competitive activities. The information processing device identifies the candidate behavior information by referring to the index information based on the parsing results of the prompt statement and the video information during the viewing.
[0616] (Note 3) According to the information processing system described in Appendix 1, the portable information processing device calculates the priority of multiple candidate behavioral information based on the evaluation information obtained from the generative artificial intelligence model, and controls the reproduction speed, reproduction order and explanatory granularity according to the priority, thereby dynamically changing the prompting method of the behavioral sequence performed by the virtual display object.
[0617] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A processing unit that runs on an information processing device and performs processing using computing resources; A unit for generating a display screen and providing a user interface through the display screen, wherein the user interface is used to allow a user to set selectable display objects; An image analysis unit for receiving captured data from a portable information processing terminal with a shooting function and estimating the motion state of an object based on the captured data; An action generation unit for determining the display action of the display object based on the motion state, time information, and scene information inferred by the image analysis unit; The unit is used to access a storage device, read motion information corresponding to multiple motion categories from the storage device, and select the motion information using index information based on the motion state and the scene information, and apply the motion information to the motion information acquisition unit of the display object. An input parsing unit is used to receive voice or text information from users and to parse user input through speech recognition and natural language processing, thereby extracting the semantic content of user input. The system is used to generate a prompt statement for input to a generative artificial intelligence model based on the semantic content, the scene information, and attribute information related to the user, and to send a generation request containing the prompt statement to the prompt statement generation unit of the generative artificial intelligence model. A response acquisition unit for receiving response information from the generative artificial intelligence model and using the response information as response information for explanatory information or question-and-answer information output by the display object; An emotion estimation unit for estimating a user's emotional state based on facial expression and voice information obtained from a portable information processing terminal with camera and radio functions. An output control unit is used to determine the explanation style, including speaker characteristics, speech rate, tone and level of detail, based on the emotional state and the user's past usage history, and to combine the explanation style with the response information and the action information to adjust the output mode of the display object; A sending unit for associating the display object, the shooting data, and the response information to generate data for display, and for sending the data for display to the portable information processing terminal; The learning unit is used to associate and record the user's past input history, emotional state, and response information, extract user preference information through statistical processing or machine learning, and update the prompt statements and explanation styles based on the user preference information.
[0618] (Note 2) According to the information processing system described in Appendix 1, the action information is stored in the storage device as action information groups that are classified into combinations of multiple movement categories and multiple emotional states respectively. The action information acquisition unit selects action information for application to the display object from the action information groups based on the index information obtained from the movement category, the emotional state and the scene information.
[0619] (Note 3) According to the information processing system described in Appendix 1, the learning unit uses user preference information and behavioral history information generated for each user to update the tone specification, explanatory quantity specification and content category specification contained in the prompt statement input to the generative artificial intelligence model, and outputs the response information obtained from the generative artificial intelligence model based on the updated prompt statement in different ways for different users.
Claims
1. An information processing system, characterized in that, include: processor, The processor is configured to: Provides users with a user interface for setting virtual characters to customize their preferences; When a user uses a portable information terminal to point at a sports competition venue or a video of a sports competition, the acquired video data is analyzed, and the actions of a virtual character are generated based on the video data. The virtual character references a sports action database to obtain action data corresponding to the current sports scenario and parses input from the user so that the virtual character can engage in question-and-answer interaction on sports-related content. The parsing includes generating prompt text to instruct the generative artificial intelligence model to generate a response. The system identifies the user's emotions, adjusts the virtual character's narration style based on the identified emotions, and uses a generative artificial intelligence model to generate corresponding responses based on the prompt text.
2. The information processing system according to claim 1, characterized in that, The processor is configured to include motion data for multiple different sports in the sports motion database, and to set an index for selecting the motion data corresponding to each sports.
3. The information processing system according to claim 1, characterized in that, The processor is configured to enable the virtual character to learn through interaction with the user, record the user's past behaviors and preferences, and generate sports commentary that matches the user's preferences based on the recorded behaviors and preferences using a learning algorithm.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A