Information processing system

CN122802748APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610327024.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-17
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]在现有技术中,针对艺人或网络红人的追随者服务主要依赖于其公开发布的图文或视频内容,用户只能被动观看既有素材,难以获得沉浸式、个性化的追体验感受

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802748A_ABST
    Figure CN122802748A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system comprises a processor configured to: input a prompt word indicating generation of visual information to a generated artificial intelligence model to generate an image or a video for reproducing the perspective of an artist or a network celebrity; identify the emotion of a user and adjust the generated content according to the identified emotion; and distribute the generated content to a terminal of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] In existing technologies, follower services for celebrities or online influencers primarily rely on their publicly released text, images, or videos. Users can only passively view existing materials, making it difficult to obtain an immersive and personalized following experience. Existing systems typically lack the following capabilities: First, they cannot automatically generate realistic first-person images or videos from the celebrity's or influencer's perspective, making it difficult for users to experience their daily life or scenarios from the target's point of view; second, they lack mechanisms for recognizing and adapting to users' emotional states, failing to dynamically adjust the style and emotion of generated content based on users' current emotional needs, thus hindering the continuous improvement of user engagement and satisfaction; third, they lack precise personalized product recommendation mechanisms based on users' historical behavior and interests, making it difficult for users to promptly discover following experience content that highly matches their interests; fourth, they fail to fully utilize user feedback to iteratively optimize subsequent content generation, resulting in generated content that is difficult to consistently align with users' true preferences, limiting the sense of unity and connection between users and celebrities or online influencers. Therefore, it is necessary to provide a content generation and distribution system that comprehensively considers the perspective of celebrities or online influencers, user emotions, user historical behavior, and user feedback to solve the above problems and enhance users' immersive experience, emotional resonance, and personalized satisfaction. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes an information processing system comprising a processor configured to input cue words to a generative artificial intelligence model, instructing the generation of visual information to produce images or videos reproducing the perspective of a celebrity or online influencer. By encoding the perspective features, scene features, and style features of the celebrity or online influencer into the cue words, the processor can drive the generative artificial intelligence model to output visual content that conforms to the perspective features of the target object, thereby providing users with a follow-up experience image or video based on the perspective of the celebrity or online influencer. Furthermore, the processor is configured to recognize the user's emotions and adjust the generated content according to the recognized emotions, including but not limited to adjusting the color tone, rhythm, emotional atmosphere, and content theme, so as to provide users with a content experience more suited to their psychological needs when they are in different emotional states. The processor is also configured to distribute the generated content to the user's terminal, enabling the user to conveniently view or browse the generated follow-up experience content through their own terminal device. Furthermore, the processor is configured to analyze users' past purchase records and interests to generate personalized product recommendations. This allows users to quickly discover content that highly matches their preferences based on their long-term or short-term interests, thereby enhancing their special connection to the target audience and promoting immersion in the target audience's world. In addition, the processor is configured to collect user feedback and adjust the generated content based on this feedback during the next content generation. This includes optimizing prompts, content structure, and emotional style to continuously provide users with an experience that better meets their expectations and deepens the sense of connection between users and celebrities or online influencers. Through the above configuration, this invention can automatically generate content from the perspective of celebrities or online influencers while comprehensively considering user emotions, historical behavior, and feedback to achieve a highly personalized, immersive, and iteratively optimizable content service.

[0005] "System" refers to the overall combination of software and hardware, including at least one processor and optional memory, network interface and terminal communication interface, for performing the functions described in this invention.

[0006] A “processor” refers to an electronic computing unit capable of executing program instructions to perform data processing, model reasoning, and control flow, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a neural network processing unit (NPU), or any combination thereof.

[0007] "Generative artificial intelligence models" refer to models trained based on machine learning or deep learning techniques, used to generate new images, videos, texts or other multimedia content based on input data, including but not limited to generative adversarial networks (GANs), diffusion models, variational autoencoders (VAEs), autoregressive generative models or combinations thereof.

[0008] "Cue words" refer to the set of input information provided by the processor to the generative artificial intelligence model to instruct the model what kind of visual information to generate. They are usually in the form of natural language text, structured descriptions, or embedded vectors, and are used to specify the generation conditions such as scene content, perspective, style, and emotion.

[0009] "Visual information" refers to content data that can be presented in the form of images or videos, including still image frames, dynamic image sequences, accompanying visual effects, and visual elements related to the perspective of celebrities or internet celebrities.

[0010] "Images or videos" refer to static or dynamic sequences of images generated by generative artificial intelligence models that can be displayed on a display device, including single-frame images, multi-frame animations, short videos, long videos, and any combination thereof.

[0011] "Celebrity or Internet celebrity perspective" refers to the viewing point of view and style characteristics that are presented from the perspective of the celebrity or internet celebrity, from their subjective first-person point of view or from a shooting angle that is consistent with their image characteristics. This includes, but is not limited to, camera height, composition habits, selection of typical scenes and color preferences.

[0012] "User's emotions" refers to the emotional state of a user at a specific point in time or within a certain time range, including but not limited to happiness, calmness, tension, sadness, excitement, relaxation, etc., which can be inferred from user behavior data, physiological signals, voice tone, facial expressions, or user input.

[0013] "Generated content" refers to images, videos, or related multimedia data that are generated by a generative artificial intelligence model under the constraints of prompt words, and are then selected, adjusted, or post-processed by a processor and are available for users to view or use.

[0014] "User terminal" refers to an electronic device used by a user to access the system and receive, play, or browse generated content, including but not limited to smartphones, tablets, personal computers, smart TVs, head-mounted displays, or other devices with display and network communication capabilities.

[0015] "Past purchase records" refers to the collection of historical information associated with a user's account regarding goods or services purchased within the system or on platforms linked to the system, including purchase time, product type, price, viewing or usage details, etc.

[0016] "Interest" refers to the user's preference tendencies in terms of content type, scenario, theme, artist, etc., obtained by analyzing user behavior data (including browsing, clicking, collecting, purchasing, rating, etc.) and optional user-initiated preference information.

[0017] "Personalized product recommendations" refers to a customized product list or recommendation result that is different from other users, which is output by the processor after filtering and sorting candidate products based on the user's past purchase history, interests and optional real-time emotional state.

[0018] "User feedback" refers to information provided by users, either explicitly or implicitly, regarding the quality, level of preference, or suggestions for improvement of the content after experiencing it. This includes, but is not limited to, data such as ratings, comments, likes, dislikes, dwell time, completion rate, complaints, or suggestions.

[0019] "Target object" refers to the artist or internet celebrity entity associated with the generated content, or the object that serves as the subject of perspective reproduction and experience in a specific application scenario, and is used to constitute the core role of the user's pursuit of experience. Attached Figure Description

[0020] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0021] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0022] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0023] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0024] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0025] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0026] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0027] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0028] Figure 9 This represents an emotion map that maps multiple emotions.

[0029] Figure 10 This represents an emotion map that maps multiple emotions.

[0030] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0031] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0032] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0033] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0034] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0035] First, let me explain the terminology used in the following instructions.

[0036] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0037] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0038] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0039] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0040] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0041] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0042] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0043] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0044] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0045] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0046] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0047] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0048] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0049] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0050] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0051] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0052] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0053] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0054] Traditional computer-based content generation and distribution technologies typically rely solely on simple user input or preset templates, directly invoking generative artificial intelligence models to generate images or videos. This type of technology suffers from the following technical problems: (1) When generating visual content, the server does not systematically extract the perspective features and expressive tendencies of the information provider from large-scale, multimodal information sources, resulting in a large deviation between the generated content and the real style of the information provider, making it difficult for users to obtain a stable and consistent immersive experience. In other words, the existing system lacks an effective data processing and feature representation mechanism in terms of "how to computationally model and reproduce the perspective of a specific subject".

[0055] (2) Existing generation processes are mostly static processes of "single call". The server usually only receives static instructions from the user, converts them into simple prompts and inputs them into the generative artificial intelligence model, without uniformly modeling and dynamically utilizing the user's emotional state, preference information, historical acquisition records, and interactive feedback at the system level, resulting in: - The generated prompts lack specificity and adaptability; - The server cannot adjust the generation conditions in real time based on the user's emotions and preferences; - User feedback cannot be effectively fed back into subsequent feature extraction and model-driven processes, making it difficult for the system to "become more and more user-friendly with use".

[0056] (3) On the server side, information collection, feature analysis, prompt statement construction, model invocation, content commercialization, and distribution processing are often loosely coupled multiple independent modules, lacking a unified data flow and control flow architecture: - There is a lack of structured connection between the information source capture and analysis results and the subsequent generated prompt statements; - The product generation and e-commerce or content distribution logic are separated, and the server cannot perform fine-grained control over product recommendations based on content characteristics and user behavior at the technical level; - The lack of a weighted relearning or reanalysis mechanism based on user feedback makes it impossible to dynamically update the feature information provided by the object's perspective, making it difficult for the overall intelligence level and generation quality of the system to evolve.

[0057] Therefore, an improved computer implementation method and system is needed to build an integrated data processing and control mechanism on the server side, from information source data collection, multimodal feature extraction, automatic generation of prompts, and generative artificial intelligence model invocation, to the composition, distribution, and feedback relearning of experiential products. This would address the shortcomings of existing technologies in terms of perspective modeling accuracy, user adaptability, and system self-evolution capabilities, thereby substantially improving the performance of computer technology in the field of personalized immersive content generation and delivery.

[0058] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0059] In this invention, the server includes: a device for automatically acquiring data containing text, image, or video information from an information source related to the information provider, and performing multimodal analysis on the data to extract feature information representing the perspective and expressive tendencies of the information provider; a device for automatically generating prompt statements for a generative artificial intelligence model based on the feature information and user-acquired demand information, and inputting the prompt statements and generation conditions into the generative artificial intelligence model to generate visual information mimicking the perspective of the information provider; and a device for acquiring the user's emotional state and preference information, and adjusting the prompt statements according to the emotional state and preference information. The device includes: an apparatus for adjusting the generation conditions to control the content or presentation of the generated visual information; an apparatus for constructing experiential goods from the generated visual information, generating and recording identification information, description information, and provision conditions of the experiential goods, receiving purchase requests from user terminals, performing settlement processing, granting access permissions, and distributing the visual information via streaming media or download; and an apparatus for obtaining evaluation information or feedback information from users about the experiential goods, and performing weighted relearning or reanalysis on the updating of the aforementioned feature information and the generation of prompt statements based on the evaluation information or feedback information, so as to dynamically update the feature information and improve subsequent generated content. This allows for the formation of a closed-loop computational process within the server, from information source acquisition, multimodal feature modeling, adaptive generation of prompt statements, invocation of generative artificial intelligence models to experiential goods distribution and feedback-driven relearning. This improves the accuracy and consistency of perspective reproduction at the computer technology level, enhances the adaptability of generated content to user emotions and preferences, and enables the system to continuously optimize generation quality and recommendation effects through user interaction, significantly improving the technical performance of computer-based personalized immersive content generation and delivery.

[0060] "Information provider" refers to the subject whose perspective and expression tendencies are modeled and imitated by the system in this invention. It can be an individual, a group, or other content provider.

[0061] "Information source" refers to a data provider that is related to the information provider and is used to obtain text, image, or video information, including websites, social networking services, blog platforms, and other online or offline data providers.

[0062] "Text information" refers to content that exists in the form of characters in an information source, including textual data that can be processed by natural language, such as articles, posts, comments, titles, tags, and descriptive descriptions.

[0063] "Image information" refers to visual data that exists in the form of static images in an information source, including photographs, illustrations, screenshots, and keyframe images extracted from videos.

[0064] "Video information" refers to multimedia data that exists in the form of a continuous time sequence of images in an information source, including various dynamic video files with or without audio tracks.

[0065] "Feature information" refers to the data representation extracted through the analysis of text, image, or video information, which is used to characterize the attributes of the information provider, such as perspective, content preferences, style characteristics, and performance tendencies.

[0066] "Perspective" refers to the combination of spatial and subjective observational characteristics, such as the observation position, direction, sense of distance, and focus of the subject, as reflected in the content expression of the information provider.

[0067] "Expressive tendencies" refer to the habitual expressive characteristics repeatedly exhibited by the information provider in content creation, such as thematic preferences, emotional atmosphere, composition style, color style, and narrative style.

[0068] "Generative artificial intelligence models" refer to computational models or combinations of models that process input prompts and generation conditions based on machine learning or deep learning techniques to generate new visual, audio, or text information.

[0069] "Prompt statements" refer to textual instructions input into a generative artificial intelligence model, which describe the scene elements, style requirements, emotional atmosphere, and other generative constraints of the target content to guide the generative artificial intelligence model to output results that meet expectations.

[0070] "Generation conditions" refer to the set of parameters used to control the output of generative artificial intelligence models, excluding prompt statements. These parameters include resolution, output type, number of steps, randomness parameters, style intensity, and other model control parameters.

[0071] "Visual information" refers to image or video data output by generative artificial intelligence models based on prompts and generation conditions, used to reproduce or imitate scenes from the perspective of the information provider.

[0072] "Experience products" refer to virtual goods that consist of visual information and related metadata, which are available to users for acquisition and consumption. This includes access permissions to the visual information, associated descriptions, and service content related to its display method.

[0073] "Identification information" refers to the information used to uniquely identify the product being experienced, including the product number, name, category, and identification data corresponding to the index in the backend database.

[0074] "Description information" refers to the textual descriptions associated with the product being experienced, including product introduction, scene description, style description, and usage instructions.

[0075] "Conditions for provision" refers to the technical and business constraints applicable when providing experiential products to users, including conditions such as price, access duration, number of views, available terminal types, and distribution methods.

[0076] "User terminal" refers to the computing device used by users to access servers and obtain and experience products, including portable terminals, fixed terminals and other information processing devices with network communication and content presentation capabilities.

[0077] "Emotional state" refers to the emotional characteristics of a user over a certain period of time, which the system infers based on the user's interactive behavior, input content, or external sensor data. These include emotional types such as pleasure, calmness, tension, and excitement, or combinations thereof.

[0078] "Preference information" refers to data that reflects users' long-term or short-term tendencies in terms of content type, style, theme, duration, and interaction methods, including explicitly set interest tags and interest characteristics inferred from behavioral analysis.

[0079] A "purchase request" refers to a request sent by a user to a server through a user terminal to apply for access to a certain product for trial. It usually includes product identification information, user identification information, and payment-related information.

[0080] "Access permission" refers to the right granted by the server to a user to access and use the visual information in a product under certain conditions, including the right to view, replay, download, or display it in a limited environment.

[0081] "Evaluation information" refers to the quantitative or qualitative evaluation data submitted by users after using and experiencing a product, including ratings, tag selections, and brief evaluation comments.

[0082] "Feedback information" refers to more detailed or open-ended reaction data provided by users during or after experiencing the product, including text comments, interaction records, and questionnaire answers.

[0083] "Relearning" refers to the process by which a server updates the model or parameter set used to generate feature information based on newly added text, image, or video information, combined with evaluation or feedback information.

[0084] "Reanalysis" refers to the process by which the server recalculates, reweights, or re-clusters the extracted features based on new data and evaluation or feedback information without completely retraining the model.

[0085] In this invention, the server operates as the core data processing node. The server can consist of one or more computing devices, each including at least one central processing unit (CPU), graphics processing unit (GPU), main memory, non-volatile memory, and a network interface. The server can run a general-purpose operating system, such as Linux, and deploy web server components (e.g., a combination of a reverse proxy server and an application server), application logic components, database components, and generative artificial intelligence model inference components on it. The terminal can be a smartphone, tablet, personal computing device, or head-mounted display device, running a web browser or native applications. Users interact with the server through the terminal.

[0086] In one implementation, the server uses a relational database management system (e.g., a relational model-based database) to store structured data, such as information provider identifiers, information source URLs, feature vectors, product metadata, and user account information. The server also uses object-oriented storage services (e.g., a storage system based on a distributed object storage protocol) to store large amounts of image and video files. The server communicates between multiple functional modules via message queues or internal APIs to achieve high-concurrency data processing.

[0087] When acquiring and parsing information sources, the server uses an HTTP client library to send requests to various information service platforms. In a preferred embodiment, the server uses an HTTP library such as Python (e.g., requests) or equivalent components in other languages ​​to access web page resources and open interfaces. For information sources in web page form, the server uses an HTML parsing library (e.g., BeautifulSoup or lxml) to extract the main text, image links, and any existing geographic and time tags. For information sources from social networking services, the server uses the service's application programming interface (API), combined with an authentication token, to batch retrieve posts, comments, image links, and video links related to the information provider. The server records each piece of raw data along with the source platform type, timestamp, geographic location identifier, and information provider identifier in the database.

[0088] When analyzing text information, the server uses natural language processing frameworks, such as pre-trained language models based on the Transformer architecture. In one implementation, the server loads a language model based on a multi-layer self-attention structure, which includes multiple encoding layers, each containing a multi-head self-attention sublayer and a feedforward network sublayer. The server segments the text information obtained from the information source into words and encodes them into a sequence of sub-word IDs, then inputs this sequence into the language model to obtain a context vector representation for each sub-word. The server uses structures such as linear classifiers and conditional random fields in the vector space to identify location entities, time entities, sentiment words, and topic-related words. By statistically analyzing the entities and descriptive words that are frequently mentioned for the same information in a specific city or scene, the server generates corresponding topic distribution vectors and sentiment distribution vectors. These vectors constitute part of the text-side feature information.

[0089] When analyzing image and video information, the server uses a visual neural network model. In one implementation, the server uses a model based on a convolutional neural network or a visual Transformer as the basic feature extractor. The server uniformly adjusts the images (or video keyframes) obtained from the information source to a predetermined resolution and performs normalization processing. The server inputs these images into the visual model to obtain a high-dimensional feature vector for each image. The server then uses a multi-label classifier on top of the feature vectors to determine scene categories (e.g., city streets, indoor venues, natural landscapes), the presence or absence of landmarks, crowd density, lighting conditions, etc. The server also analyzes pixel brightness histograms and color histograms to quantify the proportion of warm and cool colors and the overall brightness distribution to obtain style-related visual features.

[0090] The server uses a preferred approach, employing a multimodal alignment model, such as a jointly trained text encoder and image encoder, to map text descriptions and image content into a unified embedding space. The server concatenates topic and sentiment vectors from the text side with scene and style vectors from the visual side, and further compresses them using dimensionality reduction methods (e.g., Principal Component Analysis (PCA)) or multilayer perceptrons to obtain feature vectors representing the perspective and expressive tendencies of an information provider in a specific environment. The server stores this feature vector, along with the information provider's identifier and scene conditions (city name, time period, typical landmarks, etc.), in a database to form a "perspective feature table."

[0091] When generating the prompt, the server first receives the request information from the user's terminal. The user can enter free text in the terminal interface, such as "want to experience the perspective of travel information provider B in Paris," or specify parameters such as city, time period, content type (still image or video), and desired atmosphere through a selection interface. The terminal sends this information to the server via a secure channel. The server performs word segmentation and keyword extraction on the user input and matches it with records in the perspective feature table. If the server detects that the user has not specified a time period, it uses the time period (e.g., nighttime) that appears most frequently in the perspective features of the information provider in the target city as the default value; if the user has not specified a landmark, the server selects one or two typical landmarks based on the high-weight landmark labels (e.g., tower-like buildings, squares) corresponding to the feature vector.

[0092] After obtaining user requirements and perspective characteristics, the server uses a template generation system to construct prompt statements. In one implementation, the server pre-stores multiple language templates, such as: "Please generate an image of city Z in Y, viewed from the perspective of information provider X. The viewpoint is A, the overall color tone is B, the atmosphere is C, and the image contains elements such as D." The server replaces X with the information provider's identifier or type description, Y with the city name, Z with the specific scene (e.g., streets, near landmarks), A with the viewpoint height and upward / downward angle information, B with a description of warm or cool colors, C with a description of the emotional atmosphere, and D with a description of landmarks or crowd elements. Examples of such generated prompt statements include: "Please use a realistic style to generate an image of the Eiffel Tower at night in Paris as seen from the perspective of travel information provider B. The viewpoint should be close to pedestrian height, with the Eiffel Tower slightly tilted upwards. Use highly saturated warm-colored city lights in the image, with a few pedestrians on the street, creating a romantic and quiet atmosphere." In another example, the server generates the following prompt based on the user-specified city and mood: "Please generate an image of a New York street scene at night from the perspective of information provider A. The street is wet with rain, and the surface reflects neon lights. The perspective is from the sidewalk, looking slightly upward at the surrounding high-rise buildings. The overall style is realistic and the atmosphere is slightly lonely but full of urban charm." In configuring the generative AI model, the server uses a generative network based on a diffusion model or an autoregressive model. In one embodiment, the server uses a diffusion model containing a U-shaped network structure and time-step embeddings as the visual generation backend. The server uses large-scale image-text pairing data as training samples when internally training or fine-tuning the model. During training, the server maps text-side prompts to text feature vectors through a text encoder (e.g., a Transformer-based text embedding network), inputs noisy images into the U-shaped network, and introduces text features at each time step through conditional mechanisms (e.g., cross-attention or feature concatenation), allowing the network to learn to generate images semantically consistent with the text through progressive denoising. The server defines a loss function that can be the noise prediction error (e.g., mean squared error) and updates the network parameters using stochastic gradient descent or adaptive optimization algorithms. The server can also employ data augmentation strategies during training, such as random cropping, color perturbation, and horizontal flipping, to improve model robustness.

[0093] During the inference phase, the server inputs the prompts into the text encoder to obtain text feature vectors. Then, based on user settings or system default generation conditions (such as resolution, number of denoising steps, and random seed), it runs a diffusion process to generate images that conform to the specified viewpoint and style. When generating video content, the server can input prompts from multiple time segments into the model separately to generate a series of image frames. These frames are then combined into a video file based on frame interpolation algorithms or video encoding tools.

[0094] The server can infer user emotional state and preference information based on data uploaded from the user's terminal. Users can actively select their current emotion tag (e.g., "relaxed," "excited," "quiet") on their terminal, or their preferences can be indirectly reflected through historical behavioral data. In one embodiment, the server statistically analyzes the themes, colors, and scene types of previously viewed products to generate a user preference vector. Before generating prompts, the server weights and fuses this preference vector with a perspective feature vector, making the prompts more closely aligned with user preferences while maintaining the style of the information provided. For example, if a user prefers bright colors and scenes with many people, the server will increase the probability of descriptions such as "high crowd density" and "bright lighting" appearing when filling in the template.

[0095] In terms of the composition of experiential products, the server packages the generated visual information and its related descriptive information into a logical entity. The server generates a unique identifier for each experiential product and records metadata such as title, summary description, associated information (object identifier), city name, main scene elements, price, available access duration, and available terminal types. The server establishes an index in the database, enabling quick retrieval and recommendation of products based on information provided by object, city, theme, or sentiment tag. The server stores visual files in object storage and associates the access path with the experiential product identifier.

[0096] When a user purchases a trial product, the server receives a purchase request from the terminal. The terminal displays the product title, thumbnail, and price information in the user interface. The user confirms the purchase and selects a payment method through the terminal. After authenticating the purchase request, the server interacts with the payment processing system. Once the payment is successful, the server adds an access record for the product to the user's permissions table. When a user subsequently requests to view the product, the server generates a time-limited access link and returns it to the terminal via a secure communication protocol. Upon receiving the link, the terminal displays images or plays videos through its built-in player (such as a browser's multimedia element or a local multimedia component).

[0097] After acquiring user ratings and feedback, the server associates and stores this information with the corresponding product experience and the corresponding perspective feature vector. In one implementation, the server uses a weighted relearning scheme: it assigns different weights to samples based on their rating levels, treating high-rated samples as positive examples and low-rated samples as negative examples. When analyzing new information sources later, the server incorporates these weights to adjust the clustering or dimensionality reduction process, making the new feature vectors closer to the regions that received high ratings. When generating prompts, the server also considers this feedback, reducing the frequency of descriptive patterns that frequently lead to low ratings (such as extreme color tones or overcrowded scenes) in the template. This approach is not a simple manual rule change, but rather automatically adjusts the parameters for prompt generation by statistically analyzing the relationship between feedback data and the feature space distribution, thereby improving the success rate of content generation at the algorithmic level and reducing results that do not meet user expectations.

[0098] Through the aforementioned specific data structures and algorithmic processes, the server achieves a technological path that is distinctly different from the traditional manual editing of prompts and screening of information sources. Instead of requiring humans to read and filter information line by line, the server automatically constructs and updates perspective representations through multimodal feature modeling and weighted relearning mechanisms, and generates visual content in conjunction with a template system and a generative artificial intelligence model. Because the server internally manages information source data, feature vectors, prompt parameters, and user feedback in a unified manner, and stores and indexes them in a structured database, the entire system maintains high processing speed and scalability when handling large amounts of information and its multi-scenario data.

[0099] Through the above-described structure, the system of this invention not only provides users with experiential products at the business level, but more importantly, it achieves the following improvements at the computer technology level: The server, by decoupling feature vectors and template parameters, enables the generation of prompt statements to be quickly updated based on local changes in the feature space, without needing to redesign the overall business logic; the server, through GPU-oriented generative artificial intelligence model inference services, unifies the scheduling of complex image generation calculations, avoiding latency caused by insufficient computing power on the terminal side; the server, through multimodal alignment and weighted relearning, improves the accuracy of perspective modeling and reduces the deviation between the generated results and the realistic style of the information provided object; the server, through a unified data structure and indexing strategy, reduces retrieval and recommendation latency in large-scale information source scenarios, thereby achieving a dual improvement in processing speed and generation accuracy at the system level. These technical effects go beyond simply automating the human creative process; they optimize computer technology itself at the levels of algorithm design, data structure design, and system architecture design.

[0100] use Figure 11 The processing flow is explained.

[0101] Step 1: Users input their experience requirements on the terminal. Users enter text information related to their experience in the terminal's interface and select additional parameters.

[0102] Input: The text entered by the user in the terminal interface (e.g., "Want to experience the perspective of information provider B in Paris"), city selection, time period selection, content type (image / video), and mood or atmosphere preference.

[0103] The specific actions of the terminal include: - Combine the contents of text input boxes with the values ​​of drop-down options, checkboxes, and other controls into structured data; - Generate a local request ID for this request; - A predefined interface that sends requests containing user ID, request ID, text description, and parameters to the server via HTTPS.

[0104] Output: The experience request data packet sent to the server, and the request ID cached locally on the terminal.

[0105] Step 2: The server receives and parses the user request. The server receives request messages sent by the terminal from the network interface and parses out each field.

[0106] Input: An HTTP request data packet sent by the terminal, which contains fields such as user ID, request ID, request text, city, time period, content type, and sentiment preference.

[0107] The specific actions of the server include: - Parse HTTP headers and message bodies at the application layer, mapping JSON or form data into internal data structures; - Perform basic cleaning of the required text (remove HTML tags, extra spaces, etc.); - Validate the content type to ensure it is valid and check if required fields exist; - Create an experience request record in a relational database, recording the request ID, user ID, timestamp, and request status (e.g., "pending").

[0108] Data processing and computation: The server performs length checks and simple sensitive word filtering on text fields, and range checks on numeric fields, to generate a valid unified request object that can be used for subsequent processing.

[0109] Output: Standardized experience request records stored in the database, and internal request objects for subsequent module calls.

[0110] Step 3: The server identifies and collects information source data. The server provides an object identifier or name based on the information in the requested object, determines the relevant information source, and collects raw data.

[0111] Input: Internal request object (including information provider object identifier or name, city, etc.) and information source configuration records already saved in the database.

[0112] The specific actions of the server include: - Search the information source configuration table to see if there are blog URLs, social media account IDs, etc. associated with the information provider; - If no existing record is found, call an external search or platform API to query the relevant page or account based on the object name provided in the information; - Use an HTTP client library to access each URL sequentially to download HTML pages or call the platform API to get a list of posts; - Use an HTML parsing library to extract body text, titles, image links, video links, and time and tag information; - Write each data entry along with the information provider's ID, the source platform type, and the timestamp into the original data table.

[0113] Data processing and data computation: The server performs unified structuring processing on raw data from multiple platforms and in various formats, mapping unstructured HTML and API responses to standard fields (text fields, media URL fields, time fields, etc.).

[0114] Output: A list of raw text and media records written to the database, and a set of raw data IDs associated with this experience request.

[0115] Step 4: The server performs text information analysis to extract text features. The server reads the text information related to this request from the original data table and performs natural language processing.

[0116] Input: A collection of text records (including fields such as body text, tags, time, and geographic information) related to the target information provider and the target city.

[0117] The specific actions of the server include: - Use a word segmenter to divide the text into sequences of words or subwords; - Each text is encoded using a Transformer-based language model to obtain a sequence of context vectors; - Detect location entities, time expressions, activity verbs, etc. from vector sequences using the named entity recognition module; - Use the sentiment analysis module to calculate the sentiment distribution (such as positive, neutral, negative, and their intensity) of the entire text; - Use topic modeling algorithms to cluster the vectors of multiple texts to obtain common topics (such as "city night view", "coffee shop", "natural landscape").

[0118] Data processing and data computation: The server maps each text to a set of numerical features (location one-hot vector, time period vector, sentiment vector, topic distribution vector, etc.), and then performs statistical summation, averaging or weighting on all records to obtain the text-side feature vector of the information provider in a specific city.

[0119] Output: Text feature vectors stored in the feature information table, along with associated information such as object ID and scene conditions.

[0120] Step 5: The server performs image and video information analysis to extract visual features. The server performs visual feature analysis on the image and video keyframes related to this request.

[0121] Input: URLs of image files and video files related to the target information provider and the target city, as well as keyframe images extracted from the video.

[0122] The specific actions of the server include: - Download remote media to local cache or stream it directly; - Resize and normalize each image to fit the input size of the visual model; - Input the image into a convolutional neural network or a visual Transformer model to extract high-dimensional feature vectors; - Use a multi-label classification head to predict scene categories, landmark presence, and crowd density; - Use pixel-level statistical methods to calculate the brightness histogram and color histogram of the image to determine the dominant color tone and lighting conditions; - Use geometric analysis or additional networks to estimate the shooting angle (looking up, looking down, looking at eye level) and shooting distance (close-up, medium shot, long shot).

[0123] Data processing and data computation: The server maps each image or keyframe to a set of numerical features (scene category probability vector, hue index, light intensity value, viewing angle, etc.), and aggregates the features of multiple images of the same information provider in the target city in a weighted manner to form a visual feature vector.

[0124] Output: Visual feature vectors stored in the feature information table, and the corresponding set of media record IDs.

[0125] Step 6: The server fuses multimodal features and generates viewpoint feature vectors. The server fuses textual and visual features to obtain comprehensive features that represent the perspective and expressive tendencies of the information provided.

[0126] Input: The text feature vector obtained in step 4 and the visual feature vector obtained in step 5 are both related to the same information provider and scene conditions.

[0127] The specific actions of the server include: - Align the text feature vectors with the visual feature vectors in the same dimension (e.g., by linear transformation or multilayer perceptron). - By concatenating or weighting the two vectors, multimodal joint features can be obtained; - Use principal component analysis or dimensionality reduction networks to compress joint features to remove redundant dimensions and highlight key differences; - Normalize the result vector to facilitate subsequent similarity calculation and retrieval.

[0128] Data processing and computation: The server performs vector weighting, matrix transformation and feature dimensionality reduction operations in the numerical vector space to compress a large amount of raw text and image / video data into a representative feature vector.

[0129] Output: View feature vectors recorded in the view feature table, and their association with information provider ID, city, time period, and other conditions.

[0130] Step 7: The server generates initial prompts based on user needs and perspective characteristics. The server uses the aforementioned perspective features and user needs information to construct the initial prompt text.

[0131] Input: Viewpoint feature vector, user experience request object (including city, time period, content type, atmosphere preference, etc.), and predefined prompt statement template set.

[0132] The specific actions of the server include: - Convert the scene labels, color labels, and emotion labels with higher weights in the viewpoint feature vector into language descriptions (such as "night", "warm-colored lights", "romantic and quiet", etc.). - If the user does not specify a parameter (such as a time period), the default value will be automatically selected based on the maximum weight of the corresponding category in the feature vector; - Select a template from the prompt template set that matches the content type (image or video); - Fill in the template placeholders with the identifier description of the information provider, city name, scene elements, viewpoint, color tone and atmosphere, etc., and generate natural language prompts.

[0133] Data processing and data computation: The server internally uses rule mapping and placeholder replacement to map numerical features into phrases, and then combines them into complete sentences through string concatenation.

[0134] Output: One or more initial prompt text statements, such as: "Please generate a realistic image of the Eiffel Tower at night from the perspective of travel information provider B. The viewpoint is close to pedestrian height, slightly looking up at the Eiffel Tower. The image uses highly saturated warm-colored city lights, with a few pedestrians on the street, creating a romantic and quiet atmosphere." Step 8: The server adjusts the prompts and generation conditions based on user emotions and preferences. The server fine-tunes the initial prompts and generation conditions based on the user's emotional state and preferences.

[0135] Input: initial prompt statement, user emotional state (e.g., emotional label inferred from user selection or behavior), user preference vector (obtained from historical viewing records and rating statistics).

[0136] The specific actions of the server include: - Calculate the similarity between the user preference vector and the current view feature vector, and use a numerical value to measure the degree of style matching between the two; - If the similarity is low, certain descriptive components may be selectively added or removed (e.g., increasing or decreasing crowd density, changing light intensity, etc.). - Adjust the atmosphere vocabulary according to the user's emotional state, such as adding descriptions like "quiet" and "gentle" when the user is in a "wanting to relax" state; - Configure generation conditions according to preferences, such as increasing resolution, adjusting randomness parameters, or increasing the number of generation steps, to meet high-quality requirements.

[0137] Data processing and data operation: The server incorporates user preferences into the construction of prompt statements through vector similarity calculation and rule mapping, and replaces or weights the keywords in the prompt statements.

[0138] Output: The final, adjusted prompt and a set of associated generation conditions (such as resolution, number of steps, style intensity parameters, etc.).

[0139] Step 9: The server calls a generative artificial intelligence model to generate visual information. The server sends a generation request to the generative artificial intelligence model inference service and obtains visual output.

[0140] Input: Final prompt statement, generation conditions (including image size, number of generation steps, random seed, output type, etc.), and a generative artificial intelligence model instance deployed on a GPU.

[0141] The specific actions of the server include: - Encode the prompt statement into a text embedding vector (via a text encoder network). - Initialize the random noise tensor according to the generation conditions; - Denoising is performed on the text embedding vector under conditional constraints using a diffusion model or autoregressive model in several iterations. - Perform operations such as matrix multiplication, convolution, and attention calculations on the GPU to gradually obtain a clear image or a sequence of consecutive frames; - If the output is video, the multi-frame results are encoded and combined into a video file.

[0142] Data processing and computation: The server performs large-scale floating-point operations within the model, including forward propagation graph calculation, activation function calculation, and tensor updates, mapping text conditions and random noise into pixel matrices.

[0143] Output: One or more generated image files, or a generated video file, and the corresponding file's storage path or URL.

[0144] Step 10: The server constitutes the experience product and records metadata. The server packages the generated visual information into physical experiential goods that can be purchased.

[0145] Input: URL of the generated visual file, prompt statement, viewpoint feature vector, and user request parameters.

[0146] The specific actions of the server include: - Assign a unique product identifier to each generated content; - Automatically generate product titles and descriptions based on prompts and feature vectors, such as "Night view experience of the Eiffel Tower in Paris from the perspective of information provider B"; - Set conditions such as price, access duration, and allowed number of views; - Insert records into the product table that include product identifier, title, description, file URL, price, and availability conditions.

[0147] Data processing and data computation: The server aggregates various information related to the generated content into a structured record and creates an index to support fast retrieval and recommendation.

[0148] Output: The completed experience product records in the database, and a list of product entries available for subsequent display and purchase.

[0149] Step 11: Users can view and purchase trial products through the terminal. Users browse the generated product information on the terminal and submit a purchase request.

[0150] Input: A list of products or individual product details (including title, description, thumbnail, and price) sent by the server, as well as user clicks and selections on the terminal interface.

[0151] The specific actions of the terminal include: - Display product titles, thumbnails, and descriptions on the screen; - Collect the selected product ID and payment method when the user clicks to purchase; - Send the purchase request to the server's order interface via HTTPS.

[0152] Output: Purchase request data sent from the terminal to the server (including user ID, product ID, and payment parameters).

[0153] Step 12: The server processes the purchase request and issues an access link. The server processes the order, verifies the payment, and returns access information.

[0154] Input: Purchase request sent by the terminal, user account information, product records, and payment result returned by the payment gateway.

[0155] The specific actions of the server include: - Create an order record and associate it with the user ID and product ID; - Send a payment request to the payment system and wait for a callback indicating whether the payment was successful or failed; - Upon successful payment, update the order status to "Paid" and add the user's access record for the product to the permissions table; - Generate a time-limited media access URL for this user; - Returns the access URL and related information to the terminal.

[0156] Data processing and computation: The server performs insert and update operations on the order table and permission table, and signs or encrypts the URL to implement access control.

[0157] Output: The access link and confirmation information returned to the terminal, as well as the order and permission records stored on the server.

[0158] Step 13: The terminal acquires and presents the generated visual information. The terminal uses the access link obtained from the server to retrieve and display content.

[0159] Input: The media access URL returned by the server, and the user's display device parameters (screen resolution, bandwidth, etc.).

[0160] The specific actions of the terminal include: - Retrieve image or video data from a server or object storage via HTTP or streaming protocols; - Decode the received media data and render it to fit the terminal screen resolution; - Provides users with playback controls (start, pause, full screen, volume adjustment, etc.).

[0161] Data processing and computation: The terminal performs decoding and rendering operations locally, converting compressed video or image data into pixel signals that can be directly displayed by the display device.

[0162] Output: Images or videos displayed on the user's screen, providing visual content from the user's perspective, representing the actual user experience.

[0163] Step 14: Users submit ratings and feedback; the server performs feature updates. After watching, users submit ratings and feedback through the terminal, and the server updates the feature information accordingly.

[0164] Input: ratings, tags, and text reviews filled in or selected by the user on the terminal, as well as view feature vectors and product records currently stored on the server.

[0165] Specific actions taken by users and terminals include: - Users can select a rating star level in the terminal interface, check description tags (such as "very close to the original style" or "I like the colors"), and fill in their comments. The terminal packages these evaluations and feedback into a data packet and sends it to the server.

[0166] The specific actions of the server include: - Link and store the evaluation data with the corresponding experience products and perspective feature vectors; - Assign greater weight to high-scoring records and less or negative weight to low-scoring records; - In subsequent batch processing or online updates, these weights are used to adjust the cluster centers or dimensionality reduction mappings of the feature vectors, causing the feature space to shift toward the high-evaluation regions.

[0167] Data processing and computation: The server performs weighted averaging, re-clustering, or retraining of low-dimensional mapping functions on historical feature vectors and evaluation weights to update feature information.

[0168] Output: An updated set of viewpoint feature vectors, and parameters reflecting the impact of user feedback on the feature space, used to improve the generation of subsequent prompts and visual content.

[0169] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0170] With the widespread application of generative artificial intelligence models in image and video generation, related systems are now capable of generating still images or short videos based on simple text instructions. However, existing technologies still have significant shortcomings in several aspects, thus limiting the effectiveness of computer-generated content in immersive experiences and personalized services: First, existing systems typically generate visual content based solely on one-time input text commands, driving generative AI models to do so. This lacks the ability to structure and model based on the target object's real-world behavioral trajectories and multi-source information. In other words, computer systems often cannot automatically extract and integrate multi-dimensional features such as location, time, behavior, atmosphere, and emotion from scattered online document and image information. This results in a significant discrepancy between the generated visual content and the target object's real-world experience, leading to limited immersion.

[0171] Second, most existing generation systems operate on a "one-way batch processing" model, meaning that the user inputs instructions and the system outputs results all at once, lacking real-time interaction and dynamic linkage mechanisms with the user's terminal. Traditional solutions typically do not utilize location, gaze, head posture, and operational information from the user's terminal to update prompts in real time during the generation process, nor do they dynamically adjust generation conditions during the inference process of the generative AI model. Consequently, it is difficult to generate first-person perspective content that changes synchronously with the user's actions, resulting in low granularity and timeliness of the computer's response to user interaction.

[0172] Third, existing personalized recommendation or product suggestion systems are typically separated from content generation systems, merely providing list-based recommendations based on user history without deep coupling with "immersive content generation from a user's perspective." Computer systems lack a mechanism to integrate the target object's information source analysis results, user history and preference information, and user evaluations and reactions during the experience into a single generation process, thereby driving the generation of personalized experience and suggestion information simultaneously with content creation. This prevents servers from achieving a synergistic improvement in content generation quality and recommendation accuracy within the same technical architecture.

[0173] Fourth, at the content distribution level, existing systems generally transmit generated content as ordinary video files without dedicated post-processing and streaming optimization for the characteristics of generative AI model outputs. In other words, servers typically do not perform integrated processing such as frame generation, frame interpolation, encoding, and streaming format conversion on the generated visual information, making it difficult to guarantee a stable high frame rate and low latency immersive experience even with bandwidth or terminal performance limitations. The overall performance of the computer in the end-to-end content generation and distribution path is thus underutilized.

[0174] Therefore, it is necessary to provide a new computer implementation method that improves the computer's ability to generate immersive first-person perspective content and provide personalized experiences from both the system architecture and data processing flow levels by introducing structured feature extraction of information sources, automatic generation and dynamic updating of prompts, interactive control of generative artificial intelligence models, and low-latency streaming linkage with user terminals on the server side. This addresses the shortcomings of existing technologies in terms of real-time performance, immersion, personalization, and system integration.

[0175] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0176] In this invention, the server includes means for acquiring chronologically ordered document information and image information from information sources related to an object, extracting features related to location, time, behavior, atmosphere, and emotion from the document information and image information, and recording them as structured data; means for retrieving the structured data based on user request conditions, and automatically generating prompt statements using the features contained in the search results to instruct a generative artificial intelligence model to generate first-person perspective visual information; means for executing the generative artificial intelligence model using the prompt statements as input to generate image information or video information as a simulation of a character's perspective; and means for performing frame generation processing, frame interpolation processing, encoding processing, and distribution format conversion processing on the visual information. The apparatus includes post-processing to generate distribution data for streaming; an apparatus for sequentially sending the distribution data to a user terminal via a communication network and presenting the visual information in a first-person immersive experience via a display device or head-mounted display device connected to the user terminal; an apparatus for dynamically changing the generation conditions of the prompt statement or the generative artificial intelligence model and updating the generation or distribution of the visual information based on location information, gaze information, head posture information, and operation information obtained from the user terminal; and an apparatus for parsing user-related historical information and preference information and combining it with user evaluation information and reaction information to make user-level adjustments to the content of the prompt statement and the visual information to generate personalized experience information and suggestion information. This enables an integrated processing flow within the same computer system, encompassing multi-source information parsing, structured modeling, automatic generation and dynamic updating of prompts, interactive control of generative artificial intelligence models, and streaming optimization. This allows the server to generate and distribute continuously changing first-person immersive content based on the user's real-time actions and preferences, while simultaneously constructing personalized experience and suggestion information associated with the content during the generation process. This substantially improves the computer's technical performance in terms of immersive content generation quality, real-time interactive response capabilities, and personalized service accuracy.

[0177] "Information source" refers to the medium that provides data related to the object, including but not limited to websites, social networking services, blog platforms, database systems, etc., used to output documentary information and image information.

[0178] "Document information" refers to text-based data content, including but not limited to the main text of an article, titles, comments, tags, explanatory text, descriptions of time and location, and textual records related to the behavior of the subject.

[0179] "Image information" refers to data content that is primarily visual, including but not limited to still images, moving images, video frames, thumbnails, and other digital representations that can depict the appearance of scenes, people, or objects.

[0180] "Features" refer to structured data elements extracted from documentary and image information to describe attributes such as location, time, behavior, atmosphere, and emotion. These include labels, numerical values, category markers, or embedding vectors.

[0181] "Structured data" refers to a set of characteristic quantities organized and stored according to a predetermined data structure, including data in tabular, record, or key-value pair form, which facilitates computer retrieval, analysis, and generation processing.

[0182] "Request conditions" refer to the information that users input or select through their user terminals to specify the content generation requirements, including object identifiers, time ranges, location ranges, scene types, perspective types, and other constraint parameters.

[0183] "Prompt statements" refer to text instructions used to instruct generative artificial intelligence models to generate specific visual content. They are based on feature quantities and request conditions, describing generation requirements such as target scene, perspective, emotion, and style.

[0184] "Generative artificial intelligence models" refer to artificial intelligence models trained through machine learning that can automatically generate image or video information based on input prompts, including but not limited to deep learning-based text-to-image and text-to-video models.

[0185] "First-person perspective" refers to a visual viewpoint originating from the eyes of a virtual observer or character, which simulates what the character actually sees in terms of spatial position and direction.

[0186] "Visual information" refers to image or video information used to present a scene to a user, especially digital image data generated by generative artificial intelligence models that reflects the perspective of a person or the content of a scene.

[0187] "Frame generation processing" refers to frame-level generation or supplementary processing of visual information, including generating intermediate frames, extending frame sequences, or constructing new time-series frames to meet the requirements of video continuity.

[0188] "Frame interpolation" refers to the process of generating new intermediate frames between adjacent original frames to improve frame rate and smoothness of the video.

[0189] "Encoding processing" refers to the process of converting visual information into a compressed encoding format, including processing that uses predetermined video or image encoding standards to reduce data volume and adapt it for network transmission or storage.

[0190] "Distribution format conversion processing" refers to the process of converting encoded visual information into a container format suitable for streaming or adapted to different terminal playback environments, including slicing, packaging, and generating index files.

[0191] "Distribution data" refers to a collection of visual information data that has been encoded and converted into a distribution format and is suitable for streaming transmission over a communication network, including media clip files and their descriptive information.

[0192] "Streaming" refers to a transmission method that continuously sends and plays data to and from a user's terminal while the data is still being received or generated, allowing the user to start watching without waiting for the entire content to download.

[0193] "User terminal" refers to an electronic device used by a user to receive distributed data and perform interactive operations, including but not limited to smartphones, tablet computing devices, personal computing devices and wearable computing devices.

[0194] "Display device" refers to an output device used to present visual information on a user terminal, including integrated displays, external displays and projection devices.

[0195] "Head-mounted display devices" refer to devices worn on a user's head and displaying visual information at a position close to the eyes, including virtual reality display devices, augmented reality display devices, and mixed reality display devices.

[0196] "Operation information" refers to interaction-related data input by the user through the user terminal, including touch operations, key input, gesture commands, controller operations, and menu selections.

[0197] "Attitude information" refers to data related to the spatial position and orientation of the user or user terminal, including head posture, device posture, orientation angle, acceleration information, and gyroscope output.

[0198] "Location information" refers to data that represents the geographical location of a user terminal or user in physical space, including global satellite positioning information, network positioning information, and indoor positioning information.

[0199] "Gaze information" refers to data that indicates the user's gaze direction or gaze point location, including eye movement parameters or gaze coordinates obtained through eye-tracking devices.

[0200] "Generation conditions" refers to the set of parameters used to control the generation behavior of generative artificial intelligence models, including prompts, style settings, resolution, duration, random seed, temperature parameters, and other model control parameters.

[0201] "Historical information" refers to data related to a user's past behavior and usage records, including past viewing records, browsing records, interaction records, shopping records, and preference settings.

[0202] "Preference information" refers to parameters or characteristics that indicate a user's preferences, including preferences for specific objects, scene types, styles, durations, and interaction methods.

[0203] "Experience information" refers to content description data related to the user's immersive viewing or interactive experience, including personalized scene descriptions, guidance information, emotional enhancement information, and experience path suggestions.

[0204] "Proposal information" refers to suggested data generated based on visual information and user-related data, used to recommend content related to the experience object to users, including product information, service information or follow-up experience plans associated with the experience object.

[0205] "Evaluation information" refers to subjective evaluation data given by users on the generated content or experience process, including ratings, comments, tag selections, and satisfaction feedback.

[0206] "Reaction information" refers to behavioral and physiological response data collected automatically or semi-automatically during the user experience process, including dwell time, viewing order, interaction frequency, eye trajectory, posture changes, and other objective indicators used to reflect user responses.

[0207] The embodiments of this invention will explain how to specifically implement the system functions defined in each appendix, combining hardware structure, software modules, data structures, and the internal processing of generative artificial intelligence models. In the following embodiments, the subject is limited to one of server, terminal, or user.

[0208] I. Overall System Composition The server uses a computing device with a graphics processing unit (GPU) as its hardware foundation, such as a general-purpose computer with multiple GPUs. The server can use an architecture with a multi-core CPU and GPU, where the GPU can be a general-purpose parallel computing accelerator. The server runs an operating system (such as a general-purpose Unix-like system) and runs the following software components on it: The server runs web server software (such as a general HTTP server), application server frameworks (such as web frameworks based on interpreted languages), database management systems (such as relational databases or document databases), and inference services for deploying generative artificial intelligence models (such as model services based on deep learning frameworks).

[0209] A terminal is a computing device with communication and graphics display capabilities, such as a smartphone, tablet, personal computing device, or wearable computing device. The terminal runs a mobile or desktop operating system and has client applications or a web browser installed to exchange data with the server via secure communication protocols such as HTTPS or WebSocket.

[0210] Users select scenes, input parameters, and interact through the terminal, and view first-person perspective visual information generated by the server through the display device or head-mounted display device connected to the terminal.

[0211] II. Server-side functional modules and data structures 1. Information Source Acquisition and Feature Extraction Module The server includes an information acquisition module, a text parsing module, an image parsing module, and a feature extraction module.

[0212] The server uses an information retrieval module to access information sources, which may include websites, social networking service interfaces, blog platform interfaces, or other data service interfaces. The server uses an HTTP client library to send requests to retrieve literature information (such as article content, titles, descriptive text, commentary text, etc.) and image information (such as still image files, video frame screenshots, etc.) from the information sources.

[0213] After storing the acquired raw data in a temporary storage area, the server calls the text parsing module to perform regularization processing on the document information, including character encoding standardization, removal of control characters, removal of HTML tags, sentence segmentation, word segmentation, and stop word filtering. The server uses natural language processing algorithms (such as sequence labeling models based on recurrent neural networks or Transformer structures) to extract location entities, time expressions, action verbs, and sentiment terms from the text. The server can also employ named entity recognition models to map the text into a sequence of entity labels, generating feature quantities such as location labels and time labels.

[0214] The server uses an image parsing module to process image information. This module can include a scene classification model and an object detection model based on a convolutional neural network. After normalizing and resizing the image, the server inputs it into the image classification model to identify scene categories (e.g., outdoor street scenes, indoor cafes, stage environments, etc.), time features (day / night), and compositional features related to people. The server can further use a feature embedding model (such as a visual embedding network) to extract image feature vectors, which can then be used as style and scene constraints when generating prompts.

[0215] The server stores the features obtained from text and image analysis into structured data storage. The server can use relational tables or key-value pairs to store fields such as: object identifier, timestamp, location code, behavior type, atmosphere tag, sentiment tag, image feature vector, and text summary. This structured data representation allows the server to perform efficient indexing and filtering during subsequent feature retrieval and combination, thereby improving query speed and the accuracy of generated conditions.

[0216] This structured modeling eliminates the need for the server to rescan all original document and image information each time it generates data. Instead, it combines the data directly at the feature level, which significantly reduces the response time of generation requests and decreases the frequency of access to external information source interfaces, thus reducing communication load.

[0217] 2. Prompt Statement Generation Module The server includes a prompt generation module, which takes the request conditions and structured data entered by the user through the terminal as input and outputs prompts adapted to the generative artificial intelligence model.

[0218] The server first receives the request conditions sent by the terminal. These conditions may include parameters such as target object identifier, desired location, time range, scene type, desired emotional atmosphere, and whether first-person perspective video is required. The server then performs a retrieval operation in the structured data storage based on these parameters, using an index to perform a joint query on fields such as location, time, behavior, and emotion to obtain a set of records that match or are similar to the request conditions.

[0219] After obtaining the candidate record set, the server uses statistical or weighted algorithms to determine the combination of features that best represents the user's needs. For example, the server can weight and sort location, time, and behavioral features according to the user's recent viewing preferences or historical feedback, selecting the features with the highest weights as core features. The server can then use language templates or conditional text generation models to transform these features into natural language prompts.

[0220] For example, the server can generate the following Chinese prompt: "Please generate a scene of sunset along the Seine in Paris from the first-person perspective of internet celebrity A: the river is flowing slowly in the foreground, there are pedestrians and outdoor seating at cafes on the riverbank, and the Eiffel Tower can be seen in the distance. The overall atmosphere is romantic and relaxing, and the scene is realistic and detailed." The server can also be generated based on another scenario: "Please generate a first-person perspective view of the Tokyo concert backstage rehearsal scene from artist B: The scene shows the back side of the stage, lighting equipment, staff, and musicians adjusting their instruments. The overall atmosphere is slightly tense but full of anticipation." When generating prompts, the server can use a text generation sub-model based on the Transformer architecture. This sub-model takes structured feature vectors as input and generates coherent natural language descriptions through a multi-head attention mechanism. The server can perform rule checks on the generated prompts to ensure that they contain key components such as "first-person perspective," "location description," and "time or atmosphere," so that the generative AI model can produce visuals that meet expectations.

[0221] By generating prompts through structured features, the server avoids relying entirely on users to manually write long text instructions, reducing user errors and inadequate expression. It also automates and standardizes the construction of prompts, thereby improving the consistency and controllability of the generated content.

[0222] 3. Generative Artificial Intelligence Model Inference Module The server includes a generative AI model inference module, used to generate first-person perspective visual information based on prompts. The generative AI model can employ a diffusion model architecture or other image / video generation architecture.

[0223] When using the diffusion model, the server can employ a network structure consisting of a text encoder and an image decoder. The server inputs the prompt into the text encoder, which uses multiple layers of Transformers to encode the text, generating a fixed-dimensional text embedding vector. The server then uses this text embedding as a condition in the conditional network of the diffusion model, guiding the noise to gradually evolve into an image that matches the prompt.

[0224] The server performs the following typical operations during the image generation process: The server initially generates a random noise image that conforms to a Gaussian distribution. In each diffusion inversion step, the server estimates the noise residual using a conditional network and updates the image state according to the diffusion equation. After several iterations, the server obtains a clear image that is approximately noise-free. The server can control the diversity and stability of the images by setting different seed values ​​and sampling scheduling strategies.

[0225] When video needs to be generated, the server can use a video diffusion model or generate consecutive frames sequentially and perform time interpolation using a frame interpolation network. The server can call the same prompt statement multiple times to generate frames at different time points, or introduce time change descriptions in the prompt statement and control the generation of a sequence of images that change over time through conditional sequences.

[0226] During model inference, the server uses the graphics processing unit (GPU) to perform high-load operations such as matrix multiplication, convolution, and attention calculations, and leverages techniques like batch processing and tensor parallelism to improve throughput. By centrally deploying generative AI models on the server side, the terminal does not need to bear large-scale deep learning computations, thus achieving centralized computational load and efficient utilization of computing resources.

[0227] 4. Visual Information Post-processing and Streaming Distribution Module The server includes a post-processing module and a streaming distribution module. After generating raw image frames or frame sequences, the server uses the post-processing module to perform multi-stage processing on the visual information.

[0228] The server can supplement the frame sequence using frame generation algorithms to meet the target frame rate requirements. The server can use algorithms based on optical flow estimation or interpolation neural networks to generate intermediate frames, increasing the video frame rate from 12 frames per second to 24 frames per second or higher. The server then calls encoding components (such as an encoder based on a general multimedia processing library) to compress the video into a standard encoding format, setting parameters such as bitrate, resolution, and keyframe interval to reduce the bitrate while maintaining image quality.

[0229] The server uses a distribution format conversion module to slice the encoded video into media segments suitable for streaming and generates description files (such as playlists). The server generates media streams with multiple bitrate levels based on network conditions and terminal capabilities, allowing the terminal to adaptively select the appropriate stream based on bandwidth availability, thereby reducing playback stuttering and bandwidth consumption.

[0230] Through the post-processing and streaming distribution described above, the server transforms the static or low-frame-rate images output by the generative artificial intelligence model into a stable, high-frame-rate video stream that can adapt to the network environment. This achieves technical optimization from the model output domain to the network transmission domain, significantly improving the viewing experience on the user side.

[0231] 5. User terminal presentation and interaction module The terminal includes a user interface module, a playback module, and a sensor acquisition module. The user interface module presents options such as an object list, scene type, and time range, allowing users to select the desired experience via touch or commands. The terminal then packages the selection results into request conditions and sends them to the server.

[0232] After receiving the streaming media address or media segment returned by the server, the terminal uses the playback module to buffer, decode, and display it. The terminal can use a hardware-accelerated decoding unit to decode compressed video and output it to a local display screen or to a head-mounted display device via an interface.

[0233] The terminal acquires location, posture, and gaze information from an accelerometer, gyroscope, geolocation device, and optional eye-tracking device via a sensor acquisition module. The terminal sends this information to the server at a certain frequency, enabling the server to adjust prompts or generation conditions based on the user's head movements, gaze direction, and positional shifts. For example, when a user is watching a scene of "a walk along the Seine at dusk in Paris," if the user turns their head to look at a riverside café, the terminal can upload the change in gaze direction. The server can then update the description of the scene's focus in the prompts accordingly, resulting in subsequent images that better depict the café's interior or outdoor seating area.

[0234] Through this real-time interaction between the terminal and the server, the generated content is no longer a one-time static result, but a continuous stream that can be dynamically adjusted according to the user's actions, thereby achieving closed-loop control of introducing user sensor data into the condition space of the generative artificial intelligence model.

[0235] III. Examples of Learning Methods and Internal Model Processing Before system deployment, the server can train generative AI models and auxiliary models using a large amount of historical data. The server uses a training set to construct text-image pairing data, where the text includes descriptions of location, time, behavior, atmosphere, and sentiment, and the images are corresponding scene images. The server uses contrastive learning or conditional generation training to enable the model to learn the correspondence between text features and image features.

[0236] During training, the server defines a loss function, which can include reconstruction loss, contrastive loss, and perceptual loss. The server uses backpropagation to update the gradients of the weights in each layer of the neural network based on the loss function. The server can employ adaptive learning rate optimization algorithms to improve training convergence speed. Through this training method, the model can generate high-fidelity images based on prompts during the inference phase, and automatically adjust the composition and viewpoint when the prompts mention "first-person perspective" or a specific location.

[0237] The server can also train a sequence labeling model for feature extraction. This model takes literature information as input and outputs a label for each word (such as location, time, behavior, emotion, etc.). The server uses cross-entropy loss to train the model, enabling it to achieve high recognition accuracy on multi-source text. Since the feature extraction results directly determine the construction and generation conditions of the prompt statement, the improvement of model accuracy will directly lead to the improvement of the semantic accuracy of the generated content.

[0238] IV. Multiple Implementation Methods and Alternative Structures The server can employ different generative artificial intelligence model structures in different implementations. For example, in one implementation, the server uses an image generator based on a diffusion model and a separate video frame interpolator; in another implementation, the server uses a dedicated video generation model to generate a complete video sequence at once based on time conditions.

[0239] The server can generate prompts using either template rules or conditional text generation models. Template rules require less computation and are suitable for resource-constrained environments; conditional text generation can produce more natural and nuanced descriptions, making it suitable for scenarios with high content quality requirements.

[0240] In terms of presentation, the terminal can use only a flat-panel display device or combine it with a head-mounted display device to achieve an immersive experience. In the implementation using only a flat-panel display device, the first-person perspective is mainly reflected in the composition of the image and the movement trajectory; in the implementation using a head-mounted display device, the terminal can map the user's head posture to the direction of the virtual camera, and the server can adjust the generation conditions or select different perspective frames based on the posture information, thereby achieving a stronger sense of presence.

[0241] V. Technical Effects and Causal Relationship The server achieves deep parameterization of the generation process through the aforementioned structured feature extraction, automatic prompt generation, and dynamic condition control. Instead of simply passing user-input text directly to the generative AI model, the server uses feature layer abstraction and rule / model-driven prompt construction to uniformly map user intent, historical preferences, and object behavior data into the prompt space. This approach technically yields the following effects: The server is capable of high-speed querying and combining structured features, thus maintaining a low response time even in the context of large-scale data. The server can fine-tune the prompts or generation conditions based on the posture and operation information uploaded by the user in real time, so that the content generated by the generative artificial intelligence model can continuously respond to the user's actions, reducing the inconsistency and delay caused by static generation. The server reduces the amount of data that is repeatedly encoded and transmitted invalidally through post-processing and adaptive streaming, thereby reducing network load while ensuring image quality. The server improves system integration through a unified information source parsing, model inference, and streaming distribution process, reducing the overhead of format conversion and manual configuration between different systems.

[0242] Since these processes are all implemented within the computer using specific data structures and algorithmic flows, and directly affect the temporal, spatial, and network transmission characteristics of the generated content, this invention is not merely a simple automation of the human creative process, but rather a systematic improvement to the computer generation and transmission pipeline, thereby achieving technical performance improvements in multiple dimensions such as generation accuracy, processing speed, and communication efficiency.

[0243] use Figure 12 The processing flow is explained.

[0244] Step 1: The server obtains the raw data from the information source.

[0245] The server takes a pre-configured object identifier, information source URL, or API parameters as input and accesses websites, social network service interfaces, or blog platform interfaces via HTTP requests to obtain literature information (text data) and image information (images or video frames) related to the target object. The server performs basic validation on the response data (status code check, data integrity check) and temporarily stores the raw JSON, HTML, and image files in the server's local storage as output. Through this acquisition process, the server prepares the raw data source for subsequent feature extraction and generative artificial intelligence model processing.

[0246] Step 2: The server parses and cleans the document information.

[0247] The server takes the raw document information (HTML strings and text content in JSON fields) temporarily stored in step 1 as input and calls the text parsing module to perform data processing operations such as HTML tag removal, special character filtering, encoding standardization, sentence segmentation, and word segmentation. The server further uses natural language processing algorithms to perform part-of-speech tagging and sentence boundary recognition, removing advertising paragraphs and context-irrelevant noise text. The output is a standardized plain text sequence and its word segmentation results, which are stored in an intermediate data structure for subsequent feature extraction.

[0248] Step 3: The server performs preprocessing and feature extraction on the image information.

[0249] The server takes the original image files or video frames obtained in step 1 as input and performs image preprocessing operations such as resizing, center cropping, and pixel normalization on each image. Then, the server feeds the preprocessed images into a convolutional neural network or visual embedding model to calculate scene category labels (e.g., "city street," "indoor space"), temporal features (day / night), and image feature vectors. The server can also identify people, objects, or landmarks using an object detection model. The output is a set of features for each image, including scene category, possible location clues, and feature vectors. The server writes these features into structured data storage.

[0250] Step 4: The server extracts features such as location, time, behavior, atmosphere, and emotion from the text.

[0251] The server takes the standardized text sequence output from step 2 as input and calls a sequence labeling-based natural language processing model to assign labels such as location, time, behavior, and emotion to each word or phrase. The server extracts place names, dates, and times using named entity recognition, identifies action verbs (such as "take a walk," "drink coffee," and "rehearse") using pattern matching, and determines the overall or sentence-level emotion of the text (such as "romantic," "nervous," and "relaxed") using a sentiment analysis model. The server organizes these extraction results into feature records, outputting structured fields containing location, time, behavior, atmosphere, and emotion, and stores them in association with the image features from step 3. This data processing transforms unstructured text into a searchable set of features.

[0252] Step 5: The server builds and stores structured data records.

[0253] The server takes the features generated in steps 3 and 4 as input, associates text and image features according to object identifier, timestamp, and information source number, and generates a unified structured record. The server assigns a unique ID to each record and writes fields (object ID, time, location code, behavior type, atmosphere tag, sentiment tag, image feature vector, etc.) into a database table or key-value store. The output is a structured data set that can be retrieved at high speed. During this process, the server performs index creation and field normalization operations to enable efficient searching based on user request conditions.

[0254] Step 6: Users select the experience object and scene conditions on the terminal.

[0255] Users access the terminal's graphical user interface (GUI) as input, selecting a target object (such as a specific creator), a scene category (such as "travel," "performance," or "daily life") from a list, and further choosing the location (such as "Paris"), time (such as "dusk"), and experience format (such as "first-person perspective video"). The terminal converts the user's clicks, swipes, and inputs on the interface into structured request parameters (object identifier, scene type, time and location constraints, etc.), outputting a set of request condition data ready to be sent to the server.

[0256] Step 7: The terminal sends a request to the server.

[0257] The terminal takes the request condition data generated in step 6 as input, encapsulates it into an HTTP request or WebSocket message, and sends it to the interface address specified by the server via the communication network. The terminal includes user identification and authentication information in the request so that the server can verify permissions. The output is a request message received from the server. After sending, the terminal enters a waiting state, and a loading prompt can be displayed on the interface, waiting for the server to return the generated result or streaming media address.

[0258] Step 8: The server retrieves structured data that matches the request criteria.

[0259] The server takes the request conditions received in step 7 as input, accesses structured data storage, and performs queries on the object ID, location code, time range, behavior type, and sentiment tag fields. The server utilizes database indexes to accelerate retrieval and sorts candidate records according to preset similarity metrics (e.g., location similarity, time proximity, and sentiment matching). The server can perform weighted summation of multiple feature fields, selecting a set of records with higher weights as the basis for generating prompt statements. The output is a sorted set of feature records, including features such as location, time, behavior, atmosphere, and sentiment.

[0260] Step 9: The server automatically generates prompts based on feature records.

[0261] The server takes the feature record set obtained in step 8 as input and combines location, time, behavior, atmosphere, and sentiment features into a natural language description. The server can first select one or more representative records and fill the location name, time period, behavior description, and sentiment label into a predefined language template, or input them into the text generation sub-model to generate coherent sentences.

[0262] For example, the server generates the following prompt: "Please generate a scene of sunset along the Seine in Paris from the first-person perspective of internet celebrity A: the river is flowing slowly in the foreground, there are pedestrians and outdoor seating at cafes on the riverbank, and the Eiffel Tower can be seen in the distance. The overall atmosphere is romantic and relaxing, and the scene is realistic and detailed." The server outputs this text and translates it into other languages ​​as needed to adapt to the input requirements of different generative AI models. This step involves data processing including template filling, syntactic adjustment, and keyword checking to ensure that the prompts contain core constraints such as "first-person perspective."

[0263] Step 10: The server calls a generative artificial intelligence model to generate visual information.

[0264] The server takes the prompt generated in step 9 as input and feeds the text into the text encoder of a generative AI model to calculate the text embedding vector. The server then uses the text embedding as conditional input to a diffusion model or other generative network, performing multiple iterations on a graphics processor to progressively generate a clear image matching the prompt from a random, noisy image. If video is needed, the server calls the model multiple times to generate frames at different time points, or uses a video generation model to directly generate a sequence of frames. The output is a set of image frames or a complete video file. The server stores this visual information in the form of files or buffers for subsequent post-processing.

[0265] Step 11: The server performs post-processing and encoding on the generated visual information.

[0266] The server takes the image frames or video sequence output from step 10 as input and optimizes the video quality and smoothness. The server can invoke frame interpolation algorithms to estimate optical flow between adjacent frames and generate intermediate frames, thereby increasing the frame rate. Subsequently, the server uses the video encoding module to compress the frame sequence into a specified encoding format, setting parameters such as resolution, bit rate, and keyframe interval. The server then invokes the distribution format conversion module to divide the encoded video into multiple small segments and generate a playlist to adapt to streaming protocols. The output is a set of encoded media segment files and their index information, which constitute distribution data that can be played continuously by the terminal.

[0267] Step 12: The server provides a streaming entry point to the terminal and begins distribution.

[0268] The server takes the distribution data generated in step 11 as input and registers the corresponding access entry on the streaming media server, for example, generating a URL that can be accessed by an HTTP player or streaming media client. The server sends this URL or related session information as a response to the terminal. Upon receiving the response, the terminal uses this URL as input and initiates a segmentation request through its local player module. The server then sends the media segments to the terminal via the network in chronological order. The output is a continuously received media data stream on the terminal. During this process, the server can adjust the transmission bitrate based on the network conditions reported by the terminal to achieve adaptive transmission.

[0269] Step 13: The terminal decodes and presents visual information on a display device or head-mounted display device.

[0270] The terminal takes the media segment received in step 12 as input and sends it to the decoder for video decoding to restore continuous image frames. Depending on the user's current display device, the terminal renders the decoded image onto its local display screen or outputs it to the head-mounted display device via an interface. The terminal can perform geometric transformations or stereoscopic mapping on the image based on the head-mounted display device's field of view and resolution. The output is a first-person perspective view presented in the user's field of vision; the content seen by the user comes from visual information generated by the server in real time.

[0271] Step 14: The terminal collects the user's location information, posture information, and operation information and sends them back to the server.

[0272] The terminal takes inputs from its own sensors (such as a geolocation module, gyroscope, accelerometer, and optional eye tracker) and user interface events (clicks, touches, and gamepad actions), organizing this data into posture information, gaze information, location information, and operation information. The terminal packages this information and sends it to the server at regular time intervals. The output is a stream of messages reflecting the user's dynamic state, providing the server with the foundational data for subsequently dynamically adjusting prompts and generation conditions.

[0273] Step 15: The server dynamically updates the prompts or generation conditions based on user interaction and generates subsequent content.

[0274] The server takes the posture, gaze, location, and operation information received in step 14 as input and analyzes the user's current area of ​​focus and interaction intent. For example, when the server detects that the user's head is continuously turned to one side and their gaze is on a scene element in the image, the server can infer that the user expects to view that area in more detail. Based on this, the server updates the descriptions of scene focus, viewpoint direction, or behavior in the prompt, for example, changing "The Seine is ahead" to "Please move slowly towards the riverbank and street from the perspective of the outdoor seating area of ​​the café." The server then uses the updated prompt to re-invoke the generative AI model to generate new image frames or video clips and repeats the post-processing and distribution process from steps 10 to 12. The output is a visual information stream that continuously changes with the user's actions and interactions, thus achieving a dynamic and personalized first-person perspective experience.

[0275] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0276] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0277] With the development of generative artificial intelligence models, providing users with immersive experiences using first-person perspective images or videos has theoretically become possible. However, existing technologies still face the following challenges at the computer implementation level: (1) Existing systems typically rely on manually written prompts, directly inputting simple text descriptions into generative artificial intelligence models. They lack the ability to automatically extract behavioral patterns and scene patterns from large-scale heterogeneous data and structure them, resulting in the inability of the computer side to stably generate highly relevant and consistent first-person visual content.

[0278] (2) Existing generation systems often follow a "single generation - single presentation" process. Computer systems lack a mechanism to integrate image processing, natural language processing, and user behavior analysis into a single pipeline, making it impossible to achieve a cycle of "information source data → behavioral feature modeling → automatic generation of prompts → content generation → experience product management → personalized recommendation → feedback closed-loop optimization" under the same computing architecture. This makes it difficult for the system to dynamically update the generation logic based on the actual daily patterns of the subjects being utilized and the behavior of users.

[0279] (3) In terms of engineering implementation, most existing solutions regard information capture, content generation and distribution as independent modules. There is a lack of a technical solution that coordinates information acquisition software, image processing software, natural language processing software and generative artificial intelligence models on the server side, resulting in low efficiency of computing resources, redundant data flow, high interface coupling, and difficulty in expansion and maintenance.

[0280] (4) On the user side, traditional recommendation systems only recommend content based on browsing or purchase records. They rarely unify the behavioral characteristics of the subject being used with the user's perception feedback into the same data model. The computer cannot actively optimize the content structure during the generation stage to make the generated results more in line with the behavioral scenarios that the user wants to "immerse" in, thereby reducing the sense of immersion and unity.

[0281] Therefore, it is necessary to provide a new computer implementation method that integrates information source capture, feature extraction, automatic generation of prompts, invocation of generative artificial intelligence models, experience product management and personalized recommendations by building an end-to-end data processing pipeline on the server side. This will improve the data processing flow from the perspective of computer technology, enhance the consistency between generated content and the daily behavior of the subject being used, and form a closed loop of sustainable optimization within the same technical architecture.

[0282] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0283] In this invention, the server includes a unit for automatically acquiring image data, video data, and text data from an information source containing data related to the exploited subject using information acquisition software, and uniformly storing and managing the data; a unit for using image processing software to determine scene categories, detect target objects, and detect the presence of people in the image and video data, and using natural language processing software to extract location, behavior, and emotion from the text data, thereby generating feature information representing the daily behavior patterns of the exploited subject; and a unit for automatically generating prompt statements based on time, location, behavior, and emotion information in the feature information to define the first-person visual experience from the perspective of the exploited subject, and inputting the prompt statements into the generator. The system includes a generative artificial intelligence model (CAIM) for generating visual content that reproduces the perspective of the user, an information management and processing unit for converting the visual content output from the CAIM into a predetermined encoding format and aspect ratio, registering it as an experiential product, and managing the distribution of the experiential product, a unit for sending a list and detailed information of experiential products to the terminal, receiving user selection and acquisition requests for experiential products, and distributing visual content to the terminal according to the acquisition requests, and a unit for parsing user purchase history and key interest information, combining the feature information to generate experiential product recommendation information, and further collecting user evaluation information and usage information of the visual content to update the prompt generation conditions and feature information generation conditions. This creates an integrated data processing pipeline on the server side, encompassing information source data acquisition, behavioral feature modeling, automatic prompt generation, generative artificial intelligence model invocation, experiential product management, personalized recommendations, and feedback optimization. This improves the automation and consistency of data processing and content generation at the computer technology level, reduces manual intervention, enhances system resource utilization efficiency, and continuously improves the matching degree between the generated first-person visual content and the user's daily behavior and expectations.

[0284] A "system" refers to an overall technical solution consisting of one or more computing devices and the programs running on them, used to perform a series of functions such as information acquisition, data processing, content generation, content management, and content distribution.

[0285] A "server" refers to a computing device or group of devices that provides computing resources and data processing capabilities in a network environment, used to perform tasks such as information acquisition, feature extraction, prompt generation, invoking generative artificial intelligence models, and managing and distributing experiential products.

[0286] "Information source" refers to a data provider or data storage location that contains data related to the subject being used and is accessible via the network, including but not limited to social media platforms, video sharing platforms, web pages, and other data providing platforms.

[0287] "The subject being exploited" refers to an individual or group whose daily behavior, perspective, or life scene is used as the basis for generating first-person visual content, including but not limited to content creators, display subjects, or other information providers.

[0288] "Image data" refers to visual information data recorded in the form of still images, including but not limited to photographs, screenshots, and keyframe images extracted from videos.

[0289] "Video data" refers to visual information data recorded in the form of continuous frames in a time series, including but not limited to short videos, long videos, and other dynamic image data.

[0290] “Text data” refers to information recorded in text form that is related to the subject being used, including but not limited to text information contained in titles, descriptions, comments, and metadata.

[0291] "Information acquisition software" refers to a collection of programs used to automatically acquire image data, video data, and text data from information sources, including but not limited to page parsing programs, network request programs, and automatic operation programs.

[0292] "Image processing software" refers to a collection of programs used to analyze and identify the content of images and video data, including but not limited to scene recognition programs, face detection programs, and object detection programs.

[0293] "Natural Language Processing Software" refers to a collection of programs used to analyze and process text data, including but not limited to word segmentation programs, keyword extraction programs, and sentiment analysis programs.

[0294] "Scene category" refers to the type of environment or location determined based on the content of an image or video, including but not limited to abstract classifications such as indoor, outdoor, coffee shop, sports venue, and residential place.

[0295] "Target object" refers to the category of object identified by object detection in an image or video, including but not limited to abstract object types such as beverages, electronic devices, furniture, and sports equipment.

[0296] "Presence of people" refers to the attribute of whether a human body or face is detected in an image or video, and the estimated number of people.

[0297] "Location" refers to abstract information representing a spatial location or place extracted through the analysis of text or image data, including but not limited to city, building type, and place type.

[0298] "Behavior" refers to an abstract representation of the types of activities that the subject being used engages in within a certain scenario, including but not limited to work activities, leisure activities, sports activities, and travel activities.

[0299] "Emotion" refers to an abstract description of a psychological state or atmosphere obtained through text data or other data analysis, including but not limited to positive, negative, neutral and their intensity levels.

[0300] "Feature information" refers to a data set that provides a structured representation of the daily behavioral patterns of the subject being utilized, based on the results of image processing and natural language processing. This includes time features, location features, behavioral features, and emotional features.

[0301] "Daily behavior patterns" refer to recurring patterns of behavior and emotions in the subject being used, derived from statistical analysis and summarization of multiple data samples.

[0302] "Prompt statements" refer to textual descriptions used to instruct generative AI models on the characteristics of the content to be generated, including constraints on elements such as time, location, behavior, perspective, and emotion, to guide the generation of first-person visual content.

[0303] "Generative artificial intelligence models" refer to models that are based on machine learning algorithms and can automatically generate images or videos based on input prompts, including but not limited to still image generation models and video generation models.

[0304] "First-person visual experience" refers to a type of experience that presents visual content from the perspective of the subject being used, making the viewer feel as if they are observing the scene from that subject's point of view.

[0305] "Visual content" refers to the visual data output by generative artificial intelligence models and processed for presentation to users, including first-person experience images and first-person experience videos.

[0306] "Encoding format" refers to the data compression and encapsulation format used to digitally represent visual content, including but not limited to video encoding standards and image encoding standards.

[0307] "Aspect ratio" refers to the aspect ratio parameter of visual content, used to adapt to the display characteristics of different display terminals.

[0308] "Experience products" refer to digital services or content units that are primarily visual content, registered, managed, and provided to users for access and viewing on a platform in the form of commodities.

[0309] "Information management and processing" refers to the process of registering, updating, retrieving, authorizing, and distributing experiential products and their related data, including data storage management and access control management.

[0310] "Terminal" refers to an electronic device operated by a user, capable of communicating with a server and presenting visual content, including but not limited to mobile terminals, tablet devices, and computer devices.

[0311] "User" refers to an individual or organization that accesses the system through a terminal, browses and experiences products, makes purchases, and views visual content.

[0312] "Purchase history information" refers to data information that records the history of products purchased or used by users in the system, including purchase time, product identification, and usage records.

[0313] "Key interest information" refers to abstract information used to characterize a user's interest tendencies, derived from analysis of user behavior data, browsing preferences, and purchase history.

[0314] "Recommendation information" refers to structured data generated based on feature information and user-related information, used to prompt or recommend one or more products to users.

[0315] "Evaluation information" refers to feedback information provided by users regarding visual content or product experience, including but not limited to ratings, reviews, and tags.

[0316] "Usage information" refers to data representing a user's viewing behavior of visual content, including the number of times it is played, the duration of viewing, the jump position, and the interruption position.

[0317] "Generation conditions" refer to the parameters, rules, or thresholds used when generating prompt statements or characteristic information, which are used to control the generation process and the characteristics of the generated results.

[0318] In one embodiment of the present invention, the server serves as the main computing node, the terminal serves as the content presentation and interaction node, and the user interacts with the server through the terminal. The three work together to realize the complete technical process from information source data acquisition, feature extraction, prompt statement generation to generative artificial intelligence model invocation and experience product distribution.

[0319] A server can consist of one or more computer devices, which can be rack servers, cloud computing instances, or edge computing nodes. The server preferably includes a multi-core central processing unit (CPU), a graphics processing unit (GPU), main memory, and a high-speed network interface. The operating system running on the server can be a Unix-like operating system, such as a Linux distribution; the applications running on the server can be implemented using a general-purpose programming language, such as Python or other scripting languages.

[0320] The server's software architecture can include several functional modules such as information acquisition, image processing, natural language processing, feature modeling, prompt generation, generative artificial intelligence model invocation, content post-processing, experience product management, and recommendation and feedback. These modules can be deployed in a microservice architecture on different processes or nodes and communicate through remote procedure call interfaces.

[0321] Within the information acquisition module, the server can run page parsing software and automated operation software. Specifically, the server can use a webpage parsing library to perform tree-structure parsing of HTML documents to extract media links and text data from social media or video platform pages. The server can also use an automated browser control library to drive a headless browser to perform page scrolling, button clicks, and other operations, thereby acquiring dynamically loaded media resources without relying on manual input. When acquiring media resources, the server downloads the media files to a local temporary storage area, then uploads them to a cloud storage service via an object storage interface, and records the media identifier, storage path, and associated metadata in a relational database or document database.

[0322] In the image processing module, the server can use image processing libraries to process image data and image frames extracted from video data. The server can load a pre-trained convolutional neural network model as a scene classifier. The model structure can consist of multiple convolutional layers, pooling layers, and fully connected layers. The input is a uniformly sized image tensor, and the output is a probability distribution of multiple scene categories. Based on the category label with the highest probability, the server records the scene category corresponding to the image as a higher-level category such as "coffee shop," "sports venue," or "workplace."

[0323] In object detection processing, the server can use either an anchor-bound convolutional neural network model or a feature pyramid-based detection model to detect objects in the image. The server then performs post-processing on the detection results, including confidence thresholding and overlapping box suppression, to obtain a set of object category labels and corresponding location boxes. The server abstracts the detected object labels into higher-level concepts, such as "drinking utensils," "computing devices," and "sports equipment," thereby reducing category sparsity and facilitating statistical analysis.

[0324] In determining the presence of people, the server can use a human or face detection model based on a convolutional neural network to detect people in each image or video frame. Based on the number of detected people or faces, the server records a presence marker and an estimated number within the detected range. Through this image processing, the server generates a scene category vector, an object category vector, and a presence marker associated with each media sample.

[0325] Within the natural language processing module, the server can load dictionaries and language models to perform word segmentation and shallow syntactic analysis on text data. The server can use statistical methods or sequence labeling models based on pre-trained language models to tag each word with labels such as location, behavior, and emotion, extracting location names, action verbs, and emotional adjectives from the text. The server groups synonymous expressions into unified superordinate concepts; for example, unifying "coffee shop" and "coffee shop" into "coffee place," "video editing" and "video editing" into "content editing behavior," and "happy," "pleasant," and "comfortable" into "positive emotion." The server can also perform emotion classification on the entire text, using a neural network model based on a bidirectional encoder to encode sentence-level text, outputting emotion category and emotion intensity scores, and associating these results with media samples.

[0326] In the feature modeling module, the server associates image processing results with natural language processing results according to media identifiers. The server constructs a feature record for each media sample, including time features (discrete time intervals after shooting or uploading time), location features (abstract location type), behavioral features, object features, human presence features, and emotional features. The server establishes behavioral time series for the target entity in the database, sorts the feature records of multiple media samples according to the timeline, and performs statistical clustering on them.

[0327] The server can use algorithms based on frequent pattern mining or clustering algorithms to extract typical combinations of daily behaviors from feature records. For example, if the server finds that the combination of "coffee shop + content editing behavior + beverage utensils + computing device + positive emotions" occurs significantly more frequently than other combinations during the morning, then the server records this combination as a "morning work scene pattern" for the user. Similarly, the server can identify a "sports scene pattern" of "evening exercise location + running behavior + exercise equipment + positive emotions".

[0328] In the prompt generation module, the server uses the aforementioned behavioral pattern as input to generate prompts for invoking the generative AI model. The server can use a template generation method, which involves predefining a set of natural language templates and filling in elements such as time, location, behavior, objects, and emotions according to grammatical rules to generate complete and detailed prompts. For example, the server can generate the following prompt: "Please generate a 10-second video from the information provider's first-person perspective: a quiet coffee shop in the morning, the information provider sits by the window, drinking a latte and editing a video on a silver laptop, with soft sunlight and pedestrians passing by outside the window, the overall atmosphere is relaxed and focused." The server can also generate another type of prompt statement: "Please generate an image taken from the perspective of the information provider: a gym at dusk, the information provider is running on a treadmill, and the control panel of the treadmill and the mirror in front of him can be seen in the image, with a few people exercising around him vaguely reflected in the mirror." In another implementation, the server can use a small text generation model as the core of the prompt generation module. This model can employ an encoder-decoder structure to encode the input discrete feature vector (representing time, location, behavior, emotion, etc.) and generate natural language descriptions word by word according to the semantic sequence during the decoding stage. When training this text generation model, the server can use training corpus constructed from historical human descriptions, employing a cross-entropy loss function to compare the decoded output with the target sentence, and backpropagating the error to update the parameters. In this way, the server can automatically generate diverse prompts while maintaining coverage of core behavioral pattern elements.

[0329] In the generative AI model invocation module, the server sends the generated prompts and generation parameters (output type, resolution, duration, etc.) to the generative AI model service deployed on the graphics processor via a network interface. The generative AI model can be an image generation network or video generation network based on a diffusion process. Its basic structure may include a spatiotemporal coding module, a multi-layer self-attention module, and a denoising generation module. The server converts the prompts into vector representations, which are then used as conditional inputs and fed into the generation network along with random noise vectors. This process generates high-resolution visual content that conforms to the semantics of the prompts during multi-step iterative denoising.

[0330] During model training, the server can use contrastive loss or conditional generation loss functions to update network parameters by comparing the distance between generated and real images in the feature space. During inference, the server only performs forward generation computation. By structuring the behavioral patterns extracted from the information source into conditional features and encoding them into the generative AI model through prompts, the server enables the generative network to converge in a direction highly consistent with the daily behavior of the subject being utilized during multiple rounds of denoising, thereby improving the matching degree between the generated results and the source behavior.

[0331] In the content post-processing module, the server encodes and converts the images or videos output by the generative artificial intelligence model. The server can use multimedia processing tools to encode the raw video stream into a specific video encoding standard and adjust the resolution to fit the aspect ratio of mainstream terminal displays. The server can also add metadata to the video as needed, including product identifiers, entity identifiers, and generation time, for subsequent retrieval and management. The server generates corresponding thumbnails for previewing on the terminal interface.

[0332] In the product management module, the server registers each piece of visual content as a product entry. The server records fields such as title, brief description, media address, thumbnail address, price parameters, and visibility permission settings for each product in the database. When managing product experiences, the server can set different visibility ranges based on business strategies, such as public visibility or visibility limited to specific user groups. The server can also write inverted index entries to the indexing service based on generation time and associated behavioral patterns, thereby supporting fast retrieval by time period, scene category, or emotional atmosphere.

[0333] In the recommendation and feedback module, the server parses the user's purchase history and key interest information. The server can aggregate and analyze user behavior data such as browsing history, click behavior, viewing duration, and number of repeated views on the terminal, mapping this data to an abstract interest space, such as "preferring coffee scenes," "preferring sports scenes," and "preferring nighttime atmospheres." The server combines this interest representation with the behavioral pattern characteristics of the user to calculate a set of suitable experiential products for that user and generate a ranked recommendation list. When generating recommendations, the server can use recommendation algorithms based on matrix factorization or graph neural networks, outputting a recommendation score based on the similarity between the user vector and the experiential product vector.

[0334] In terms of feedback collection, the server receives user ratings, textual reviews, and viewing interruption points for visual content. The server performs statistical analysis and modeling on this feedback, such as analyzing the relationship between the time, location, and emotional combination of the prompt statements and user ratings, thereby updating the weight parameters or rule weights in the prompt statement generation module. The server can adjust the emphasis of different scene elements in the prompt statements; for example, when users generally rate scenes containing "excessive noise descriptions" poorly, the server can reduce the frequency of such elements when generating prompt statements to improve user satisfaction.

[0335] The terminal can be a smart communication device, a mobile computing device, or a desktop computing device. It includes a display screen, an audio output device, and a network communication module. When running a dedicated application, the terminal receives a list of products and detailed data from the server, displaying product thumbnails and descriptions through a graphical user interface. After the user selects a specific product, the terminal sends a request to the server and, upon obtaining the media address, loads visual content from a content delivery network via a streaming media protocol. The terminal uses a local decoder to decode the video or image data and plays or displays it on the screen from a first-person perspective.

[0336] Users browse, search, purchase, and watch content through the terminal interface. During viewing, users can pause, drag the progress bar, or switch resolutions via the control interface. After watching, users can submit ratings and written reviews, which are sent to the server via the terminal and incorporated into the feedback dataset. With repeated user interaction, the server gradually develops a more refined user interest model, allowing for more targeted adjustments to prompts and generation criteria in subsequent new product experiences.

[0337] This invention, by constructing the aforementioned modular structure and data flow path on the server side, creates a closed loop encompassing information acquisition, feature modeling, prompt generation, and generative artificial intelligence model invocation. Unlike simple manual prompt writing, the server transforms large amounts of heterogeneous data into standardized feature information and behavioral patterns, then generates prompts based on this, thereby reducing noise caused by human subjective bias and forming a repeatable and scalable automated generation process within the computer. By statistically analyzing the correspondence between behavioral patterns and user feedback, the server dynamically adjusts the prompt generation rules and model parameters, gradually converging the generation process to a state with better technical indicators, such as higher scene matching, longer user dwell time, and a lower proportion of invalid generation.

[0338] Because the server employs structured data storage, batch media processing, and parallel job scheduling technologies, the system can achieve high throughput and high scalability when handling a large number of utilized subjects and large-scale media data. The server transmits feature information between different modules through a unified data structure, reducing redundant parsing and format conversion, and lowering computational redundancy and communication burden. Therefore, this invention not only provides users with an immersive first-person perspective experience at the content level, but also, at the computer technology level, improves data processing efficiency, generation accuracy, and overall system performance through specific data structure design and specific feature extraction and generation process design.

[0339] use Figure 13 The processing flow is explained.

[0340] Step 1: The server retrieves raw multimedia data from the information source. The server takes the identifier of the exploited entity as input and reads a list of information source addresses associated with that entity from the database (e.g., social media platform homepage URLs, video platform channel URLs). The server invokes information retrieval software to send HTTP requests to the information sources via a network interface, or loads the target webpage through an automated browser control program. After loading the webpage, the server parses the HTML structure and script-generated content, extracting image links, video links, title text, and descriptive text. For the obtained media links, the server sequentially downloads the corresponding image and video files to its local cache directory. Subsequently, the server uploads these files to a cloud storage service and generates a record for each file in the database, storing the exploited entity identifier, original URL, storage path, and associated text. The input to this step is the exploited entity identifier and the list of information source addresses; the output is the original image data and video data stored in cloud storage, and the text data and media metadata stored in the database. The operations performed by the server during data processing include network data retrieval, HTML structure parsing, string extraction, and file transfer operations.

[0341] Step 2: The server performs visual feature analysis on image and video data. The server takes the storage paths of the image and video data stored in step 1 as input and downloads media files from cloud storage to the local working directory of the computing node. The server loads each image using image processing software, scaling or cropping it into a tensor of uniform size. This image tensor is then input into a pre-trained convolutional neural network model to perform forward inference, obtaining probability distributions for multiple scene categories. The category with the highest probability is selected as the scene label for that image. Simultaneously, the server uses an object detection model to detect objects in the image, performing convolutional feature extraction, candidate region generation, and classification regression operations to obtain a set of object categories and corresponding bounding boxes. Targets with confidence levels below a preset threshold are discarded, and overlapping bounding boxes are merged using non-maximum suppression, ultimately outputting a set of abstract object labels (e.g., beverage containers, computing devices). For video data, the server uses multimedia processing tools to extract keyframes at fixed time intervals. These keyframes are input as images into the aforementioned scene classification and object detection models. The frequency of each category is calculated, and the category with the highest frequency is used as the main scene label for the video. The server also uses a face or human detection model to determine the presence and estimated number of people in the scene. The input for this step is image and video files, and the output is scene category labels, object category labels, and human presence markers for each media sample. The specific data processing performed by the server in this step includes multidimensional tensor operations, convolution operations, probability calculations, threshold comparisons, and statistical induction.

[0342] Step 3: The server extracts natural language features from the text data. The server takes the title text, description text, and optional comment text recorded in step 1 as input and reads these strings in batches from the database. The server first uses natural language processing software to segment the text, dividing the continuous character sequence into word sequences. Then, the server uses part-of-speech tagging and entity recognition algorithms to assign category labels to each word, such as location words, action verbs, and emotion adjectives. The server maps synonyms or near-synonyms to unified higher-level concepts through table lookups and rule matching; for example, it maps specific place names to abstract location categories such as "city place" and "coffee place," and specific behavior descriptions to abstract behavior categories such as "work behavior" and "exercise behavior." The server also uses an emotion classification model to encode the entire text, calculating its probability in three categories: positive, neutral, and negative, and using the category with the highest probability and emotion intensity as the emotion feature of the text. The output of this step is the location feature, behavior feature, and emotion feature associated with each piece of text data. The specific data operations performed by the server in this step include word segmentation, sequence labeling inference, vocabulary normalization, probability calculation, and emotion score generation.

[0343] Step 4: The server builds characteristics of the daily behavioral patterns of the subjects being exploited. The server takes the visual features output from step 2 and the text features output from step 3 as input. It merges features from the same media sample using media identifiers to form feature records containing fields such as time, scene category, location category, behavior category, object category, presence of people, and emotion category. The server categorizes the time into predefined time periods (e.g., morning, afternoon, evening, night) based on the media's shooting or uploading time. The server groups the database by the subject being exploited and performs statistical analysis on all feature records for that subject. The server can use frequent itemset mining algorithms to calculate the frequency of combinations of fields such as time period, scene category, behavior category, and emotion category, identifying combinations with support higher than a preset threshold; alternatively, it can use clustering algorithms to map feature records to a high-dimensional vector space, cluster them based on distance metrics, and extract representative feature combinations from each cluster. The server uses these high-frequency combinations or cluster centers as the subject's daily behavior patterns, recording them as pattern identifiers and their corresponding feature combinations. The output of this step is a set of daily behavior pattern features for each subject being exploited. The main data operations performed by the server in this step include counting statistics, association rule mining, vector distance calculation, and cluster partitioning.

[0344] Step 5: The server automatically generates prompts for generative artificial intelligence models. The server takes the daily behavior pattern features constructed in step 4 as input and selects elements such as target time period, scene category, behavior category, object category, and emotion category. In template generation mode, the server uses a preset language template, such as a sentence structure containing "time description," "location description," "behavior description," "object description," and "emotion description." The server converts these elements into natural language fragments based on the pattern features and fills them into placeholders in the template to generate a complete prompt text. For example, after inputting the "morning coffee" scene pattern, the server generates the following prompt: "Please generate a 10-second video from the information provider's first-person perspective: a quiet coffee shop in the morning, the information provider sits by the window, drinking a latte and editing a video on a silver laptop, with soft sunlight and pedestrians passing by outside the window, the overall atmosphere is relaxed and focused." The server can also generate the following prompt after the motion scene mode is entered: "Please generate an image taken from the perspective of the information provider: a gym at dusk, the information provider is running on a treadmill, and the control panel of the treadmill and the mirror in front of him can be seen in the image, with a few people exercising around him vaguely reflected in the mirror." In another implementation, the server encodes pattern features into discrete feature vectors, inputs them into a text generation model, and performs forward inference using an encoder-decoder structure to generate prompts word by word. The output of this step is a set of prompts for each behavior pattern. The operations performed by the server in this step include string concatenation, template replacement, feature-to-word mapping, and sequence generation inference.

[0345] Step 6: The server invokes a generative artificial intelligence model to generate visual content. The server takes the prompt message generated in step 5 and preset generation parameters (output type, resolution, duration, frame rate, etc.) as input and sends the prompt message to the generative AI model service deployed on the graphics processor via a network interface. When sending the request, the server packages the prompt message into a text field and attaches the generation configuration parameters. Upon receiving the request, the generative AI model encodes the prompt message into a vector representation and inputs it along with a random noise vector into the conditional generation network. This network performs multiple rounds of linear transformation, nonlinear activation, and denoising operations to gradually generate image or video frames that meet the conditions. After the generation is complete, the server receives the raw output data stream from the generative AI model service. The server performs integrity verification on the received data and saves it as an image file or video file. The output of this step is a first-person visual content file corresponding to the prompt message. The main data operations performed by the server in this step are request packaging and unpacking, data transmission control, and file writing.

[0346] Step 7: The server encodes and converts the generated visual content and registers the experiential products. The server takes the visual content file obtained in step 6 as input and calls a multimedia processing tool to encode and convert it. According to the system's preset encoding format and target resolution, the server performs transcoding operations on the video content, including compression according to specified encoding standards, adjustment of frame rate and screen size; and performs resolution adjustment and compression quality control on the image content. The server generates one or more thumbnails for each visual content item for quick preview in a list. Subsequently, the server uploads the processed visual content file to cloud storage and registers a new experience product record in the database, writing the main identifier, visual content address, thumbnail address, product title, description, price, and tag information. The output of this step is a list of distributeable experience product items and their associated media files. The specific actions performed by the server in this step include multimedia encoding calculations, image scaling calculations, object storage uploads, and database write operations.

[0347] Step 8: The server provides the terminal with a list and details of the products to experience. The server takes the user identifier and optional filter criteria as input and reads matching product records from the database. It filters out invisible products based on the user's access permissions and organizes the titles, descriptions, thumbnail URLs, prices, and tags of the remaining products into structured response data. The server then sends this response data to the requesting terminal via a network interface. The output of this step is a list of product experiences that the terminal can parse. The operations performed by the server during this process include conditional querying, sorting, result set assembly, and network data encapsulation. The terminal takes the list data returned by the server as input, requests thumbnail images from the content delivery network, and displays the product entries in the graphical user interface for the user to browse and select.

[0348] Step 9: Users select and purchase trial products on the terminal. Users input the list of products displayed on the terminal and select a desired product via touch or pointer. The terminal packages the selected product identifier and user account information and sends it to the server as a purchase request. After verifying the user's identity and product status, the server sends a payment instruction containing the order amount and product identifier to the payment service. After the user confirms the payment on the terminal, the payment service notifies the server of the result. The server updates its database based on the payment result, creating an authorization record between the user and the product, indicating that the user has obtained the right to view the product. The output of this step is the updated authorization information and the purchase completion status. After receiving confirmation from the server, the terminal displays a "Purchased" status on the interface and provides a "View" entry.

[0349] Step 10: The terminal retrieves visual content from the server and presents it to the user. The terminal sends a content retrieval request to the server, taking the user's viewing request and the target experience product identifier as input. The server verifies whether the user has viewing authorization. If the verification is successful, the server reads the media address and access token corresponding to the experience product from the database and returns the actual playback URL to the terminal. The terminal uses this URL as input, loads the visual content data stream from the content delivery network via a streaming media protocol, and decodes the video or image data using a local decoder. The terminal plays or displays the decoded image on the screen, and the audio is output through speakers or headphones. The user sees visual content generated from the first-person perspective of the utilized subject on the terminal. The output of this step is a first-person visual experience perceptible to the user. The specific actions performed by the terminal in this step include network data reception, buffer control, decoding operations, and image frame rendering.

[0350] Step 11: The server collects user feedback and updates the generation and recommendation logic accordingly. The server takes user feedback data uploaded from terminals as input, including ratings, text reviews, viewing duration, exit time and location, and whether the content was viewed repeatedly. The server associates and stores the feedback data with corresponding product experiences and prompt labels. The server performs statistical analysis on feedback from multiple users, such as calculating average ratings across different scenario categories, completion rates for content across different time periods, and the relationship between different emotional descriptions in prompts and user dwell time. Based on the analysis results, the server adjusts the parameter weights in the prompt generation module, for example, increasing the probability of high-rated scenario elements appearing in prompts and reducing combinations of elements associated with low completion rates. The server can also update the threshold for mining daily behavior patterns, making the model more inclined to select patterns validated by positive user feedback. The output of this step is the updated prompt generation conditions, feature modeling parameters, and recommendation weights. Through this feedback-based data processing, the server continuously optimizes the conditions called by the generative AI model, thereby improving matching accuracy and user experience quality in subsequent generation.

[0351] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0352] As generative artificial intelligence models are increasingly used in image and video generation, existing content generation systems based on these models still face the following technical challenges: First, existing systems mostly use static, manually written prompts to issue instructions to generative artificial intelligence models. They lack automatic parsing and feature extraction of publicly available electronic information (images, videos, text) of the target object. Therefore, it is difficult to accurately reconstruct the viewpoint and scene attributes of the represented object at the computer level, resulting in a disconnect between the generated content and the real context, and insufficient immersion.

[0353] Second, existing systems lack a closed-loop, calculable linkage mechanism between generated content and user emotional state. Emotional signals such as facial expressions and voice generated by users during playback are typically not collected and analyzed in real time by the system. The generation module also fails to dynamically update prompts or adjust the content editing structure based on these signals. This prevents the system from adaptively optimizing generated content according to changes in user emotions during runtime, thus limiting the depth of personalized experiences.

[0354] Third, existing systems largely rely on simple recommendations when utilizing user viewing history, operation history, purchase history, and feedback information. They lack the computer-based means to model and learn these historical data in relation to the conditions for generating prompts and editing content. Therefore, the system cannot continuously learn the correspondence between user preferences and emotional responses on both the input side of the generative AI model and the content editing side, making it difficult to iteratively improve the quality of user-generated content.

[0355] Fourth, in highly immersive scenarios such as virtual reality, traditional systems often simply project ordinary videos onto display devices without integrating format conversion and viewpoint editing logic for head-mounted displays into the system architecture. They also lack a unified technical framework for structuring generated content on the server side and outputting regenerated data adapted to different display terminals, resulting in low rendering efficiency and poor content reusability.

[0356] Based on the above problems, a system architecture is needed that integrates electronic information parsing, feature extraction, automatic generation and dynamic updating of prompt statements, closed-loop feedback of sentiment analysis, and multi-terminal adaptive output on the server side, thereby achieving this at the computer technology level. (1) Automatically construct viewpoint and scene representation of the object being represented using publicly available electronic information; (2) Incorporate the user's emotional state into the prompt generation and content editing process in real time; (3) Update generation rules and editing conditions based on user history behavior and feedback; (4) Generate playable data in a unified manner for virtual reality and ordinary display, thereby improving the overall system’s intelligence and operating efficiency.

[0357] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0358] In this invention, the server includes a processing unit for automatically acquiring image information, video information, and text information from publicly available electronic information of the object being represented, and extracting feature information representing the viewpoint and scene attributes of the object being represented through image analysis and natural language processing; a generation unit for generating prompts for inputting into a generative artificial intelligence model based on the feature information and prompts received from the user, and calling the generative artificial intelligence model to generate content containing reproduced viewpoints; a format conversion unit for converting the generated content into a virtual reality format or a common display format and generating playback data that can be presented in a head-mounted display device or display terminal; an emotion analysis unit for parsing facial expression information and voice information from the user terminal or head-mounted display device to determine the user's emotional state; and a content control and distribution unit for dynamically updating the prompts input to the generative artificial intelligence model and / or editing the viewpoint, scene, and edit structure of the generated content according to the emotion analysis results and playback status to generate re-experience content adapted to the emotional state and distributing it to the user terminal through a communication network. This allows for the formation of an integrated technical chain within the computer system, encompassing the parsing of publicly available electronic information, generation of prompts, reasoning using generative artificial intelligence models, closed-loop feedback of emotional states, and multi-terminal adaptive output. This enables viewpoint reconstruction, emotionally adaptive editing, and individualized optimization of generated content. While reducing reliance on manual intervention and hard-coded rules, it improves the matching degree between generated content and user emotions and preferences, enhances system generation efficiency and the quality of immersive experiences, thereby improving the content generation and distribution mechanism based on generative artificial intelligence models at the computer technology level.

[0359] "System" refers to the overall technical structure of a system consisting of at least one server, at least one user terminal, and optional electronic devices such as head-mounted displays, which are interconnected through a communication network and work together to perform information acquisition, parsing, generation, editing, and distribution to provide users with re-experience content.

[0360] "Information processing device" refers to an electronic device that includes hardware resources such as processor, memory and communication interface, and is capable of executing program code to collect, analyze, generate and control electronic information.

[0361] A processor is a computing unit in an information processing device that executes program instructions, performs logical operations on input data, executes control flow, and processes data read and write. It can be a central processing unit, a graphics processing unit, or other programmable logic devices.

[0362] "Subject to representation" refers to the target individual or entity whose publicly available electronic information is used as a reference for content generation, including but not limited to content providers, information publishers, or other individuals with a certain influence.

[0363] "Public electronic information" refers to digital information that is publicly available through network services or storage media and can be obtained by information processing devices, including image information, video information, text information and their related metadata.

[0364] "Image information" refers to static visual data recorded in the form of a pixel matrix, including photographs, illustrations, or other still images.

[0365] "Video information" refers to dynamic visual data consisting of multiple consecutive image frames and optional audio streams, including short videos, long videos, or streaming media data.

[0366] "Text information" refers to text data recorded in the form of character sequences, including titles, body text, descriptive text, comments, and tag information.

[0367] "Image analysis and processing" refers to the automated processing of frame data in image and video information, such as feature extraction, object detection, scene classification, and style analysis, to obtain structured data representing the content and visual attributes of the image.

[0368] Natural Language Processing (NLP) refers to the process of performing operations such as word segmentation, part-of-speech tagging, entity recognition, sentiment analysis, and semantic understanding on textual information in order to extract key semantic elements and contextual relationships.

[0369] "Feature information" refers to a structured data set extracted from publicly available electronic information through operations such as image analysis and natural language processing, used to represent the viewpoint, scene attributes, contextual features, and semantic elements of the represented object.

[0370] "Viewpoint" refers to the observation position and direction presented in the generated visual content, used to reproduce the spatial positional relationship of the object being represented in its subjective view of the world and the first-person or specific third-person visual angle.

[0371] "Scene attributes" refer to characteristic information related to the environment, time, location, atmosphere, activity type, etc. in the generated content, used to describe the contextual characteristics of the visual content.

[0372] "Prompt statements" refer to text instructions input into a generative artificial intelligence model, which specify the subject, viewpoint, scene, style, emotion, and other generation conditions of the generated content.

[0373] "Generative artificial intelligence models" refer to artificial intelligence models that are based on machine learning or deep learning technologies and can automatically generate new images, videos or other data based on input prompts and optional auxiliary features.

[0374] "Generated content" refers to the set of visual information output by a generative artificial intelligence model after receiving prompts and feature information, including images or videos that reproduce the viewpoint of the represented object.

[0375] "Virtual reality format" refers to a content encoding method suitable for presentation in head-mounted display devices that can update the perspective with head movements, including panoramic video, stereoscopic video, or other three-dimensional scene representation formats.

[0376] "Common display format" refers to the standard video or image encoding format used for playback on flat panel display terminals (such as smart terminals, tablet terminals, and monitors).

[0377] "Playback data" refers to digital media data that has been encoded or packaged and can be decoded and presented by display terminals or head-mounted display devices, including file format and streaming media format.

[0378] "Head-mounted display devices" refer to display devices worn on the user's head that can update the displayed content in real time according to the head's posture, providing the user with an immersive visual experience.

[0379] "Display terminal" refers to an electronic device that has display function and can receive and play playback data, including portable terminals, fixed terminals and other devices with display screens.

[0380] "User terminal" refers to a terminal device operated by a user for sending requests, receiving content, presenting an interface, and collecting facial and voice information.

[0381] "Face information" refers to facial images or facial feature data of a user acquired through an image acquisition device, used to reflect the user's current facial expression state.

[0382] "Voice information" refers to the sound signals and characteristic data of a user obtained through an audio acquisition device, which are used to reflect the user's voice content, tone, and emotional characteristics.

[0383] "Emotional state" refers to the type and intensity of a user's emotions, determined by analyzing facial and vocal information, including categories such as joy, excitement, sadness, anxiety, and calmness.

[0384] "Emotional analysis methods" refer to the combination of hardware and software used to receive facial and voice information and infer the user's emotional state through methods such as pattern recognition or machine learning.

[0385] "Playback status" refers to the status information related to the playback process of generated content on the user terminal or head-mounted display device, including playback progress, playback duration, pause / resume operations, and switching behavior.

[0386] "Re-experience content" refers to visual content that is generated based on sentiment analysis results and editing, and is used to allow users to have an immersive experience from the perspective of the represented object.

[0387] "Communication network" refers to wired or wireless communication infrastructure that transmits data between servers and user terminals and head-mounted display devices, including the Internet, mobile communication networks, and local area networks.

[0388] "Viewing history" refers to records of information about the types, durations, order, and interactions a user has viewed in the past.

[0389] "Operation history" refers to the recorded information of user actions such as selection, clicking, swiping, and switching during system interface or playback.

[0390] "Purchase history" refers to the transaction records and purchase preference information of users in the system for purchasing re-experienced content or related products.

[0391] "Evaluation information" refers to user ratings, comments, tags, or other subjective feedback data regarding re-experiencing content.

[0392] "Generation conditions" refer to a set of parameters and rules used to control the characteristics of the output results when a generative artificial intelligence model performs content generation processing, including the structure of prompt statements, model parameter settings, and constraints.

[0393] "Editing conditions" refers to a set of parameters and rules used to determine the editing method, splicing strategy, and display logic when editing generated content in terms of viewpoint, scene, and time sequence.

[0394] "Generation rules" refer to a set of predefined or learned logic used to automatically generate prompt statements, set generation conditions, and adjust the expression of prompt statements, which guide the input construction of generative artificial intelligence models.

[0395] In this embodiment of the invention, the server, terminal, and user each assume different functional roles, collaboratively realizing the generation of re-experience content and the adaptive presentation of emotions based on a generative artificial intelligence model and prompts. The following provides a detailed description of the hardware configuration, software modules, data structures, algorithm processing flow, and technical effects.

[0396] I. Overall Hardware and Software Composition of the System In one implementation, the server employs a computing device with graphics processing capabilities, such as a general-purpose computer including a multi-core central processing unit, a graphics processing unit, main memory, and solid-state storage. In one example, the server uses an x86 architecture processor and a general-purpose GPU, runs an operating system (such as a UNIX-like operating system), and deploys a relational database management system, web server software, and various artificial intelligence frameworks and multimedia processing tools.

[0397] The server, at the software level, includes: a natural language processing module, an image / video parsing module, a feature management module, a prompt generation module, a generative artificial intelligence model inference module, a format conversion module, a sentiment analysis interface module, a content editing module, and a content distribution module. In one specific implementation, the server uses the following open-source or commercial software as implementation tools: natural language processing libraries (such as a word segmentation and dependency parsing library), deep learning frameworks (such as a tensor computation framework), image processing libraries (such as a computer vision library), multimedia transcoding tools (such as an audio / video processing tool), and virtual reality engines (such as a type of 3D engine).

[0398] In one embodiment, the terminal is a smart terminal device, such as a control terminal for a smartphone, tablet computer, or head-mounted display device. The terminal has hardware including a processor, memory, display unit, camera, microphone, and wireless communication module. The terminal runs an operating system (e.g., a mobile operating system) and dedicated applications to send user selection and prompt statements to the server, receive re-experience content, and capture user facial images and voice signals.

[0399] Users operate the device through a graphical user interface, including selecting the object to be displayed, entering prompts, issuing play / pause / switch commands, and moving the viewpoint on the head-mounted display.

[0400] II. Server-side data modeling and feature extraction In one implementation, the server obtains publicly available electronic information about the objects to be represented from the network service provider through an information acquisition module. Internally, the server maintains an "object file" data structure for each represented object, including a unique identifier, account identifier, content index list, and statistical information. The server assigns a record structure to each piece of electronic information (image, video, text), including fields such as: content ID, object ID, media type, storage path, timestamp, platform source, text content, and initial tag.

[0401] When performing natural language processing, the server represents text information as a sequence of tags. The server first segments and tags the text, then maps each tag to a vector space using pre-trained word vectors or sub-word encoding. In one implementation, the server encodes the entire text into a context-dependent hidden vector sequence using a bidirectional attention network or a Transformer-like structure. Based on this, it uses a classification head to output scene categories (e.g., "morning," "performance," "beach") and contextual labels (e.g., "quiet," "excited"). The server treats these results as "text features" and stores them as fixed-length vectors and corresponding label lists.

[0402] In image / video parsing, the server inputs each image or video frame into a pre-trained convolutional neural network or visual Transformer model. The server obtains multi-level feature maps through forward propagation and outputs labels such as scene category, subject category, environment type (indoor / outdoor), lighting features (bright / dim), and color tone features (cool / warm) at the top level through a classification head. The server also generates fixed-length visual feature vectors through global average pooling of the feature maps. For videos, the server performs temporal pooling on multiple frames to obtain the temporal integrated features of the entire video.

[0403] The server integrates textual and visual features in the feature management module. In a specific manner, the server combines textual and visual feature vectors from the same time window or topic into a "scene feature vector" through concatenation or weighted summation. The server records a scene description structure for each scene in the object archive, including the scene feature vector, scene tag list, time range, and representative media paths.

[0404] Through the above modeling method, the server transforms the originally unstructured text and image content into a structured and indexable feature set, which can be directly queried and synthesized when generating subsequent content, reducing the repeated parsing calculations during each generation and improving the overall processing speed and resource utilization efficiency.

[0405] III. Server-side prompt statement generation and dynamic updating In the prompt generation module, the server constructs prompts suitable for generative artificial intelligence models using a combination of template-based and language model-based approaches. In one implementation, the server first retrieves the records with the highest matching degree to the scene tag from the scene feature database based on the user-selected object and target scene, and extracts its scene tag list, context tag, and time features.

[0406] The server then calls the language template module to populate the key elements into predefined sentence templates. For example, the server generates the following prompt statement in one instance: "Based on the breakfast-related content posted by Subject A on social media, please generate a video of a breakfast scene from Subject A's first-person perspective. The video should include details such as preparing ingredients in the kitchen, setting out tableware, and sitting at the table to eat. The overall style should be realistic and natural, with soft morning sunlight." In another example, the server generates prompts for a virtual reality scene: "Please generate a 360-degree panoramic video from the perspective of object B, showing a walk on a beach on a summer evening, including waves, sunset, and distant city silhouettes, suitable for playback on a head-mounted display." In emotion-adaptation scenarios, the server will also dynamically modify or add prompts based on the emotion analysis results. For example, when the emotion analysis result is excitement, the server generates: "The user is currently very excited. Please generate a clip taken from the perspective of object C in the center of the performance stage, highlighting the scene of lights, fireworks and audience singing together, with a duration of about 15 seconds." When the sentiment analysis result is slightly melancholic, the server generates: "The user appears slightly melancholy. Please generate a heartwarming video of the user interacting with their pet in a quiet room from the perspective of object D, with soft colors and relaxing background music." In one implementation, the server uses a sequence-to-sequence language model to refine and expand the initial template prompts, thereby improving naturalness and detail while maintaining the structured elements. This approach allows the server to automatically adapt to different object and scenario combinations while ensuring controllable content standards are met, reducing the burden of manual rule writing.

[0407] IV. Generative Artificial Intelligence Model Structure and Inference Processing The server deploys a diffusion-based image / video generation model or other generative model in the generative artificial intelligence model inference module. In one implementation, the server employs a diffusion generation architecture based on a U-shaped network and a temporal convolution / attention module. This architecture uses Gaussian noise as initial input and gradually generates image or video frames through a multi-step inverse diffusion process.

[0408] During inference, the server first encodes the prompts into a sequence of context vectors using a text encoder (e.g., a multi-layer self-attention network). The server can also use scene feature vectors as additional conditional inputs, mapping them to the same semantic space through a conditional encoding layer. The server inputs both textual conditional embeddings and scene feature embeddings into the conditional control channel of the generative model, subjecting the generation process to the combined constraints of multiple features such as viewpoint and scene attributes.

[0409] When generating video, the server adds a temporal attention module or a 3D convolution module to the model to ensure the continuity of the generated results between frames. The server specifies the number of frames and frame rate for the video output and encodes and concatenates the single-frame image sequence into a standard video format file after generation.

[0410] The server improves generation throughput through batch inference and memory reuse techniques on GPUs. Since the server has already encoded and cached the features of objects and the scene in the preceding steps, the generation module can directly access these encodings, avoiding repeated execution of heavy parsing models and thus significantly reducing the average generation latency.

[0411] V. Server-side content format conversion and editing The server, in its format conversion module, is responsible for converting the generated content into either a virtual reality format or a standard display format. In one implementation, the server maps the generated 2D video frames onto spherical or cubic surfaces in the virtual scene, renders them using a 3D engine, and outputs a sequence of equidistant rectangular panoramic images. The server then calls a video encoding tool to compress the panoramic images into a panoramic video file suitable for head-mounted displays.

[0412] In the content editing module, the server controls viewpoint changes and segment editing based on prompts and emotional states. In one implementation, the server maintains a timeline and viewpoint switching table for each generated segment, specifying when to use which viewing direction and viewing distance. Upon receiving new emotional feedback, the server can select a suitable segment from the candidate segment library and splice it to the end of the current timeline, or replace the subsequent part of the currently playing segment, thereby achieving adaptive emotional adjustment while maintaining playback continuity.

[0413] This editing method enables the server to technically achieve non-linear playback path selection and dynamic viewpoint control based on emotional state, unlike traditional linear video playback. Because the server uses pre-calculated features and pre-generated segments during the editing process, the system can complete switching decisions and segment splicing in a short time, resulting in a virtually lag-free real-time adaptation effect on the terminal side.

[0414] VI. Terminal-side data collection and presentation In one implementation, the terminal allows the user to select the object to be represented and the target experience type through a graphical user interface. The user can input prompts such as, "I want to experience a daily routine of preparing breakfast from the perspective of object A," or "Please generate a summer beach evening scene seen from the first-person perspective of object B." The terminal then sends this text and selection information to the server via a secure communication protocol.

[0415] During the presentation phase, the terminal receives playback data of the generated content from the server and calls the system's multimedia interface for decoding and display. In scenarios involving head-mounted displays, the terminal transmits panoramic video files to the head-mounted display or acts as a network relay, allowing the head-mounted display to directly retrieve streaming media from the server. When the user wears the head-mounted display, they change their viewing angle through head movements. The head-mounted display adjusts its rendering direction based on the output of its built-in sensors, thereby achieving an immersive experience of observing the surrounding environment from the user's perspective.

[0416] When collecting emotional data, the terminal periodically captures facial image frames from the user's camera and acquires audio clips through the microphone. The terminal can either run a lightweight facial expression recognition network locally or upload the features to a server-side emotional analysis module for classification. Within each time window, the terminal sends the obtained emotional tags and intensity values, along with the current playback position, to the server, allowing the server to construct a trajectory of the user's emotional changes over time.

[0417] VII. Server-Side Sentiment Analysis and Generation Rule Learning The server receives facial expression information, voice features, or pre-classified emotion tags uploaded by the terminal in the emotion parsing interface module. In one implementation, the server employs a multimodal emotion recognition model that combines facial key point features, convolutional features of expression images, and audio time-frequency features, outputting a multidimensional emotion probability distribution through a multilayer perceptron or attention fusion network. The server determines the dominant emotion state and its intensity based on the maximum probability or a preset threshold.

[0418] In the content control module, the server uses emotional state and playback progress information to select appropriate prompt templates and content editing strategies. After multiple interactions, the server associates and stores the user's emotional trajectory with the prompts used, editing decisions, and user feedback information (such as ratings and dwell time) to form training samples.

[0419] In the rule learning module, the server updates the prompt generation rules and editing conditions based on these samples. The server can use gradient descent to optimize the parameterized policy network, enabling it to automatically select prompt structure and content combinations that better improve user satisfaction based on user history and current context. In this way, the server technically constructs an adaptive rule system, allowing the system to gradually improve the matching accuracy between content and user preferences as usage time increases.

[0420] VIII. Explanation of Technical Effects and Causal Relationship By using the structured feature representation and caching mechanism described above, the server reduces the repeated parsing of the original image and text data, allowing the inference calls of the generative artificial intelligence model to focus on the necessary condition combinations. This increases the number of generation requests that can be processed per unit time under the same hardware resources, directly improving processing speed and system throughput.

[0421] By introducing scene feature vectors and multimodal condition control, the server not only relies on the text description of prompts, but also internally uses vector space constraints to generate results. This makes the generated content more closely match the real activity patterns and visual style of the represented object, thereby reducing errors that deviate from the context and improving the accuracy and consistency of the generated content.

[0422] By dynamically updating prompts and editing structures based on emotional feedback, the server transforms the generated content from a one-time static entity into an adjustable media stream coupled with the user's state in real time. This mechanism differs from traditional recommendation methods that rely solely on historical clicks or browsing records. It introduces emotional control factors during the generation phase, thereby altering the content generation path and resource allocation strategy within the computer, enhancing the consistency and personalization of the overall experience.

[0423] By recording and learning from user behavior and feedback over a long period, the server gradually evolves the rules for generating and editing prompts from fixed templates into a parameterized policy network. The system can automatically discover the optimal prompt structure and scenario combination for different user groups without requiring manual adjustments to each rule. This learning mechanism achieves a balance between computational efficiency and knowledge representation capabilities, avoiding the maintenance difficulties of purely manual rule-based systems.

[0424] The terminal performs partial feature extraction and preprocessing locally or at the edge, sending compressed features instead of the complete original video / audio data to the server, thereby reducing network bandwidth pressure and improving real-time performance. Simultaneously, since the server can perform emotion recognition and content decision-making based on the compressed features, the overall system communication load is significantly reduced.

[0425] The above embodiments illustrate how the server, terminal, and user collaborate around the generative artificial intelligence model and prompts, enabling the system not only to automatically generate and play visual content, but also to achieve structured feature management, emotion-adaptive generation rule learning, and cross-terminal format adaptation within the computer. This technically improves the processing efficiency, generation accuracy, and user experience quality of the content generation system. The above embodiments can be used individually or in combination, with various modifications and substitutions made without departing from the spirit of the invention.

[0426] use Figure 14 The processing flow is explained.

[0427] Step 1: The user selects an object on the terminal and enters an initial prompt statement.

[0428] Input: The list of selectable objects to be displayed on the interface, the search box, and the text input box.

[0429] Output: Request data containing object identifiers and user-input text.

[0430] Users launch the application on their devices, click on a displayed object in the list, or enter a name in the search box to search. The device then displays the search results locally. Users enter their requirements in the text input box, such as "I want to experience a daily routine of preparing breakfast from the perspective of object A," or "Please generate a summer beach evening scene from the first-person perspective of object B." The device packages the user-selected object identifier and the original prompt into structured data (e.g., a record containing object ID, original text, and timestamps) and sends it to the server via a network protocol.

[0431] Step 2: The server obtains raw data related to the object from publicly available electronic information sources.

[0432] Input: Object identifier from the terminal, platform account information, or link information.

[0433] Output: A raw data set containing images, videos, text, and metadata.

[0434] The server queries its internal configuration based on the object identifier to obtain the object's account or page links on multiple content platforms. The server calls the interfaces provided by each platform or executes web scraping programs to request and download image files, video files, and text descriptions one by one, saves the files to the storage system, and generates a record for each piece of content in the database. The record includes content ID, object ID, media type, file path, timestamp, and source.

[0435] Step 3: The server performs natural language processing on the text information and generates text features.

[0436] Input: The set of text information obtained in step 2.

[0437] Output: Text feature vector and scene label list for each text.

[0438] The server schedules the natural language processing module to read each piece of text. It first performs word segmentation and part-of-speech tagging, then converts the word sequence into a set of hidden vectors using a pre-trained text encoding network. The server performs pooling operations on the entire text to obtain a fixed-length text feature vector. It then uses a classification head to calculate the probability of belonging to each scene category and emotion category, selecting the tags with the highest probabilities as scene labels, such as "morning," "kitchen," and "performance." The server writes the text feature vector and labels into the corresponding records in the database.

[0439] Step 4: The server performs visual analysis on images and videos and generates visual features.

[0440] Input: Image files and video frame data obtained in step 2.

[0441] Output: Visual feature vector and visual label for each image or video clip.

[0442] The server invokes the image / video parsing module, inputting keyframes from each image or video into a convolutional neural network or visual Transformer model. Forward propagation is performed to obtain multi-layered feature maps and a top-level classification output. The server selects labels such as subject category (person, object), scene category (indoor / outdoor), lighting, and color style from the classification output, and performs global pooling on the feature maps to obtain fixed-length visual feature vectors. For videos, the server averages or performs attention-weighted summation on features from multiple frames along the timeline to generate video-level visual features. The server stores the visual features and labels in a database and associates them with the corresponding media records.

[0443] Step 5: The server combines textual and visual features to construct a scene feature vector.

[0444] Input: Text features output from step 3, visual features output from step 4, and time and object association information.

[0445] Output: Scene feature vector and comprehensive scene label for each scene.

[0446] The server pairs text records with image / video records based on timestamps and object IDs to form a candidate scene set. For each scene, the server combines the corresponding text feature vector and visual feature vector according to preset rules (such as concatenation or weighted summation) to form a higher-dimensional scene feature vector; at the same time, it merges text tags and visual tags into a comprehensive tag list. The server generates a scene record for each scene, which includes the scene feature vector, tag list, and representative media path, and stores it in a scene feature table for subsequent generation.

[0447] Step 6: The server generates basic prompts for generative artificial intelligence models.

[0448] Input: User's original prompt statement, scene feature vector, and comprehensive scene label.

[0449] Output: Structured and expanded base message text.

[0450] The server reads the user's original prompt, parses keywords such as time, location, and activity, and compares them with the corresponding scene tags for the user, selecting scenes with high matching scores. The server then inputs the selected scene tags, such as "breakfast," "kitchen," and "morning light," into a template to generate a more complete basic prompt, for example: "Based on breakfast-related content posted by user A on social media, please generate a video of a breakfast scene from user A's first-person perspective. The video should include details such as preparing ingredients in the kitchen, setting out tableware, and sitting at the table to eat, with a realistic and natural overall style and soft morning sunlight." The server outputs the constructed prompt as text and records it in the generation task log.

[0451] Step 7: The server adds prompts based on the target output type and sets the generation parameters.

[0452] Input: basic prompts, user-selected experience type (normal video or virtual reality), and system default parameters.

[0453] Output: The build task configuration, including detailed prompts and build parameters.

[0454] The server determines whether the user selects a virtual reality experience. If virtual reality is selected, the server adds panoramic and viewpoint requirements to the prompt, for example: "Please generate a 360-degree panoramic video showing a walk on a beach on a summer evening from the perspective of object B, including waves, sunset, and distant city outlines, suitable for playback on a head-mounted display." The server simultaneously generates corresponding generation parameters, such as image resolution, video duration, frame rate, viewpoint type, and whether it is stereoscopic, and stores these parameters along with the prompt in a generation queue as input configuration for the generative artificial intelligence model.

[0455] Step 8: The server encodes the prompts and prepares them as input for the generative artificial intelligence model.

[0456] Input: Detailed prompt text, scene feature vector, and generation parameters.

[0457] Output: Model input tensors (text encoding vector, multimodal condition vector, control parameters).

[0458] The server invokes the text encoder module to convert the prompt statement into a sequence of subwords and embed them as word vectors. A multi-layer self-attention network is then used to compute the context encoding vector for the entire sentence. The server projects the scene feature vector onto the same feature space as the text encoding using a linear mapping, and combines the text encoding and scene features to form a multimodal conditional vector. Based on the generation parameters, the server generates a control tensor, such as target resolution and time duration, and packages all conditions into an input batch, sending it to the generative AI model's inference module.

[0459] Step 9: The server performs inference in a generative artificial intelligence model to generate images or videos.

[0460] Input: The model input tensor output in step 8.

[0461] Output: Generated image or video frame data.

[0462] The server loads a diffusion-based generative model or other generative AI model onto the GPU, feeding the input conditions into the model's conditional channel. The server first initializes the latent space with a noise tensor, then performs an inverse diffusion process at multiple time steps. At each step, based on the current noise state, text conditional vector, and scene feature vector, it calculates a noise estimate using a U-shaped network and updates the latent representation. After multiple iterations, the server maps the final latent representation to the image space through a decoder, obtaining image frames that match the prompt description. If generating video, the server repeats the above process in the time dimension, obtaining a series of consecutive frames. The server caches the generated image or frame data as a result.

[0463] Step 10: The server encodes the generated results and saves them as media files.

[0464] Input: Image / frame data generated in step 9, generation parameters.

[0465] Output: Playable image or video file.

[0466] The server selects an appropriate encoding format based on the generation parameters, encoding a single image into an image file and compressing multiple frames sequentially into a video file. The server calls a multimedia encoding library to encapsulate the frame sequence into a standard container format and writes metadata such as duration and frame rate. The server saves the media file to the file system or object storage and creates a corresponding record in the database, recording the file location, format, object ID, and task ID, for subsequent download and use by the terminal.

[0467] Step 11: The server converts regular videos into virtual reality playback formats (when virtual reality is required).

[0468] Input: The generated video file and the target virtual reality format parameters.

[0469] Output: Panoramic virtual reality video file or stereoscopic virtual reality video file.

[0470] The server loads the generated video in the virtual reality format conversion module, maps each frame as a texture map onto the internal surface of a virtual sphere or cube, renders the scene using a 3D engine, and generates a sequence of equidistant rectangular panoramic images. The server adds left and right eye views as needed to support stereoscopic display, then uses a video encoding tool to compress the panoramic image sequence into a virtual reality format video and writes it into standard metadata recognizable by the head-mounted display device. The server updates the virtual reality file path of the media corresponding to this task in the database.

[0471] Step 12: The terminal requests and downloads the generated content from the server.

[0472] Input: The user's content selection request on the terminal, the content list returned by the server, and the file address.

[0473] Output: Media files stored locally on the terminal or a playable streaming media connection.

[0474] The terminal sends a request to the server based on the user's selected object and experience type. The server returns a list of available content and corresponding file links. The terminal displays the list for the user to select. When the user selects content, the terminal retrieves the corresponding image or video file from the server or content distribution node via the communication network and stores it in its local storage area or buffer. If streaming media is used, the terminal saves the playback address and starts a network streaming playback session.

[0475] Step 13: The terminal collects user facial expressions and voice in real time during playback.

[0476] Input: media content in playback, video frames captured by the terminal's camera, and audio clips captured by the microphone.

[0477] Output: Facial expression feature data and speech feature data or emotion tags.

[0478] When playing generated content, the terminal activates its front-facing camera and microphone to capture the user's facial images and short audio data at fixed time intervals. The terminal can execute a lightweight facial expression recognition network locally to extract keypoint coordinates and facial feature vectors from the facial images; it can also extract spectral features from the audio for acoustic emotion recognition. The terminal encodes or compresses these features and uses them as input for emotion analysis, or directly calls an external emotion recognition library to return preliminary emotion labels and confidence scores.

[0479] Step 14: The terminal sends the emotional data and playback status to the server.

[0480] Input: Facial features, voice features or emotion tags obtained in step 13, as well as the current content ID and playback progress.

[0481] Output: Emotional feedback data packets sent to the server.

[0482] At each predetermined time slice, the terminal packages the current emotion data along with the playback position, content identifier, and user identifier into a small data packet and uploads it to the server via the communication network. The terminal can employ compression and throttling strategies to reduce bandwidth usage, such as sending only when there is a significant change in emotion, or averaging the features of consecutive frames before sending.

[0483] Step 15: The server parses the sentiment data and determines the user's current sentiment state.

[0484] Input: The emotional feature data or emotional tags reported in step 14 and the playback information.

[0485] Output: Dominant sentiment state and intensity value for each time slice.

[0486] The server receives emotional features from the terminal in the emotion analysis module, inputs facial and audio feature vectors into the multimodal emotion recognition network, and calculates the probability distribution of each emotion category through forward propagation. The server selects the emotion category with the highest probability as the current dominant emotion and records the probability value as the emotion intensity. The server stores the emotion state along with the corresponding playback time and position and content ID in the emotion trajectory record.

[0487] Step 16: The server dynamically adjusts the prompts and content strategies based on the emotional state.

[0488] Input: Current emotional state, historical emotional trajectory, and information about the currently playing content.

[0489] Output: New sentiment-adaptive prompts and content editing decisions.

[0490] The server analyzes emotional trajectories in the content control module. If it detects high levels of excitement over a continuous period, it determines that a high-energy segment needs to be added; if it detects increased sadness or anxiety, it selects a soothing scene. Based on these judgments, the server generates new prompts, such as "The user is currently very excited. Please generate a segment shot from the perspective of object C standing in the center of the stage, highlighting the lighting, fireworks, and audience singing along, approximately 15 seconds long." or "The user is currently slightly melancholic. Please generate a heartwarming video shot from the perspective of object D interacting with a pet in a quiet room, with soft tones and relaxing background music." The server also formulates a splicing strategy, such as at what point to insert new segments and which existing segments to retain.

[0491] Step 17: The server uses the new prompts to invoke a generative artificial intelligence model to generate supplementary fragments.

[0492] Input: Emotional adaptation prompts generated in step 16, relevant scene feature vectors, and supplementary fragment parameters (short duration, high adaptation, etc.).

[0493] Output: Short video clips or image sequences that are coherent with the current content.

[0494] The server generates short video clips or image sequences using new prompts and scene features through an encoding and reasoning process similar to steps 8 and 9. To reduce generation time, the server employs fewer diffusion steps or lower resolution to meet real-time requirements. After generation, the server encodes the clips into video files, records the start and end times of the clips, and registers the clip's purpose and suggested insertion points in the content editing log.

[0495] Step 18: The server sends the new supplementary fragments and editing instructions to the terminal.

[0496] Input: Path to the generated supplementary fragment file, insertion time point, and switching conditions.

[0497] Output: Addresses of supplementary fragments and control information that can be called by the terminal.

[0498] The server sends a notification message to the terminal via a control interface, containing the download address of the new segment, the suggested insertion position on the current playback timeline, and switching conditions (such as seamless splicing after the current segment ends). The server can pre-generate multiple bitrate versions for the new segment so that the terminal can select the appropriate stream based on network conditions.

[0499] Step 19: The terminal preloads supplementary clips and switches playback accordingly.

[0500] Input: The address of the supplementary fragment and the switching command issued by the server.

[0501] Output: A continuously rendered, adjusted playback stream on the terminal.

[0502] Upon receiving the notification, the terminal downloads or buffers supplementary segments from the server in the background. When the current playback progress reaches a designated switching point, it stops the current mainstream media source and seamlessly switches to the playback of the supplementary segment. For virtual reality content, the terminal ensures that the viewpoint direction of the new segment is similar to the viewpoint at the end of the previous segment to reduce discomfort during the transition. The terminal presents this dynamically adapted playback stream to the user, allowing the user to visually experience the content changing in accordance with their own emotional state.

[0503] Step 20: Users complete the experience and can provide optional feedback.

[0504] Input: The final content displayed on the terminal and the terminal feedback interface.

[0505] Output: Feedback data such as ratings, reviews, or preference selections.

[0506] After watching the entire re-experience content, users can rate the content, write comments, or select preference tags such as "more exciting" or "more relaxing" on the terminal interface. The terminal uploads this feedback data, along with the corresponding content ID and user ID, to the server to train the prompt generation rules and editing strategies, thereby providing more personalized content to the same or similar users in subsequent interactions.

[0507] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0508] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0509] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0510] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0511] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0512] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0513] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0514] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0515] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0516] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0517] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0518] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0519] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0520] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0521] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0522] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0523] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0524] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0525] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0526] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0527] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0528] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0529] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0530] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0531] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0532] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0533] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0534] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0535] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0536] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0537] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0538] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0539] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0540] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0541] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0542] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0543] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0544] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0545] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0546] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0547] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0548] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0549] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0550] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0551] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0552] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0553] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0554] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0555] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0556] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0557] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0558] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0559] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0560] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0561] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0562] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0563] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0564] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0565] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0566] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0567] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0568] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0569] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0570] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0571] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0572] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0573] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0574] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0575] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0576] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0577] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0578] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0579] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0580] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0581] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0582] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0583] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0584] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0585] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0586] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0587] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0588] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0589] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0590] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0591] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0592] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0593] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0594] In addition, the following notes are provided in response to the above explanation.

[0595] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring data from information sources related to an information provider, and analyzing the text, image, or video information contained in the data to extract feature information representing the perspective and expressive tendencies of the information provider. A device for automatically generating prompt statements for a generative artificial intelligence model based on the feature information and the demand information obtained from the user, and inputting the prompt statements and generation conditions into the generative artificial intelligence model to generate visual information that mimics the perspective of the object provided by the information. An apparatus for acquiring a user’s emotional state and preference information, and adjusting the prompt statement or the generation conditions based on the emotional state and preference information, thereby controlling the content or presentation of the generated visual information. A device for constructing an experiential product from the generated visual information, and for generating and recording identification information, explanatory information, and conditions provided about the experiential product; An apparatus for receiving a purchase request for the experience product from a user terminal, performing a settlement process and granting the user access to the experience product, and distributing the generated visual information to the user terminal in a streaming media or download manner. A device for obtaining evaluation or feedback information about the experienced product from the user and reflecting the evaluation or feedback information in the updating of the feature information and the generation of the prompt statement.

[0596] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for obtaining information from the user is further configured to obtain the user's historical acquisition records and key interest information. The device for analysis is configured to select experiential products that make it easier for the user to obtain a special sense of unity and immersion based on the correspondence between the feature information and the generated visual information and the historical acquisition records and key interest information, and to provide the user with product recommendation prompts for the experiential products.

[0597] (Note 3) The information processing system according to Appendix 1 is characterized in that, The device for analysis is configured to perform weighted relearning or reanalysis processing on additional text information, image information, or video information obtained from the information source. The weighting is based on the evaluation information or feedback information, and the feature information representing the object's perspective and expressive tendency of the information is dynamically updated according to the result of the relearning or reanalysis. The updated feature information is then used to improve the content of subsequently generated prompts and visual information.

[0598] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: An apparatus for acquiring chronologically ordered documentary and image information from information sources related to an object, extracting features related to location, time, behavior, atmosphere, and emotion from the documentary and image information, and recording them as structured data; A device for retrieving structured data based on user-requested conditions and automatically generating prompts to instruct a generative artificial intelligence model to generate visual information from a first-person perspective, using features about location, time, behavior, atmosphere, and emotion contained in the retrieval results. A device for taking the prompt statement as input to execute the generative artificial intelligence model, thereby generating visual information as image or video information simulating a human perspective. A device for performing post-processing on the visual information, including frame generation, frame interpolation, encoding, and distribution format conversion, to generate distribution data for streaming transmission. A device for sequentially sending the distribution data for streaming transmission to a user terminal via a communication network, and for presenting the visual information in an immersive first-person perspective by a display device or head-mounted display device connected to the user terminal. An apparatus for dynamically changing the generation conditions of the prompt statement or the generative artificial intelligence model based on operation information or posture information obtained from the user terminal, and updating the generation or distribution of the visual information. An apparatus for parsing user-related historical and preference information, adjusting the content of the prompts and visual information at the user level, thereby generating personalized experience and suggestion information; A device for collecting evaluation and reaction information from users and incorporating it into the subsequent generation and processing of the prompt statements and the control conditions of the generative artificial intelligence model.

[0599] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for dynamically changing the prompt statement or the generation conditions based on operation information or posture information obtained from the user terminal is configured to obtain location information, gaze information, head posture information, and operation information from the user terminal, and update the descriptions of location, viewpoint direction, and behavior contained in the prompt statement successively based on the obtained information, and re-execute the generative artificial intelligence model using the updated prompt statement, thereby generating visual information from a first-person perspective that changes with the user's actions.

[0600] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating personalized experience information and suggestion information is configured to generate suggestion information for an experience object associated with visual information from a human perspective based on the analysis results of document information and image information obtained from the information source, historical information and preference information related to the user, and evaluation information and reaction information from the user, and to personally prompt the user terminal with the suggestion information as product information or service information about the experience object.

[0601] Example 2 (Note 1) An information processing system, characterized in that it comprises: A unit for connecting to an information source containing data related to the subject being exploited, and for automatically acquiring image data, video data, and text data related to the subject being exploited from the information source using information acquisition software; The image and video data are processed using image processing software to determine the scene category, target object, and presence of people. The text data is processed using natural language processing software to extract location, behavior, and emotional expression. Based on the determination results and extraction results, a unit representing feature information of the daily behavior pattern of the subject being used is generated. The system uses the time information, location information, behavioral information and emotional information contained in the feature information to generate prompt statements that define the first-person visual experience from the perspective of the subject being exploited, and inputs the prompt statements into the generative artificial intelligence model to instruct the generation of units that reproduce the visual content from the perspective of the subject being exploited. An information management processing unit for converting visual content output from the generative artificial intelligence model into a predetermined encoding format and aspect ratio, registering the visual content as an experiential product, and managing the experiential product as a distribution target; A unit for sending a list of information and detailed information related to the product to the terminal, receiving the user's selection and acquisition request for the product via the terminal, and distributing the visual content to the terminal according to the acquisition request; A unit for analyzing purchase history information and key interest information related to users, and generating recommendation information for the experience products based on the analysis results and the feature information to improve the sense of unity between the utilized subject and the user; A unit for collecting evaluation information and usage information from the user regarding the visual content, and updating the generation conditions of the prompt statement or the generation conditions of the feature information based on the information, thereby adjusting the content of the subsequently generated visual content.

[0602] (Note 2) According to the information processing system described in Appendix 1, the information acquisition software includes a program for parsing the content displayed on the screen and extracting media data, and a program for performing automatic operations; the image processing software includes a general image recognition program and a target detection program; and the natural language processing software includes a program for performing word segmentation, word extraction, and sentiment analysis.

[0603] (Note 3) According to the information processing system described in Appendix 1, the generative artificial intelligence model includes at least one generative processing device among models for generating still images and models for generating videos, and by including time periods, scene categories, and emotional atmospheres corresponding to the typical behavioral patterns of the exploited subject in the prompt statements, the generative artificial intelligence model generates first-person experience videos or first-person experience images that enable users to have an immersive experience of the world of the exploited subject.

[0604] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A means for extracting feature information representing the viewpoint and scene attributes of the object being represented by acquiring image information, video information and text information from publicly available electronic information of the object being represented by a processor in an information processing device, performing image parsing processing on the image information and the video information, and performing natural language processing on the text information. Means for generating prompts for inputting into a generative artificial intelligence model based on the feature information and prompts received from the user, and inputting the prompts into the generative artificial intelligence model to generate generated content containing visual information that reproduces the viewpoint of the represented object; Means for converting the generated content into at least one of a virtual reality format or a common display format, and generating playback data that can be presented in a head-mounted display device or display terminal; An emotion analysis method for determining a user's emotional state by analyzing facial and voice information obtained from a user terminal or the head-mounted display device. A means for dynamically updating the prompts input to the generative artificial intelligence model and / or changing the viewpoint, scene, and editing structure of the generated content based on the emotional state determined by the emotional analysis means and the playback status of the generated content, thereby generating re-experience content that is adapted to the user's emotional state. Means for distributing the re-experience content to a user terminal via a communication network, and enabling the user terminal or the head-mounted display device to play the re-experience content.

[0605] (Note 2) According to the information processing system described in Appendix 1, the processor records user viewing history, operation history, purchase history, and evaluation information obtained from the user terminal or the head-mounted display device, and learns or updates the generation conditions of prompt statements input to the generative artificial intelligence model and the editing conditions of the re-experience content based on the records, thereby serving as a means to continuously provide each user with individually optimized re-experience content with different viewpoint and scene compositions.

[0606] (Note 3) According to the information processing system described in Appendix 1, the processor records the correspondence between the user's emotional state obtained by the emotion analysis method, the prompt statements input to the generative artificial intelligence model, and the user feedback information on the re-experience content, and adjusts the vocabulary, expression intensity, and scene-specific content of the prompt statements based on the correspondence, thereby updating the generation rules of the prompt statements to improve the immersion and sense of unity of the subsequently generated re-experience content.

Claims

1. An information processing system, characterized in that, Includes a processor, the processor being configured to: Input prompt words into the generative artificial intelligence model to generate visual information, so as to generate images or videos that reproduce the perspective of celebrities or internet celebrities; Identify the user's emotions and adjust the generated content accordingly; The generated content is distributed to the user's terminal.

2. The information processing system according to claim 1, characterized in that, The processor is also configured to analyze a user’s past purchase history and interests to generate personalized product recommendations, enabling the user to feel a special connection and immerse themselves in the world of the target object.

3. The information processing system according to claim 1, characterized in that, The processor is also configured to collect user feedback and adjust the generated content based on the collected user feedback during the next content generation, in order to provide users with a valuable experience and deepen the sense of connection between users and the target audience.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A