Information processing system
Patent Information
- Application Number
- CN202610271596.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-06
- Publication Date
- 2026-09-22
AI Technical Summary
本发明通过以生成式人工智能模型为核心,对人格对象生成、选举预测模拟以及结果输出进行一体化自动处理,不仅能够利用小规模事前调查和人口统计数据构建大规模虚拟选民群体,提升预测精度和覆盖度,还能根据媒体类型和使用场景对预测结果进行格式化与结构化输出,从而有效克服现有技术中预测能力有限、更新不及时以及输出形式单一的问题
[0004]To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to: input prompts to generate personality objects based on small-scale pre-conference questionnaire results and demographic information of each constituency into a generative artificial intelligence model, so that the generative artificial intelligence model constructs multiple personality objects reflecting the demographic structure and voting tendencies of the constituency based on the questionnaire results and demographic information; input prompts to perform predictive simulations of elected candidates in an election using the generated personality objects into the generative artificial intelligence model, so that the generative artificial intelligence model outputs simulation results related to the candidate election prediction based on the matching relationship between personality object attributes and candidate policies; and input prompts to analyze the prediction results and output the prediction results in a form suitable for sale to various information media into the generative artificial intelligence model, so that the generative artificial intelligence model generates prediction result data adapted to different media needs. This invention integrates and automatically processes personality object generation, election prediction simulation, and result output using a generative artificial intelligence model as its core. It can not only construct a large-scale virtual voter group using small-scale pre-surveys and demographic data to improve prediction accuracy and coverage, but also format and structure the prediction results according to media type and usage scenario, thereby effectively overcoming the problems of limited prediction capabilities, untimely updates, and single output format in existing technologies.
Smart Images

Figure CN122797724A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.
[0003] Existing election result predictions typically rely on manually designed statistical models and limited-scale opinion poll data. Due to limited sample sizes, discrepancies between sample structure and actual population structure, and the inability to promptly reflect dynamic changes in voter preferences as candidates' policies evolve, existing technologies suffer from several significant shortcomings: First, they struggle to perform refined modeling and simulation of large-scale voter groups while maintaining controllable costs, leading to insufficient prediction accuracy. Second, traditional statistical methods often require professionals to manually set model structures and parameters, resulting in long update cycles, poor flexibility, and an inability to quickly respond to changes in the election situation. Third, prediction results are mostly provided in the form of static tables or simple statistical values, lacking customized output formats for different media (such as national newspapers, local newspapers, and television media), limiting the application value of prediction results in media reporting and election strategy formulation. Therefore, a system is needed that can utilize generative artificial intelligence models to automatically generate personality objects based on small-scale pre-conference questionnaire results and demographic data from various constituencies, automatically simulate election result predictions, and output the prediction results in a format adapted to different media needs, thereby improving the accuracy, efficiency, and practicality of predictions. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to: input prompts to generate personality objects based on small-scale pre-conference questionnaire results and demographic information of each constituency into a generative artificial intelligence model, so that the generative artificial intelligence model constructs multiple personality objects reflecting the demographic structure and voting tendencies of the constituency based on the questionnaire results and demographic information; input prompts to perform predictive simulations of elected candidates in an election using the generated personality objects into the generative artificial intelligence model, so that the generative artificial intelligence model outputs simulation results related to the candidate election prediction based on the matching relationship between personality object attributes and candidate policies; and input prompts to analyze the prediction results and output the prediction results in a form suitable for sale to various information media into the generative artificial intelligence model, so that the generative artificial intelligence model generates prediction result data adapted to different media needs. This invention integrates and automatically processes personality object generation, election prediction simulation, and result output using a generative artificial intelligence model as its core. It can not only construct a large-scale virtual voter group using small-scale pre-surveys and demographic data to improve prediction accuracy and coverage, but also format and structure the prediction results according to media type and usage scenario, thereby effectively overcoming the problems of limited prediction capabilities, untimely updates, and single output format in existing technologies.
[0005] "System" refers to a collection of devices and programs, including hardware and software, used to perform election result prediction-related processing. It includes at least a processor and storage devices, network interfaces, etc., that communicate with the processor, and is an overall technical solution for realizing functions such as data reception, processing, storage, and result output.
[0006] A processor is an electronic component or computing unit used to execute instructions and process input data. It can be a single central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or a combination thereof, used to input prompts into generative artificial intelligence models and process the results generated by the models.
[0007] "Hints" refer to textual information, structured information, or a combination thereof generated by a processor and input into a generative artificial intelligence model. This information is used to instruct the generative artificial intelligence model to perform specific tasks, such as generating personality objects based on questionnaire results and demographic information, performing election prediction simulations, and analyzing and formatting the prediction results.
[0008] "Generative artificial intelligence model" refers to an artificial intelligence model that is trained on large-scale data and can automatically generate corresponding output content based on input prompts. It includes, but is not limited to, generative language models and generative multimodal models, and is used to generate personality objects, election prediction results and structured data or text data related to media output based on the prompts.
[0009] "Pre-election survey results" refers to the data obtained from small-scale surveys conducted among voters in each constituency before the formal election. This data includes at least basic information about the respondents and their support intentions for candidates or political parties, and is used to reflect initial voting tendencies.
[0010] "Demographic information" refers to statistical data obtained for the population of each electoral district, including but not limited to age distribution, gender distribution, occupational category distribution, and education level distribution. It is used to describe the demographic characteristics of the electoral district so as to simulate and reconstruct the demographic structure when generating personality objects.
[0011] "Electoral district" refers to a geographical or administrative region divided according to the electoral system. Within each electoral district, voters elect representatives or candidates. In this invention, each electoral district corresponds to a set of pre-survey questionnaire results and demographic information.
[0012] "Personality object" refers to a virtual individual data unit generated by a generative artificial intelligence model based on the results of a prior questionnaire and demographic information, used to simulate the characteristics of real voters. Each personality object has at least attributes such as age, gender, occupation, education level, and support tendency for candidates, and is used to represent a certain number of real voters in election prediction simulation.
[0013] "Predictive simulation of elected candidates in elections" refers to a series of processes that use generated personality objects and generative artificial intelligence models or related algorithms to estimate the expected votes of different candidates in a given constituency, thereby determining the candidate most likely to be elected in that constituency.
[0014] "Prediction results" refers to the output data obtained from the prediction simulation based on the candidates to be elected in the election, including but not limited to the predicted vote share, probability of being elected, number of seats predicted for each candidate, and corresponding statistical analysis information.
[0015] "Information media" refers to various media organizations that release or disseminate election-related information to the public, including but not limited to national newspapers, local newspapers, television stations, and other news or communication organizations. These organizations can use the prediction results output by the system of this invention for reporting or analysis.
[0016] "Output in a format suitable for sale to various information media" refers to analyzing, organizing, and formatting the forecast results so that they are presented in one or more forms such as structured data, reports, charts, and text descriptions, in order to meet the needs of different information media in terms of content depth, layout, and data structure, thereby facilitating direct adoption or secondary processing by the information media. Attached Figure Description
[0017] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0018] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0019] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0020] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0021] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0022] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0023] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0024] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0025] Figure 9 This represents an emotion map that maps multiple emotions.
[0026] Figure 10 This represents an emotion map that maps multiple emotions.
[0027] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0028] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0029] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0030] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0031] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.
[0032] First, let me explain the terminology used in the following instructions.
[0033] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0034] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0035] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0036] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0037] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0038] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0039] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0040] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0041] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0042] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0043] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0044] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0045] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0046] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0047] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0048] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0049] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0050] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0051] With the increasing demand for predicting election results using survey information and demographic data, traditional prediction systems based on rules or fixed statistical models are significantly inadequate in handling large-scale, multi-dimensional data and complex uncertainties. Specifically, existing technologies mainly suffer from the following problems: (1) Data preprocessing relies on manual configuration. After collecting survey information and population attribute information from multiple sources, the existing system often requires manual scripting to complete missing values, remove outliers, and standardize fields. Different developers have different implementation methods, which makes it difficult to guarantee data quality, make the processing flow difficult to reuse, and result in poor overall robustness and maintainability of the system.
[0052] (2) Limited ability to generate simulated samples at the individual level. Traditional methods often construct simulated voter data through simple weighted sampling or replication and expansion based on historical samples. It is difficult to generate individual information objects with rich attribute combinations and consistent behavioral characteristics while maintaining the constraints of the overall population distribution, thus limiting the accuracy of simulated elections.
[0053] (3) Generative AI models are used in a crude manner. Although generative AI models have powerful data generation capabilities, existing systems mostly use them only as text generation tools. They lack a mechanism to automatically construct prompts based on structured population attribute distribution and support bias data, which makes it difficult for the model output to match the real population structure. The generated results are statistically uncontrollable, reducing their credibility in election prediction scenarios.
[0054] (4) The simulation voting and prediction process is disconnected from data generation. Existing systems often separate virtual sample generation and voting simulation, lacking a unified data structure and a unified processing pipeline. It is difficult to complete the end-to-end processing of "data cleaning - sample generation - simulation voting - result output" in the same system, resulting in low computational efficiency, complex system integration, and failure to fully utilize the parallel processing capabilities of modern computing devices.
[0055] (5) The output of results for various information delivery devices lacks intelligence. Traditional systems usually manually write format conversion logic for different media, such as outputting long texts for newspapers, generating brief manuscripts for radio, and generating structured data for online platforms. This manual rule-based approach is difficult to adapt quickly to changes in media needs, and it is also difficult to automatically generate expression forms suitable for different dissemination channels based on the content of the predicted results, which increases the system operation and maintenance costs.
[0056] (6) Lack of unified optimization of computer resource utilization and algorithm pipeline. In the existing technology, few systems integrate database preprocessing, generative artificial intelligence model calling, statistical / machine learning simulation and media adaptation output into a whole architecture optimized for computing resources. As a result, there is redundancy in data transmission, model call count, and storage read and write, making it difficult to fully improve throughput and response speed.
[0057] Therefore, a new computer implementation scheme is needed. Under a unified data processing architecture, this scheme can automatically preprocess structured data, automatically generate constraint prompts based on population attribute distribution to call generative artificial intelligence models, generate individual information objects that conform to statistical distribution, and complete simulated voting and multimedia formatted output within the same system. This would substantially improve the processing efficiency, scalability, and output diversity of election prediction systems at the computer level.
[0058] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0059] In this invention, the server includes: a device for acquiring survey information and population attribute information recorded according to election units from an information storage device, performing preprocessing on the survey information and population attribute information including missing value completion and outlier removal to generate preprocessed population attribute distribution information and preprocessed support tendency information; a device for automatically calculating the occurrence ratio of each population attribute category based on the preprocessed population attribute distribution information and the preprocessed support tendency information, generating a prompt statement containing constraints reflecting the occurrence ratio, and inputting the prompt statement into a generative artificial intelligence model, thereby enabling the generative artificial intelligence model to generate multiple individual information objects including age attribute, gender attribute, occupation attribute, education attribute, and support candidate attribute, with the overall population composition consistent with the occurrence ratio; and a device for using the support candidate attribute and population attribute information contained in the multiple individual information objects. The apparatus comprises: a device for performing simulated voting processing by invoking statistical processing or machine learning processing to generate predicted result information including estimated votes for each candidate and estimated results for elected candidates; a device for converting the predicted result information into tabular data or visualized image data, and updating the predicted result information according to user correction instructions, thereby generating updated predicted result information as data output for information provision; and a device for receiving media type information, generating a prompt statement based on the media type information to instruct a generative artificial intelligence model to perform format conversion processing on the data for information provision, inputting the prompt statement into the generative artificial intelligence model to obtain media-specific data for information provision including at least one of text data for reporting, summary data for broadcasting, or structured data for electronic distribution, and transmitting the media-specific data for information provision to various types of information provision devices via a communication network. This allows for the construction of an integrated processing pipeline within the same computer system, encompassing structured data preprocessing, constraint prompt generation, large-scale individual information object generation based on generative artificial intelligence models, statistical / machine learning simulation, and intelligent format output for different media. This reduces reliance on manual scripts and fixed rules, enhances data quality control capabilities and the statistical controllability of generated results, and reduces data transmission and format conversion overhead between different processing modules. Consequently, at the computer technology level, this significantly improves the automation level, processing efficiency, and adaptability to diverse information provision needs in election prediction processing.
[0060] "Survey information" refers to structured or semi-structured data obtained from questionnaires, interviews, or other forms of surveys conducted on individuals within a specific electoral unit, reflecting the attributes, opinions, and behavioral tendencies of the surveyed individuals.
[0061] "Demographic attribute information" refers to data used to characterize the demographic structure of a specific region or population, including but not limited to statistical attribute information such as age distribution, gender ratio, occupational category, education level, income level, and residential area.
[0062] "Information storage device" refers to hardware or software components used to electronically store the aforementioned survey information, demographic information and datasets derived therefrom, including database systems, file storage systems or other persistent storage media.
[0063] "Preprocessing" refers to the quality improvement and standardization process performed on the original survey information and population attribute information before subsequent analysis or generation processing, including operations such as missing value completion, outlier removal, data type conversion, encoding standardization, and field standardization.
[0064] "Missing value completion processing" refers to the process of filling in missing data items in survey information or population attribute information through statistical inference, rule setting, or default value filling, so as to reduce the impact of missing data on subsequent analysis.
[0065] "Outlier removal processing" refers to the process of identifying, deleting, or marking data records with abnormal numerical ranges, logical contradictions, or non-compliance with business constraints based on pre-set rules or statistical judgment methods.
[0066] "Population attribute distribution information" refers to statistical information representing the overall population structure distribution obtained by summarizing the frequency or ratio of occurrence of each population attribute category (such as different age groups, genders, occupational categories, etc.) based on population attribute information or its preprocessing results.
[0067] "Support bias information" refers to data in survey information or individual information objects used to characterize the degree of an individual's preference for one or more candidates or options, including explicit identification of the supported object or support strength information given in the form of probability or rating.
[0068] "Individual information object" refers to a data structure unit used to simulate or represent a single virtual individual (such as a virtual voter), which includes at least one or more of the following attributes: age, gender, occupation, education, and supporting candidate, and may further include other characteristics such as income and interests.
[0069] "Generative artificial intelligence models" refer to data processing models trained through machine learning that can automatically generate text, structured data, or other forms of output based on input data or prompts, including but not limited to deep learning models used for natural language generation, data completion, or data synthesis.
[0070] "Prompt statements" refer to the instructional or explanatory text that serves as input to a generative artificial intelligence model. These statements describe the generation goals, constraints, output format, and content scope to guide the generative artificial intelligence model in generating data or information that meets expectations.
[0071] "Simulated voting processing" refers to the process of using the supporting candidate attributes and demographic information contained in multiple individual information objects, and using predetermined statistical or machine learning methods to simulate the voting behavior of each individual in order to estimate the votes received by each candidate and their probability of being elected.
[0072] "Statistical processing" refers to the analysis, inference, and calculation of data based on statistical methods, including frequency statistics, proportion calculation, regression analysis, Monte Carlo simulation, and other computational processes used to generate predictive or intermediate results.
[0073] "Machine learning processing" refers to the process of using trained machine learning models (such as classification models, regression models, or ensemble models) to predict, classify, or generate input data, in order to support tasks such as simulated voting, vote estimation, or result correction.
[0074] "Prediction results information" refers to the set of result data obtained through simulated voting or related statistical and machine learning processing, used to characterize the estimated votes, vote share, elected candidates and their related confidence levels, etc.
[0075] "Tabular data" refers to a data representation format that is organized in a row and column structure and is suitable for display and processing in spreadsheet software or database tables, including result information stored in comma-separated value files or structured tables.
[0076] "Visualized image data" refers to digital image data that expresses prediction results through charts, graphs, or other image forms, including bar charts, pie charts, line charts, heat maps, and other graphical representations used to intuitively display data distribution and comparison.
[0077] "Information provision data" refers to structured or semi-structured data that is based on predicted results and has undergone necessary transformation or correction, and is used to transmit and display to external information provision devices. It is suitable for direct reception and utilization by different types of terminals or media systems.
[0078] "Media type information" refers to the identification information used to characterize the type of the target information providing device or target dissemination channel, including but not limited to categories such as newspaper media, radio media, television media, online media, and mobile terminal applications, which are used to guide the selection of output format and content style.
[0079] "Media-specific data" refers to output data optimized for a specific media type after format conversion, structural adjustment, and expression adaptation of information-providing data based on media type information. This includes text data for reporting, summary data for broadcasting, and structured data for electronic distribution.
[0080] "Information providing device" refers to a device or system used to receive and display or disseminate media-provided data to end users, including print publishing systems, broadcasting systems, television broadcasting systems, web publishing servers, mobile application servers, and other information publishing platforms.
[0081] In a preferred embodiment of the invention, the system includes a server, a terminal, and user-operated input / output devices. The server is preferably a data processing device equipped with an operating system (e.g., a Linux-based server operating system) and a database management system (e.g., a relational database management system). The terminal is preferably a personal computer or a smart mobile terminal, and the user interacts with the server through a browser or client program on the terminal. The server further includes a scripting environment (e.g., a Python runtime environment), a data processing library (e.g., the Pandas data analysis library), a statistical and machine learning library (e.g., a numerical computation library, a machine learning library), and a client program for communicating with generative artificial intelligence models.
[0082] The server preferably includes a multi-core central processing unit, main memory, persistent storage media, and a network interface in its hardware architecture. Through the cooperation of the multi-core processor and main memory, the server can internally perform multiple computational tasks in parallel, such as data preprocessing, prompt generation, generative artificial intelligence model request assembly, simulated voting statistics, and result format conversion. The terminal preferably includes a display device, input device, local storage device, and network interface, allowing users to view tabular data and visualized image data on the terminal and input correction instructions for the prediction results.
[0083] In this embodiment of the invention, the server preferably uses a relational database management system (such as a general-purpose relational database system) to store survey information, population attribute information, preprocessed data, individual information objects, and prediction result information. The server uses a structured query language to perform selection, insertion, and update operations on the tables in the database to access the information storage device. This invention, by standardizing data modeling at the database level, separates and stores survey information tables, population attribute information tables, cleaned survey tables, synthetic individual tables, and simulation result tables, allowing the server to independently index and cache data at different stages, thereby reducing disk I / O and improving data retrieval speed.
[0084] During preprocessing, the server uses the Pandas data analysis library to construct DataFrame objects corresponding to the data tables in main memory. It then performs missing value completion and outlier removal on the column data through vectorized operations. For missing value completion, the server can use central tendency statistics (such as the arithmetic mean or median) or calculate conditional averages based on demographic attributes, thereby improving the consistency between the completed values and the true distribution without increasing the complexity of the model. For outlier removal, the server can use range rules (such as setting the age range to 18 to 120 years old) and quantile-based outlier detection (such as using upper and lower quartiles and a 1.5 interquartile range threshold). Within the data analysis library, Boolean mask operations are used to filter outlier records all at once, reducing the performance loss caused by checking each record individually.
[0085] After preprocessing, the server generates population attribute distribution information based on the demographic attribute information. The server performs grouping, counting, and normalization operations on fields such as age, gender, occupation, and education to obtain the occurrence ratio of each population attribute category. The server stores these ratios as structured vectors or key-value pair mappings, which can then be directly referenced when generating subsequent prompt statements. By retaining this distribution representation in server memory, the server can avoid repeatedly scanning the original table, thereby reducing database access load.
[0086] To generate prompts for invoking generative AI models, the server encodes demographic distribution and support / support information into natural language descriptions. It converts the proportion of each age group into text fragments, such as "approximately 40% are 18 to 30 years old" and "approximately 35% are 30 to 50 years old," and the occupational distribution into statements like "approximately 60% are in the service industry, and approximately 20% are in the manufacturing industry." The server embeds these fragments into predefined prompt templates to form complete prompts with constraints, making it easier for generative AI models to adhere to the target distribution when generating individual information objects.
[0087] In one specific implementation, the server generates the following example prompt statement: "You are a data generation assistant. Please generate information on 20,000 virtual voters for a given constituency based on the following demographic distribution: Age distribution: approximately 40% 18-30 years old, approximately 35% 30-50 years old, and approximately 25% over 50 years old. Occupational distribution: approximately 60% service industry, approximately 20% manufacturing industry, and approximately 20% other industries. For each individual, please generate: age, gender, occupation, education level, monthly income range, and the probability of supporting candidate A / B / C (a decimal between 0 and 1, where the sum of the three is 1). Output structured text, one individual per line, without explanatory notes." In another embodiment, the server can generate prompts tailored to specific age and occupational groups, for example: "Please generate 5,000 voter profiles for a constituency primarily populated by manufacturing workers aged 30 to 40. Requirements: 1. 70% of individuals must be between 30 and 40 years old; 2. At least 50% of individuals must be manufacturing workers or technicians; 3. Each profile should include: age, gender, occupation, education level, monthly income, primary concerns (employment, prices, healthcare, etc.), and the candidate they support (one of A, B, or C); 4. Output in a text format suitable for subsequent analysis." When invoking a generative AI model, the server preferably sends a request containing prompts and parameters to a remote inference service via a network interface. The remote generative AI model can be implemented as a multi-layered transformer neural network model, including an input embedding layer, multiple self-attention encoding / decoding layers, and an output label prediction layer. During training, the model minimizes the next label prediction error or mask reconstruction error on a large-scale corpus using unsupervised or hybrid supervised methods, employing a cross-entropy loss function and updating parameters through gradient descent and adaptive learning rate optimization algorithms. The server does not perform training locally but uses the trained model for inference. During inference, the server controls model diversity and stability by specifying temperature parameters, maximum output length, and sampling strategies, ensuring that the output is both constrained and preserves individual differences.
[0088] The server internally parses the text results returned by the generative AI model. Because the server pre-defines the output field order and delimitation rules in the prompt statements, it can use positional partitioning or matching based on key field names to parse each line into individual information objects. The server maps the parsed fields to a unified data structure and assigns a unique identifier and election unit identifier to each individual information object. The server stores these individual information objects in an individual information object table and creates an index for this table to facilitate efficient subsequent queries and statistical aggregations. Through this structured parsing and index optimization, the server can maintain high query performance even with a large number of generated samples.
[0089] When performing simulated voting, the server uses statistical or machine learning processing to calculate individual information objects. The server can employ one of two modes or a combination thereof: one is a statistical mode based on expected probability, directly summing the support probabilities for candidates in each individual information object; the other is a Monte Carlo mode based on random sampling, using a pseudo-random number generator to simulate candidate selection for each individual according to the support probabilities. When implementing the Monte Carlo mode, the server can use a fixed-seed random number generation mechanism, ensuring reproducible results in repeated simulations with the same input data, facilitating debugging and model comparison.
[0090] In another implementation, the server can incorporate a machine learning model to correct voting behavior. The server can train a supervised learning model (e.g., logistic regression or gradient boosting decision tree model) based on historical election data. The model's input features include age, gender, occupation, education, income, and historical voting frequency, and its output is the probability of voting for a particular candidate. The server applies this model to generated individual data objects to obtain the corrected support probabilities, and then uses the aforementioned statistical or Monte Carlo model to simulate voting. By integrating such a machine learning model on the server side, this invention can improve the accuracy of prediction results by introducing a correction mechanism based on real historical data while maintaining the flexible generation capabilities of generative artificial intelligence models.
[0091] After receiving the voting simulation results, the server generates prediction information, including the estimated vote count, vote percentage, and estimated results for each candidate. It can further calculate confidence intervals or distribution characteristics obtained from multiple simulations. The server utilizes aggregation functions from the data analysis library to quickly group and statistically analyze a large number of individual information objects, and uses vectorized operations to reduce loop overhead. Due to the server's reasonable index design and caching strategy for the data tables, the simulation and statistical phase can be completed in a short time. Even when generating hundreds of thousands or even millions of individual information objects, it can maintain a good response time, thus demonstrating the technical effectiveness of this invention in terms of computational efficiency.
[0092] In this invention, the terminal serves as the interface between the user and the server. The terminal accesses the application programming interface (API) provided by the server via a network to obtain tabular or visualized data. The terminal can use general spreadsheet software to display tabular data or use graphics library components to display visual images such as bar charts and pie charts. Users can view comparisons of vote percentages for different electoral units and candidates on the terminal interface, and can provide correction instructions based on their domain knowledge, such as adjusting the support rate for a candidate in a specific age group or occupational group. The terminal converts these correction instructions into structured parameters and sends them back to the server, which then recalculates the prediction results or updates the information provided.
[0093] During the output phase, the server controls the output format based on media type information. The server can receive media type information, such as "newspaper media," "radio media," and "online media," and based on this, internally generate specific prompts to guide the generative AI model in summarizing, rewriting, or structurally transforming the predicted results. For example, the server can generate the following prompts for newspaper media: "Based on the following election projections, please generate a report for newspaper readers. Describe the projected vote share and strong constituencies for each major candidate in a formal and objective tone, and briefly explain the main factors that may influence the election. Do not include subjective comments." For broadcast media, the server can generate the following prompt: "Based on the following election prediction data, please generate a 30-second broadcast script. The language should be concise and clear, easy to deliver orally, and highlight the current leading candidates and changes in key constituencies." After receiving media provision data output by the generative artificial intelligence model, the server transmits it as information provision data to the corresponding information provision device, such as a newspaper layout system, a broadcast control console, or a web publishing server, via a communication network. By leveraging the language generation capabilities of the generative artificial intelligence model and combining it with structured statistical results, the server can automatically generate media content in various styles and formats, avoiding the burden of writing numerous rules and maintaining templates in traditional systems, thereby improving the system's adaptability to multimedia output needs.
[0094] This invention enables computer systems to perform complex election prediction calculations in a non-traditional, non-manually scripted manner by unifying the entire processing chain within a server, encompassing data preprocessing, constraint prompt generation, large-scale individual information object generation, statistical and machine learning simulation, and media format output. Because population attribute distribution information and support tendency information are explicitly encoded and used for prompt generation, the output of the generative artificial intelligence model is statistically closer to the true population distribution, thus reducing simulation bias. The preprocessing stage employs vectorized data structures and unified rules for automatic execution, reducing human intervention and improving data quality consistency and processing speed. Furthermore, the collaborative operation of simulated voting, correction models, and media output within the same server reduces the overhead of cross-system data transmission and format conversion, achieving an overall improvement in computational efficiency.
[0095] Furthermore, the various embodiments of this invention can be modified according to actual application scenarios. For example, the server can be replaced with elastic computing resources on a cloud computing platform, automatically scaling up to process large-scale individual information objects in parallel across multiple nodes; the generative artificial intelligence model can be replaced with neural network models of different architectures, such as models based on a convolution-attention hybrid structure or a sparse attention mechanism, to further improve the processing capability of long text prompts; the statistical simulation module can select different algorithms, from simple frequency statistics to complex Bayesian inference, based on the required accuracy and computing resources. As long as the server adheres to the core idea of this invention, namely, generating statistically constrained individual information objects based on structured preprocessed data, and completing simulation and multimedia output within the same technical framework, it should be considered to fall within the embodiments of this invention.
[0096] use Figure 11 The processing flow is explained.
[0097] Step 1: Users acquire and process raw data via a terminal. Inputs include raw data from the pre-election survey of electoral units and raw demographic data; output is a structured data file with a standardized format. Users access an online survey platform using a browser on their terminal to download a CSV file containing questionnaire responses; users also access statistical agency websites to download CSV or tabular files containing demographic attributes such as age, gender, and occupation. Users then use spreadsheet software on their terminal to standardize column names from different sources, remove obviously erroneous records, and save the data as a uniformly UTF-8 encoded CSV file, thus providing standardized input for subsequent automated processing.
[0098] Step 2: The terminal imports the organized structured data into the server database. The input is the CSV data file generated in step 1, and the output is a survey information table and a population attribute information table stored in the database. The terminal connects to the relational database management system on the server via a database client or script, executes table creation statements to create data tables for storing survey records and population attribute records, and then calls the database import command or executes the import script to read the CSV file line by line and convert it into table records. During the import process, the terminal converts text to numeric or date types according to field definitions and displays the error line number and reason to the user when import errors occur, thus ensuring that the data is correctly written to the information storage device.
[0099] Step 3: The server reads raw data from the database and performs preprocessing. The input consists of a survey information table and a population attribute information table from the database; the output is preprocessed population attribute distribution information and preprocessed support tendency information. The server sends query statements through a data processing program, loading the survey records and population attribute records of the target election unit into data frame objects in memory. The server performs missing value completion processing on these data frames, such as calculating the median of each attribute or calculating conditional averages grouped by age group, and filling missing items with these statistical values. The server performs outlier removal processing, constructing a Boolean mask to filter outlier records according to preset rules (such as age range, legal gender value set). The server also performs field standardization and encoding unification, mapping occupational descriptions from different tables to standard occupational categories and generating new standardized fields. After completing the above data processing, the server calculates the frequency and ratio of each population attribute category, forming population attribute distribution information; the server also summarizes support tendency information based on support records or rating statistics for candidates in the survey, serving as the basis for subsequent simulations.
[0100] Step 4: The server generates prompts for invoking generative AI models based on the preprocessing results. The input is the demographic distribution and support bias information obtained in step 3, and the output is a text prompt containing distribution constraints. The server converts the occurrence ratios of each age group, occupation category, and gender category into natural language descriptions, such as "approximately 40% are 18 to 30 years old, and approximately 35% are 30 to 50 years old." The server transcribes the overall support ratios for candidates A / B / C into constraints, such as "the overall support probability distribution for candidates A, B, and C should be approximately X:Y:Z." The server embeds these descriptions into predefined prompt templates to form complete text instructions. For example, the server generates text similar to "You are a data generation assistant. Please generate information for 20,000 virtual voters for a constituency based on the following demographic distribution… For each individual, please generate: age, gender, occupation, education level, monthly income range, and support probability for candidates A / B / C (a decimal between 0 and 1, and the sum of the three is 1). Please output structured text as one individual per line, without explanatory notes." The server stores the generated prompt as a string in memory for use in the next step of calling the generative artificial intelligence model.
[0101] Step 5: The server invokes a generative AI model to generate multiple individual information objects. The input consists of the prompts generated in step 4 and the model invocation parameters; the output is a collection of individual information objects in text format. The server sends a request to the remote generative AI model service via a network interface, including the prompts and control parameters (such as output size, temperature parameters, maximum length, etc.) as part of the request body. In the remote inference environment, the generative AI model progressively generates lines of text containing fields such as age, gender, occupation, education level, and supporting candidate attributes based on the prompts. The server receives the returned long text, splits it line by line, and treats each line as an individual candidate record. In this step, the server primarily handles data transmission and result reception; it does not modify the model's internal structure but records information such as invocation time and token usage for resource management and subsequent optimization.
[0102] Step 6: The server parses the generated results and stores them in a structured format as individual information objects. The input is the text-based list of individuals obtained in step 5, and the output is an individual information object table in the database. The server uses a pre-defined output format (e.g., field order, delimiters, or field labels) to parse each line of text, mapping substrings to fields such as age, gender, occupation, education level, monthly income range, and the probability of support for candidate A / B / C. The server generates a unique identifier for each record and attaches an election unit tag. The server performs a lightweight validation on the parsed records, such as checking if the support probability is between 0 and 1 and sums to 1, and if the age is within a reasonable range; rows that do not meet the conditions are discarded or corrected. Subsequently, the server batch-inserts the validated records into the individual information object table and creates indexes on key fields to improve the performance of subsequent queries and aggregations.
[0103] Step 7: The server performs simulated voting based on individual information objects and generates prediction results. The input is the individual information objects written to the database in step 6, and the output is the estimated vote count, vote percentage, and estimated candidate selection result for each candidate. The server loads all individual information objects of the target election unit from the database, constructing a dataset in memory. Based on the supporting candidate attributes or support probabilities recorded in the individuals, the server selects a simulation strategy: in the probability expectation mode, the server sums the column vectors of the support probability values for each candidate and divides by the total number of individuals to obtain the vote percentage; in the Monte Carlo mode, the server uses a pseudo-random number generator to randomly select a candidate for each individual based on the support probability, then counts the frequency of all individual voting results to calculate the vote count and vote percentage. The server can repeat the Monte Carlo simulation multiple times, recording the distribution of results in each round, calculating the mean and variance to obtain the prediction confidence interval. Finally, the server combines the estimated vote count, vote percentage, and the candidate with the highest vote percentage as the estimated candidate selection result into a prediction result information structure and writes it into the prediction result table.
[0104] Step 8: The terminal retrieves prediction results from the server and displays them to the user. The input is the prediction results provided by the server, and the output is tables and graphical visualizations displayed on the terminal interface. The terminal accesses the server's interface via the network, submitting query conditions (such as electoral unit numbers or candidate lists) and receiving structured data of the prediction results from the server. The terminal loads this data into its local display component, mapping candidate names to their corresponding vote percentages and estimated vote values as table rows. The terminal uses a chart library to create bar charts or pie charts, graphically displaying the differences in vote percentages among different candidates. The terminal can also display comparative results from different batches of simulations based on timestamps, allowing users to intuitively understand the changing trends in the predictions.
[0105] Step 9: Users analyze and correct prediction results on the terminal and send correction instructions back to the server. Inputs are the prediction result tables and graphs displayed on the terminal, and outputs are a set of parameters representing the correction content. Users decide whether to adjust a candidate's vote share among a specific group based on external intelligence or professional judgment; users can modify a candidate's weight or add correction percentages on the terminal interface, for example, increasing candidate A's overall predicted vote share among the youth group by a certain percentage. The terminal converts user operations into structured correction parameters (such as candidate identifier, target group conditions, and adjustment coefficients) and sends them to the server via the network as input for subsequent updates to prediction results.
[0106] Step 10: The server updates the prediction results and generates information provision data based on user correction instructions. The inputs are the original prediction results from step 7 and the correction parameters from step 9. The outputs are the updated prediction results and information provision data suitable for external use. The server reads the original prediction results, applies correction coefficients to affected candidates and subsets of the population, and recalculates the overall vote share and estimated results for elected candidates. During the correction process, the server maintains consistency in the total vote count or percentage, ensuring that the sum of the vote shares for each candidate is 1 after adjustment through normalization. Subsequently, the server converts the updated results into a standard output format, including tabular data for spreadsheet import and structured data objects for front-end display, and marks it as information provision data storage or caching, ready for distribution to different media devices.
[0107] Step 11: The server generates media-specific prompts based on media type information and calls a generative AI model to generate media-specific provision data. The input consists of the information provision data from step 10 and externally provided media type information; the output is media-specific provision data such as text reports, broadcast summaries, or structured distribution data. The server receives media type information (e.g., newspapers, radio, online), selects the corresponding prompt template according to the expression needs of different media, embeds key indicators from the information provision data (such as each candidate's vote share, lead margin, and changes in key constituencies) into the template, and generates media-specific prompts for the generative AI model. For example, the server generates text such as, "Please generate a report text for newspaper readers based on the following election prediction data. Please describe the expected vote share and advantageous constituencies of each major candidate in a formal and objective tone, and briefly explain the main factors that may affect the election. Do not include subjective comments." The server sends this prompt along with the information provision data to the generative AI model service, requesting the generation of text or structured descriptions that conform to the media style. The server receives the media-specific content returned by the model, performs basic format checks, caches it as media-specific provision data, and prepares it for distribution.
[0108] Step 12: The server sends media provision data to the information providing device via a communication network. The input consists of the media provision data generated in step 11 and the address or interface information of the target information providing device. The output is content data that can be directly used in the external media system. Based on the interface specifications of different information providing devices, the server selects a transmission method such as HTTP interface, file transfer, or message queue, packages the media provision data, and sends it to the newspaper editing system, broadcast control system, or network publishing server. During transmission, the server records a transmission log, including the target address, transmission time, and data summary, for subsequent tracking and troubleshooting. After receiving the media provision data, the information providing device can directly use it for typesetting, broadcasting, or online publishing, thus completing the technical closed loop from data generation to real-world information dissemination.
[0109] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0110] In existing technologies, in scenarios such as election prediction and advertising placement, target audience profiles are typically constructed manually by analysts based on limited questionnaire results or historical statistical data, and then effect prediction and content design are carried out on this basis. This approach has the following technical problems: First, the data processing chain is fragmented. Cleaning of raw survey data, demographic distribution modeling, virtual population generation, and subsequent predictive analysis are often completed through multiple independent tools and manual operations, lacking unified procedural control. This leads to redundant calculations, inconsistencies between multi-source data, and untraceable processing, hindering stable and scalable execution on computer systems. Second, the granularity of target audience modeling is limited. Traditional methods often use simple grouping or coarse labeling, making it difficult to generate a large number of fine-grained virtual individual data within the machine. This fails to fully utilize the computing power of devices to finely model complex audience structures and behavioral tendencies, thus limiting the accuracy of subsequent simulation predictions. Third, the interaction between humans and generative artificial intelligence models lacks structured control. Existing systems often rely on users directly inputting prompts into generative AI models using free text. These prompts are not tightly coupled with underlying demographic information, virtual population data, and behavioral characteristics. This results in insufficient matching between the generated content and the real target population, and makes it difficult to standardize and repeat the same generation process in a computer. Fourth, there is a lack of evaluation and feedback mechanisms for the generated results. For the various content schemes output by generative AI models, the evaluation typically relies on subjective human judgment, lacking a systematic evaluation and recording based on virtual populations and objective indicators. The computer cannot use user interaction results to iteratively update the prompt construction logic and evaluation metrics, making continuous performance optimization for specific application scenarios difficult.
[0111] Therefore, there is a need for an integrated processing mechanism within a computer system that can start from raw survey and demographic information, automatically clean the data, construct population composition and distribution, generate and cluster large-scale virtual individuals, generate structured prompts, call generative artificial intelligence models, and evaluate and optimize results. This mechanism would improve the automation, repeatability, and prediction accuracy of election prediction and advertising simulation processes on computing devices, and enable controllability and evolvability of the input and output behavior of generative artificial intelligence models. In this way, the performance of applications based on generative artificial intelligence models would be improved from the perspective of "computer technology".
[0112] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0113] In this invention, the server includes a computing device for performing the following processes in an information processing device. The computing device is implemented via program instructions: acquiring survey information and population composition information corresponding to multiple geographic units, and performing missing value completion, outlier removal, classification value normalization, and summarization processing on the survey information based on the population composition information to construct population composition distribution data corresponding to each geographic unit within the computer; performing random number generation and weighted sampling operations on a computing device based on the population composition distribution data to automatically generate a large number of virtual individual data, and attaching attribute information and behavioral tendency information to each virtual individual data; simultaneously, performing cluster learning after vectorizing the virtual individual data to divide the virtual individual data into multiple groups and assigning identification information and statistical features to each group; and, based on target conditions from the terminal device and target information related to information presentation, selecting from the virtual individual data... The system extracts statistical characteristics of the population that meet the target conditions from the body data and its clusters. It automatically generates prompts describing these statistical characteristics in natural language on the server, embeds these prompts into prompt statements, and sends these prompt statements as structured input to the generative artificial intelligence model. This allows the computer system to obtain generated text data containing content schemes and characteristic descriptions corresponding to the target virtual population. The generated text data is then parsed to extract evaluation indicators related to each content scheme. These evaluation indicators are quantified and formed into a comparable data structure. The comparison results and content schemes are sent to the terminal device for display. Simultaneously, the user's selection results and editing operations on the terminal device are recorded. Based on these records, the generation rules for the prompt statements and the calculation conditions for the evaluation indicators are updated to iteratively optimize the input control and output scheme evaluation logic of the generative artificial intelligence model on the server side. This allows for the formation of a closed-loop processing chain within the computer system, encompassing data cleaning, virtual crowd modeling, generative AI model invocation, and result evaluation feedback, without relying on decentralized manual operations. This automates and enables the construction of input prompts and the selection of output schemes for generative AI models, improving simulation accuracy and operational efficiency in applications such as election prediction and advertising simulation. Furthermore, it enhances the overall performance of computer technology from the perspectives of system architecture and data processing flow.
[0114] "Information processing device" refers to a general term for data processing equipment that includes at least one computing device and a storage device, and is capable of acquiring, processing, storing and outputting input data according to a pre-stored program, such as a server, workstation or other computer equipment.
[0115] "Computing device" refers to the general term for electronic processing components located in an information processing device, used to execute program instructions, perform calculations on data, and control the flow of processes, such as a central processing unit or other processors.
[0116] "Geographical unit" refers to a population partitioning unit formed based on geographical or administrative divisions for statistical and analytical purposes, such as electoral districts, urban areas, or other spatial division units.
[0117] "Survey information" refers to raw data obtained through questionnaires, interviews, or other data collection methods, reflecting the opinions, preferences, behaviors, etc. of individuals or groups.
[0118] "Population composition information" refers to a set of data used to represent the structural characteristics of the population within a specific geographical unit, including demographic attributes such as age, gender, occupation, education level, and income level.
[0119] “Population composition distribution data” refers to the statistical processing results based on population composition information and survey information, which are used to represent the frequency and proportion of different combinations of population attributes in various regional units.
[0120] "Missing value completion" refers to the process of inferring and filling in missing or incomplete data in survey information or population composition information using statistical methods or rules.
[0121] Outlier removal refers to the process of marking, filtering, or replacing data that exceeds a reasonable range or clearly violates statistical laws when data is detected in survey information or population composition information, in order to improve data quality.
[0122] "Category value normalization" refers to the standardization of fields representing category attributes, including name standardization, code standardization, or category mapping, so that data with similar meanings are represented in a consistent form in the system.
[0123] "Summary processing" refers to the process of grouping, counting, summing, averaging, or other statistical aggregations of raw data according to predetermined dimensions, thereby generating higher-level statistical results.
[0124] "Virtual individual data" refers to synthetic data records generated through random sampling or models based on population composition and distribution data, used to simulate the attributes and behavioral characteristics of real individuals. Each record corresponds to a virtual individual.
[0125] "Attribute information" refers to data related to a virtual or real individual that describes their demographic characteristics, such as age, gender, occupation, education level, and place of residence.
[0126] "Behavioral tendency information" refers to data inferred from survey information, rules, or models to represent the preferences and tendencies that virtual individuals may exhibit in voting behavior, consumption behavior, or information reception behavior.
[0127] "Statistical learning processing" refers to the processing of datasets containing multiple features using statistical and machine learning methods, such as pattern analysis, clustering, classification, or regression.
[0128] Clustering refers to the process of dividing multiple data samples into several subsets based on feature similarity using statistical learning methods, so that samples within the same subset have high similarity to each other, while different subsets have large differences.
[0129] "Identification information" refers to identifying data used to distinguish different data objects or different cluster groups, such as cluster numbers, label codes, or other markers.
[0130] "Terminal device" refers to a device used to communicate with an information processing device and provide an interactive interface to a user, such as a personal computer, smartphone, tablet device or other user terminal.
[0131] "Target conditions" refer to the constraint information input by the terminal device used to limit the scope of the target population or the scope of the analysis, such as screening conditions such as geographical scope, age group, gender, and occupation category.
[0132] "Target information" refers to the information set in relation to the purpose of information presentation, advertising, or content generation, such as a description of a goal to increase brand awareness, promote registration conversion, or predict election results.
[0133] "Statistical characteristics" refer to indicators used to characterize the overall features of a group after statistical analysis of its data, including mean, distribution ratio, variance, or other statistical measures.
[0134] "Natural language description" refers to the use of human-readable natural language text to describe data, characteristics, or rules, rather than using pure numerical values or program code.
[0135] "Prompt information" refers to the content elements used to construct prompt statements, including target audience characteristics, task requirements, generation style constraints, etc., which are used by generative artificial intelligence models to understand the generation target.
[0136] "Prompt statements" refer to natural language text that is automatically generated by the server according to predetermined rules and underlying data, and is used as input to generative artificial intelligence models to guide their output.
[0137] "Generative artificial intelligence models" refer to artificial intelligence models trained on large-scale data that can automatically generate text, images, or other content based on input prompts.
[0138] "Generated text data" refers to the result data presented in the form of natural language text, which is output by the generative artificial intelligence model after receiving prompts. This includes content plans, explanatory text, and related information.
[0139] "Information presentation content scheme" refers to specific text content or expression schemes generated by generative artificial intelligence models or compiled by systems for display on media, targeting specific target groups or usage scenarios.
[0140] "Evaluation metrics" refer to the metrics used to assess the quality or effectiveness of information presentation content schemes, including metrics such as relevance, attractiveness, expected click-through rate, or conversion rate.
[0141] "Comparable form" refers to a representation method that uses a unified data structure, quantitative indicators, or sorting rules to enable computer comparison of the merits or differences of multiple information presentation schemes.
[0142] "Information delivery medium" refers to the carrier or channel used to transmit information to the public or a specific audience, including print media, radio media, television media, online media, or other communication media.
[0143] "Explanatory text" refers to natural language text used to explain and summarize a group or feature, in order to explain the representative characteristics of the group to the user or subsequent processing modules.
[0144] "Generation rules" refer to a set of predefined logic and parameters used in the server to construct prompt statements or other text structures, including templates, variable mapping relationships, and combination strategies.
[0145] "Calculation conditions for evaluation indicators" refers to the parameters and rules used when calculating evaluation indicators, such as weight settings, threshold settings, data source selection, and calculation methods.
[0146] "Iterative optimization" refers to the process of repeatedly adjusting model inputs, rule parameters, or processing flows by utilizing historical operating results and feedback information in order to gradually improve system performance or output quality.
[0147] In one implementation, the server is deployed as the core node of an information processing device within a data center. It is equipped with a multi-core general-purpose processor and an optional graphics processor, running a server operating system and application runtime environment. The server stores program modules, structured data, and the configuration files for generative artificial intelligence models in its storage device. The terminal, serving as a user interaction device, can be a personal computer, tablet, or portable communication terminal, running browser applications or native applications. Users issue commands to the server through the terminal's graphical interface, including setting conditions, adjusting prompts, and confirming results.
[0148] In a preferred configuration, the server uses a scripting language environment as the execution platform for application logic. It utilizes a data analysis library for data cleaning and statistical processing, a numerical computation library for vectorization and matrix operations, and a machine learning library for clustering and feature statistics. A database management system manages survey information, population composition information, and virtual individual data. The server calls externally provided generative artificial intelligence model services via a network interface. This model performs large-scale neural network inference on external computing resources. The server sends prompts in text format and receives generated text data.
[0149] The server stores the raw survey information in a row-column structure at the data storage layer. Each row corresponds to a real respondent, and each column corresponds to a demographic attribute or survey question answer. The data table includes at least age, gender, occupation, education level, geographic unit identifier, and survey response fields. The server also stores a separate population composition information table within the same database instance, recording the population proportions of each geographic unit, such as age distribution, gender ratio, and occupational category ratio. The server links the survey information table with the population composition information table using key fields, enabling joint analysis based on geographic units in subsequent statistical learning processes.
[0150] In the data preprocessing module, the server performs missing value completion and outlier removal on the survey information using statistical rules and threshold judgment methods. For numerical fields, the server calculates the median and standard deviation within the selected area, marking values higher than the mean by several times the standard deviation or lower than zero as outliers, which are then removed or replaced with the median in subsequent calculations. For categorical fields, the server unifies common syntax differences into standard category codes using a string mapping table, storing them in the feature matrix in integer or one-hot encoding form. This unification process reduces the dimensionality sparsity of the feature space, improving the convergence speed and numerical stability of subsequent clustering algorithms when processing large-scale virtual individual data.
[0151] In the virtual individual generation module, the server converts population distribution data into probability distribution vectors, constructing probability vectors for age groups, genders, and occupational categories for each geographic unit. Based on these probability vectors, the server performs weighted random sampling to generate attribute combinations for each virtual individual. The server forms a virtual individual dataset by recording attribute vectors and geographic unit identifiers for each virtual individual, and stores this dataset using a dense matrix structure, enabling batch matrix operations on millions of virtual individuals in memory. When attaching behavioral tendency information to virtual individuals, the server uses rule functions to map attribute vectors to behavioral tendency scores. For example, it calculates behavioral characteristic values such as political awareness or price sensitivity based on a linear combination of age group, education level, and occupational category, and a threshold function. This rule-based mapping reduces reliance on manually labeled behavioral data, allowing virtual individual data to maintain discriminability in the feature space even in the absence of complete real-world behavioral data.
[0152] In the clustering module, the server vectorizes the virtual individual data and inputs it into the clustering algorithm. In one embodiment, the server uses a centroid clustering algorithm based on iterative optimization. It constructs a feature vector from the attribute and behavioral information of each virtual individual and performs clustering calculations on the feature matrix formed by all virtual individuals. In each iteration, the server repeatedly executes the two steps of "assigning virtual individuals to the nearest cluster center" and "updating cluster centers" until the change in cluster centers is below a preset threshold or the maximum number of iterations is reached. After clustering, the server adds a cluster identifier to each virtual individual and calculates the mean, variance, and behavioral characteristic distribution of the demographic characteristics of each cluster, serving as the basis for subsequent construction of statistical characteristics. By clustering, the server reduces the complexity of describing the target group, compressing the high-dimensional virtual individual space into a few representative groups. This significantly reduces the computational load of traversing data at the individual level when constructing prompt statements, improving overall processing speed and reducing memory access overhead.
[0153] The server receives target conditions and information from the terminal in the prompt statement generation module. Based on the target conditions, the server performs a filtering operation on the virtual individual dataset, extracting a subset of virtual individuals that meet criteria such as region, age group, gender, and occupation. The server calculates statistical characteristics on this subset, including mean age, age group distribution, occupational category ratio, education level ratio, and the mean and distribution range of behavioral tendencies. The server converts these statistical results into natural language descriptions, such as "This group is mainly composed of women aged 30 to 35, most of whom work in marketing-related fields and have a high frequency of social media use." The server further reads the product descriptions or candidate policy summaries and advertising or information presentation targets entered by the user on the terminal, embedding the statistical descriptions and target information into a predefined prompt statement template to generate a complete prompt statement. Through this automatically constructed prompt statement based on statistical characteristics, the server ensures that the input to the generative AI model not only includes natural language task instructions but also structured background information traceable to the distribution of virtual individuals, thereby improving the relevance between the model's output and the target virtual population.
[0154] In a specific example, when a user wants to generate advertising content targeting a demographic of "women in their 30s, college graduates, working in the marketing industry," the server, based on the conditions input by the user, filters individuals with the corresponding attributes from the virtual individual data and calculates their statistical characteristics. The server then constructs the following example of a prompt statement: "You are an advertising planning expert. Your target audience is: women in their 30s, college graduates, working in the marketing industry, living in the city, frequently using social media, and focused on career development and quality of life. Based on these virtual audience characteristics, please generate three product advertising copy in different styles, each no more than 150 words, and explain why each copy is most likely to appeal to this audience." The server sends the aforementioned prompt statement as text input to the generative AI model, which performs multi-layer neural network inference on external computing resources. In one embodiment, the generative AI model can be a deep neural network based on a self-attention mechanism. The model consists of several stacked encoding and decoding layers, each including a multi-head self-attention sublayer, a feedforward network sublayer, residual connections, and normalization operations. The model uses word embedding vectors and positional encodings to represent the input prompt statement, inputting it into the encoding layer to form a contextual semantic representation. During decoding, attention weights are used to weight the context, generating output text word by word or character by character in an autoregressive manner. During training, the model employs a cross-entropy loss function, updating network parameters using gradient descent based on a large-scale text corpus. During inference, the server controls the diversity and length of the generated content by providing parameters such as temperature and maximum generation length. In this invention, the server does not retrain the model but indirectly controls the model's output behavior by changing the prompt statement structure, content boundaries, and constraints, thus achieving a technical utilization of the model's capabilities.
[0155] The server receives the generated text data returned by the generative artificial intelligence model in the result parsing module and parses the data into multiple content scheme units using string segmentation and regular expression matching methods. Based on predefined text tags or serial numbers, the server identifies each advertising copy or information presentation scheme and parses any "attractive point descriptions" or "target audience descriptions" that the model may attach to the output as supplementary fields. When quantitative evaluation is required, the server can construct a scoring function based on virtual individual behavioral tendency information and keyword features of the content scheme. For example, it can match price discount terms, career development terms, and family-related terms in the content scheme with the corresponding behavioral tendency features of the virtual audience and calculate a matching score. Through this matching scoring function based on the virtual individual feature space, the server ensures that the evaluation of the generated results does not rely on immediate external feedback but obtains repeatable quantitative indicators internally, facilitating the system's automatic comparison and ranking of multiple content schemes.
[0156] In one implementation, the terminal is a user interface implemented using web-based technology. The terminal displays multiple candidate content options via graphical components and shows server-calculated evaluation metrics, such as relative attractiveness scores or expected conversion potential levels, in tabular or chart form. The terminal allows users to select certain options as "candidate options" or "preferred options" and provides a text editing area for users to modify content details. After the user completes their selection and editing, the terminal sends the user's selection results and edited content back to the server in a structured format. The server then associates and stores this user feedback with the original generated text data and corresponding prompts.
[0157] Users can adjust the target conditions and constraints in the prompts multiple times on the terminal, such as changing "the tone should be light and humorous" to "the tone should be professional and reliable," or adding instructions like "please provide two versions, one emphasizing emotional appeal and the other emphasizing rational appeal." The server records these changes in the prompt generation rule model and, when automatically generating prompts later, refers to user history preferences to adjust template selection, wording style, and explanatory structure. The server can also analyze the relationship between different prompt patterns and evaluation metrics; for example, if a certain type of prompt structure has a higher average score among a specific group of people, then that type of structure will be prioritized in the prompt generation stage. This data-driven iterative optimization process transforms the server's input control of the generative artificial intelligence model from static rules to a dynamic and learnable prompt generation mechanism, thereby achieving technical optimization of the prompt construction strategy at the system level, resulting in improved output quality and processing efficiency.
[0158] The server optimizes overall computing resources by integrating data cleaning, virtual individual generation, cluster analysis, prompt generation, generative model invocation, and result evaluation feedback into a closed-loop data processing chain. By modeling the target population at the virtual individual level, the server reduces reliance on large-scale real-world questionnaires and long-term online experiments, significantly lowering data collection and network communication load during the experimental simulation phase. In the candidate text selection stage, the server only displays a limited number of filtered content options to the user. Users do not need to repeatedly invoke the generative AI model for every minor change, thus reducing the number of external interface calls and indirectly lowering network latency and external computing resource consumption.
[0159] In this invention, the server employs a virtual individual clustering and statistical characteristic-driven prompt statement construction method. Compared to the traditional method where prompt statements are subjectively written by human operators, this introduces a non-conventional processing flow based on a high-dimensional feature space and clustering structure. Instead of simply relaying the user's goal, the server calculates statistical characteristic summaries to compress the distribution information of virtual individuals in the feature space into natural language descriptions. This allows the prompt statements to semantically correspond directly to the underlying data distribution, achieving a structured correspondence between "data—prompt—generated result." This structured correspondence provides a foundation for subsequent error analysis and optimization. Users and the system can trace deviations in the generated results back to the corresponding statistical characteristics and clustering features, adjusting the virtual individual generation rules or prompt construction logic, thereby technically closing the feedback loop of the generative model application.
[0160] In this invention system, users assume only high-level decision-making and minimal editing roles. The server centrally handles most data processing, feature extraction, and prompt construction, moving beyond simply proceduralizing manual workflows. Instead, it introduces a collaborative optimization mechanism based on virtual individual modeling and generative artificial intelligence models within the computer. This mechanism, by designing scoring functions, clustering rules, and prompt statement structures within the feature space, enables the system to improve prediction accuracy and generation quality in a non-linear manner. This enhances the performance of computer technology itself in applications such as election prediction and advertising simulation, resulting in technical improvements including reduced prediction errors, increased matching between generated content and target audience, optimized computational resource utilization, and fewer user interaction rounds.
[0161] use Figure 12 The processing flow is explained.
[0162] Step 1: The server receives and stores the raw data.
[0163] The server takes survey information and population composition information uploaded from external data sources or management terminals as input. The input format can be CSV files, JSON data, or database records. The server uses a data interface module to read this input data and performs preliminary validation of field names, data types, and encoding formats. The server checks each record to ensure it contains required fields such as geographic unit identifier, age, gender, occupation, and education level. Records that pass validation are written to the raw data table in the relational database, while invalid records are written to an exception log table. The output consists of a raw survey information table and a population composition information table stored in the database, providing the basic data source for subsequent data processing.
[0164] Step 2: The server performs cleaning and standardization processing on the raw data.
[0165] The server reads the raw survey information stored in step 1 from the database as input and uses a data analysis library to build a data table structure in memory. The server performs range checks and outlier detection on numerical fields (such as age and income), replaces obviously abnormal values with the median, and fills in missing values using the statistical median or mode of the same population group within the same geographical unit. For categorical fields (such as occupation and education level), the server uses a predefined mapping table to standardize different syntaxes, converting text categories to standard category codes. The server groups and aggregates all records by geographical unit, calculates the number and proportion of each region in different age groups, genders, and occupational categories, and generates a population composition distribution data table in a unified format. The output is a cleaned survey information table and a distribution data table containing the population proportion of each region, providing probability distribution input for virtual individual generation.
[0166] Step 3: The server generates virtual individual data based on population distribution.
[0167] The server takes the population distribution data table output from step 2 as input and reads statistical values such as the age group ratio, gender ratio, and occupational category ratio for each geographic unit. Based on the preset total number of virtual individuals, the server uses random number generation and weighted sampling algorithms to sample each attribute dimension according to its corresponding probability distribution, generating a set of attribute value combinations for each virtual individual. During the generation process, the server assigns a unique identifier to each virtual individual and records its geographic unit, age, gender, occupation, education level, and other attribute information, forming a virtual individual dataset. The server writes this virtual individual dataset into a database or document storage system in tabular or document format. The output is a virtual individual data table containing a large number of virtual individual records, providing input for subsequent behavioral tendency inference and cluster analysis.
[0168] Step 4: The server attaches behavioral tendency information to virtual individuals and quantifies it.
[0169] The server takes the virtual individual data table generated in step 3 as input and reads the attribute information of each virtual individual. Based on a pre-defined rule function or simple model (e.g., a weighted combination of age, occupation, and education level), the server calculates multiple behavioral tendency scores for each virtual individual, such as willingness to participate in elections, sensitivity to political advertising, attention to price promotions, and acceptance of online information. The server appends these behavioral tendency scores to the virtual individual records, forming an extended feature set. Subsequently, the server encodes the attribute and behavioral tendency information of each virtual individual into numerical feature vectors, for example, using one-hot encoding and normalization to combine the attribute and behavioral features into fixed-length floating-point vectors. The output is a virtual individual feature matrix containing feature vectors and a virtual individual data table with accompanying behavioral tendency fields, providing input for clustering and statistical analysis.
[0170] Step 5: The server performs clustering on the characteristics of virtual individuals and generates group labels.
[0171] The server takes the virtual individual feature matrix output from step 4 as input and calls the clustering algorithm module to perform clustering operations. The server pre-sets the number of clusters and a convergence threshold, repeatedly executing iterative operations of "assigning each virtual individual to the nearest cluster center" and "updating the cluster center based on the current cluster members" until the center change is below the threshold or the maximum number of iterations is reached. The server statistically analyzes the clustering results, calculating the average age, gender ratio, occupational distribution, education level distribution, and behavioral tendency mean for each cluster, and generating a brief descriptive text as the group label for that cluster. The server writes the cluster number and group label into the virtual individual data table, attaching a group identifier to each virtual individual. The output consists of virtual individual data containing cluster identifiers and group statistical characteristics, as well as a separately stored group label table for subsequent filtering and prompt statement construction.
[0172] Step 6: The terminal receives the target conditions and target information set by the user and sends them to the server.
[0173] The terminal takes the conditions entered or selected by the user in the interface as input, including target geographic area, age range, gender, occupation category, education level, and other target conditions, as well as target information presented in advertisements or information presentation (e.g., increasing awareness, promoting registration, enhancing favorability) and product or candidate description text. The terminal organizes these inputs into a structured request message and sends it to the server through the network interface. The terminal does not perform complex calculations at this step; its output is a request data packet containing the target conditions and target information, which serves as input for the server to subsequently filter virtual individuals and generate prompt statements.
[0174] Step 7: The server filters virtual individuals based on target criteria and calculates the statistical characteristics of the target population.
[0175] The server takes the request data packet received in step 6 and the virtual individual data output in step 5 as input, and filters a subset of virtual individuals that meet the user's target conditions from the virtual individual data table. The server performs statistical calculations on this subset to generate statistical characteristics of the target population, including age distribution, gender ratio, occupational category ratio, education level ratio, and the mean and distribution range of behavioral tendency scores. When necessary, the server uses clustering results to statistically analyze the distribution of the target population within each group, determining which typical groups the target population focuses on. The server converts these statistical results into a structured statistical summary and generates intermediate data for natural language description. The output is a summary of the statistical characteristics of the target population, providing input to the prompt statement construction module.
[0176] Step 8: The server constructs the prompt message and generates the prompt statement.
[0177] The server takes the target audience statistical characteristic summary generated in step 7 and the product description or candidate policy summary and target information input by the user in step 6 as input. The server calls the prompt generation module to convert the statistical characteristic summary into natural language description fragments, the target information into task descriptions, and embed the user-provided explanatory text into contextual descriptions. The server combines these fragments according to a predefined template to generate a complete prompt statement for generative artificial intelligence models, which includes elements such as target audience characteristics, generation task requirements, style constraints, and output quantity. In one example, the server generates the following prompt statement as output: "You are a seasoned advertising planning expert. Your target virtual audience characteristics are as follows: a woman in her 30s, a university graduate, working in marketing, residing in a city, frequently using social media, and focused on career development and quality of life. Product description: an online marketing tool primarily designed to help professionals enhance their personal brand exposure and customer acquisition efficiency. Advertising objective: to increase the target audience's interest in the product and their willingness to try it while maintaining a professional image. Based on the above virtual audience characteristics and advertising objective, please generate three different styles of Chinese advertising copy, each no more than 150 characters, and explain why each copy is most likely to attract this audience." The server outputs the prompt as text data and passes it to the generative artificial intelligence model invocation module.
[0178] Step 9: A server-side approach that uses a generative artificial intelligence model to generate text content.
[0179] The server takes the prompt from step 8 as input and constructs a request to connect to the generative AI model service, including control parameters such as model name, generation length, temperature parameters, and the number of returned candidates. The server sends the request to the remote generative AI model via a network interface. The model performs neural network inference on external computing resources, using its internal multi-layered self-attention structure and language modeling parameters to encode the prompt and gradually generate the output text. After receiving the data returned by the model, the server performs a completeness check on the output text, parsing out multiple advertising copy schemes and accompanying explanations. The server then provides these parsed text schemes as output to the subsequent evaluation and display modules.
[0180] Step 10: The server parses and quantifies the generated text scheme.
[0181] The server takes the multiple generated text schemes obtained in step 9 and the previously calculated target audience behavioral tendency information as input. First, the server splits the generated text according to a preset format, identifying the main body of each text and its corresponding explanation. Then, using keyword matching and feature extraction methods, the server extracts semantic features related to price discounts, career development, quality of life, and family care from the text, and matches these features with the target audience behavioral tendency feature vector to obtain a matching score for each text across different psychological appeal dimensions. The server combines these matching scores into multi-dimensional evaluation indicators, such as overall relevance score, emotional resonance score, and rational persuasion score, and normalizes and ranks all schemes. The output is a list of text schemes with added quantitative evaluation indicators for display on the terminal and selection by the user.
[0182] Step 11: The terminal displays the generated solution and evaluation results, and receives user feedback.
[0183] The terminal takes the list of text proposals with evaluation metrics sent by the server in step 10 as input and displays each proposal and its corresponding quantitative score in a list or card format on the interface. The terminal can use chart components to display a comparison of scores across multiple dimensions for different proposals, such as creating bar charts or radar charts. The terminal allows users to select one or more proposals on the interface and mark them as "candidates" or "adopted," and provides a text editing area for users to modify the content. After the user completes the operation, the terminal packages the user's selection, the edited text, and the corresponding identification information as feedback data and sends it to the server. The output is structured data containing user feedback information, used by the server to update the prompt generation rules and optimize the evaluation logic.
[0184] Step 12: The server records user feedback and updates the prompt message generation and evaluation rules.
[0185] The server takes the user feedback data sent by the terminal in step 11, along with the corresponding original generated text schemes, prompt statements, and target audience statistical characteristics as input. The server stores this information in a feedback record table, marking which prompt structures and content features were adopted by the user, and which were significantly modified or discarded. Based on accumulated feedback, the server updates the prompt statement generation rules, such as increasing the emphasis on certain statistical characteristics, adjusting the description order, or modifying commonly used wording. Simultaneously, the server adjusts the weight parameters of the evaluation indicators, increasing the weight of feature dimensions highly correlated with the user's adopted scheme in the overall evaluation. Through this iterative adjustment of internal rules and parameters, the server ensures that subsequent generation and evaluation processes better align with historical selection patterns and target scenario requirements. The output consists of updated prompt statement generation rules, evaluation weight configurations, and expanded historical feedback data. These outputs serve as input for the next round of generation and optimization, thus forming a continuous improvement cycle within the system.
[0186] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0187] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0188] Traditional election prediction techniques typically employ methods such as conducting questionnaires with real voters, expanding the sample size for inference, and relying on simple statistical models or fixed rules for prediction. This approach suffers from the following technical problems: First, at the computer implementation level, survey data and demographic information are often stored separately in unstructured or loosely structured forms, lacking a unified data representation and processing flow. This necessitates multiple manual preprocessing and rule configurations on the server side during prediction, making it difficult to perform large-scale voter behavior simulations in a timely and efficient manner, resulting in low computational resource utilization and poor system scalability. Second, existing systems struggle to utilize generative artificial intelligence models for fine-grained modeling of voter characteristics. They typically only perform coarse clustering or hierarchical statistics based on predetermined attribute fields, failing to automatically generate virtual voter objects with rich attribute combinations within the computer, thus limiting the diversity of simulation scenarios and prediction accuracy. Third, the generation, analysis, and external delivery of election prediction results often rely on manual writing of explanatory text and chart creation based on statistical tables. This process is not only time-consuming but also lacks a unified automated generation mechanism on the server side, making it impossible to dynamically adapt to the output format requirements of different media or information service platforms and failing to meet the demands for real-time and personalized information delivery. Fourth, generative AI models in existing systems are usually only used for single text generation tasks and are not tightly integrated with demographic data, candidate policy information and voting simulation logic. The lack of a computer control scheme for the overall process makes model calls scattered and redundant, increasing system complexity and maintenance costs.
[0189] Therefore, it is necessary to provide a new computer implementation scheme to manage personality objects and candidate policy information in a unified data structure in the server, and to achieve (1) automatic batch generation of personality objects, (2) simulation of voting behavior and prediction of election results based on personality objects, and (3) automatic generation of structured charts and natural language explanatory texts for different information dissemination platforms through the collaborative work of generative artificial intelligence models, databases, simulation engines and statistical processing modules. This will improve the automation, processing efficiency and prediction accuracy of the election prediction system at the computational level, and further improve the operation mode and resource utilization mode of computer technology itself.
[0190] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0191] In this invention, the server includes a unit for inputting prompt statements to a generative artificial intelligence model based on survey results of regional units and group attribute information, instructing the generation of personality objects representing virtual characters; a unit for comparing the generated personality object attribute information with candidate subject policy information, thereby simulating voting behavior for each personality object and calculating the probability of a candidate subject being elected based on the simulated voting results; a data management unit for setting a data structure on a recording medium device for storing personality objects and candidate subject policy information and performing registration, updating, and acquisition according to the data structure; a statistical processing unit for summarizing and statistically processing the voting result data obtained from the selection behavior simulation to calculate the number of votes, the vote rate, and the election prediction index; a unit for generating chart data and explanatory text data based on the statistical processing results and inputting prompt statements to the generative artificial intelligence model to instruct the output of the data in a form suitable for information dissemination platforms; and a unit for controlling the generative artificial intelligence model and inputting various prompt statements in stages for specifying the processing content and output form of personality object generation, selection behavior simulation, and automatic generation of explanatory text based on statistical results. This enables the formation of an integrated automated processing flow within the server, encompassing data acquisition, personality object generation, behavioral simulation, result analysis, and multi-style output. It achieves collaborative control of generative artificial intelligence models, databases, and statistical processing modules, reducing manual intervention and repetitive preprocessing operations. This improves the computational efficiency and scalability of large-scale simulation and prediction, and enhances the structure and adaptability of election prediction results output through a unified data structure and automated explanatory text generation mechanism. Ultimately, this improves the overall processing performance and technical effectiveness of the computer system in election prediction tasks.
[0192] "Information processing device" refers to a general term for electronic computing devices, including processors, storage devices, and communication interfaces, used to execute programs, process data, and send and receive data with external devices. It can be a server, terminal device, or computer system, etc.
[0193] "Generative artificial intelligence models" refer to artificial intelligence models built on machine learning algorithms that can generate text, structured data, or other content based on input data, including natural language generation models based on deep learning.
[0194] "Prompt statements" refer to control information input into a generative artificial intelligence model to instruct the model to perform specific processes or generate specific outputs. They can describe the required processing content, constraints, and output format in the form of natural language, structured text, or a combination thereof.
[0195] "Personality object" refers to virtual character instance data represented in a computer-readable data structure. This data includes attribute information such as age, gender, occupation, education level, residential area, and issues of interest, and is used to simulate the characteristics and behavioral tendencies of individuals in real groups.
[0196] “Survey results of regional units” refers to statistical or raw survey data obtained through questionnaires, online surveys or other data collection methods for a specific geographical area or electoral district, which is used to reflect the attitudes and inclinations of the group in that area on a specific issue or candidate.
[0197] "Group attribute information" refers to a set of data used to describe the demographic characteristics of a group or region, including but not limited to demographic information such as age distribution, gender ratio, occupational category, education level, income level, and residential area.
[0198] "Candidate entities" refer to entities that are selected as potential candidates for support during elections or similar selection processes. These entities are typically candidates, organizations, or policy proposals and are stored in the system as data records.
[0199] "Policy information" refers to information related to the candidate subject, such as policy propositions, action plans, program content, and their positions in various fields. It is usually stored in the form of text data and is used to match the concerns and attributes of the personality object in the simulation process.
[0200] "Selection behavior simulation processing" refers to the process of simulating the selection or voting behavior of each personality object in a computer based on the attribute information of the personality object and the policy information of the candidate subjects, according to a predetermined algorithm or rules, and estimating the probability of each candidate subject being supported or elected based on the simulation results.
[0201] "Recording media device" means a storage device used for non-transitory storage of program code, configuration data and business data, including hard disk devices, solid-state storage devices, network storage devices or other computer-readable storage media.
[0202] "Data structure" refers to the logical structure and physical layout used to organize and represent data in a recording medium or main memory, including table structure, field definition, index structure, and relationships, which are used to support data registration, updating, and retrieval.
[0203] "Voting results data" refers to the data set obtained during the simulation of selected behaviors, which represents the simulated selection results of each personality subject for the candidate subject. It usually includes personality subject identifiers, candidate subject identifiers, simulation round identifiers, and related status information.
[0204] "Statistical processing" refers to the process of summarizing, counting, calculating proportions, and analyzing distribution of voting results data using mathematical and statistical operations to obtain the number of votes, the vote rate, and other predictive indicators.
[0205] "Vote count" refers to the total number of times a candidate is supported or selected in the voting results data, which reflects the scale of votes for that candidate in the simulation.
[0206] "Vote rate" refers to the ratio of the number of votes a candidate receives to the total number of valid votes. It is usually expressed as a percentage and is used to reflect the relative support level of the candidate in the simulation results.
[0207] "Election prediction indicators" refer to assessment measures calculated based on the number of votes, vote percentage, and other statistics, used to indicate the likelihood of a candidate being elected in an election. These include probability values, ranking indicators, or comprehensive scores.
[0208] “Chart data” refers to structured data used to generate graphical or tabular representations, including the number of votes, vote percentage, time series data, or grouped statistical data for each candidate subject, which can be converted into bar charts, pie charts, line charts, or tables by visualization components.
[0209] "Explanatory text data" refers to natural language text automatically generated based on statistical results and simulation conditions. It is used to explain election prediction results, analyze support structures, explain trends and uncertainties, and can be used in news reports, analysis reports, or explanatory texts.
[0210] "Information dissemination platform" refers to a system or service used to publish and transmit information to an unspecified majority or a specific group of users, including websites, electronic media systems, content management systems or other information service platforms.
[0211] "Information provider" refers to an entity that uses the data and explanatory text output by this system for republication, reporting, or analysis, including media organizations, information service organizations, research institutions, or other content providers.
[0212] In various embodiments of this invention, the server, terminal, and user work together to achieve automatic generation of personality objects, simulation of selected behaviors, and structured output of election prediction results through a combination of generative artificial intelligence models, database management systems, statistical analysis modules, and visualization modules. This invention is not limited to a specific hardware or software platform, but for ease of understanding, it is described below in conjunction with typical implementations.
[0213] I. Overall System Composition The server can be a rack-mount server or a cloud computing node, comprising a central processing unit, main memory, non-transitory recording media, and a network interface. The server installs an operating system (e.g., a general-purpose server operating system) on the recording media and runs backend service programs, database management software (e.g., a relational database management system such as MySQL), and an inference service environment for generative artificial intelligence models (e.g., model services based on deep learning frameworks).
[0214] The terminal can be a smartphone, tablet computer, or personal computer. It runs a browser application or client application to display the interface, receive user input, and send requests to and receive responses from the server.
[0215] Users input prompts, configuration parameters, and view prediction results through the terminal interface. The server communicates with the terminal via a network interface, using secure communication protocols to transmit structured and text data.
[0216] II. Server-side program and module composition 1. Main software modules in the server The server includes the following logical modules, which can be implemented as processes, services, or libraries: (1) Interface Management Module: The server receives HTTP / HTTPS requests from the terminal through the interface management module, parses the prompts, parameters, and command types in the requests, and distributes them to the personality object generation module, simulation module, or statistical output module.
[0217] (2) Personality Object Generation Module: The server utilizes a personality object generation module to combine user-specified regional unit survey results, group attribute information, and user-input prompts into control information for a generative artificial intelligence model. This module is responsible for constructing model inputs, calling the model inference interface, and parsing the model output into the internal data structure of personality objects.
[0218] (3) Candidate Entity Management Module: The server receives and stores policy information for candidate entities, including policy text, domain labels, and selected area information, through the candidate entity management module. This module interacts with the database to persist the policy information.
[0219] (4) Data Management Module: The server defines data structures on the recording medium through a data management module, such as table structures in a relational database, including personality object tables, candidate subject tables, policy information tables, and simulation voting result tables. The data management module is responsible for performing data registration, updating, and retrieval operations.
[0220] (5) Select the appropriate behavior simulation module: The server utilizes a selected behavior simulation module to match and calculate the attribute information of personality subjects with the policy information of candidate subjects, generating simulated voting results. This module can combine rule-based calculation with vectorized semantic matching algorithms to achieve multi-dimensional preference calculation.
[0221] (6) Statistical Analysis Module: The server uses a statistical analysis module to aggregate and statistically process the simulated voting results, calculate the number of votes, vote rate, and election prediction indicators for each candidate, and provide structured statistical data for subsequent visualization generation.
[0222] (7) Report generation and visualization module: The server utilizes a report generation and visualization module to convert statistical results into charts and explanatory text. The explanatory text can be automatically generated using a generative artificial intelligence model under specified prompt constraints.
[0223] (8) Model control module: The server uses a model control module to control the invocation of generative artificial intelligence models in stages, constructing different types of prompt statements for personality object generation, simulation-assisted judgment, and explanatory text generation, and setting constraints on the model's temperature, maximum output length, and output format.
[0224] 2. Generative Artificial Intelligence Model Structure and Training Methods The server can deploy generative AI models based on the Transformer architecture. The model can employ a multi-layered self-attention network structure, including an encoder and decoder section or a decoder-only structure. During the model training phase, the server pre-trains using a large-scale corpus (including general language corpora and specialized corpora related to demographic information, political policies, and social issues).
[0225] When performing customized fine-tuning for this invention, the server can use a structured labeled dataset containing paired samples of input prompts and target outputs (such as JSON fragments of personality attributes or policy classification labels). The server updates model parameters by minimizing the cross-entropy loss function and uses gradient descent and its variants (such as adaptive learning rate algorithms) for weight updates. During training, the server can perform data augmentation on the samples, such as paraphrasing texts describing the same type of personality traits or mixing multiple languages, to improve the model's robustness to different expressions.
[0226] During the inference phase, when calling the model, the server can set temperature parameters to control output diversity and control generation stability through top-k or top-p sampling methods. By embedding structured placeholders and output specifications in the prompt statements, the server makes the model output more consistent with the expected data structure, reduces the post-processing burden, and thus improves overall computational efficiency.
[0227] III. Personality Object Generation Processing 1. The server generates personality objects using a generative artificial intelligence model. Users input natural language prompts on their terminals, which then send these prompts to the server. Upon receiving the prompts, the server combines them with regional survey results and group attribute information to form model input. For example, the server can encode data from regional surveys such as "support rate for education reform" and "attention level to tax policies" into feature vectors and append explanatory text before the prompts, enabling the model to generate personality objects that correspond to the population distribution of that region.
[0228] In a practical implementation, the server can construct the following sample prompt statement: "Please generate a personality object based on the following criteria: 30 years old, female, primary school teacher, holds a master's degree in education, lives in a city, and is interested in education reform and childcare policies." “Create a virtual voter: a 45-year-old male, a factory worker, with a high school education, married, with two children, and very concerned about job stability and social security.” The server inputs the aforementioned prompts into a generative artificial intelligence model, which outputs a structured description including age, gender, occupation, education level, residential area, and a list of topics of interest. The server then uses a parsing module to map this output into an internal personality object data structure and a data management module to store it in a database.
[0229] 2. The server generates a batch of personality objects. The server can automatically construct multiple prompts based on group attribute information and generate a predetermined number of personality objects by repeatedly calling a generative artificial intelligence model. Compared to manual configuration, the server automatically synthesizes multi-dimensional attribute combinations through the model, generating a broader set of personality objects in the dimensional space, thereby improving the representativeness of the simulation samples.
[0230] During the generation process, the server can control the distribution of personality object attributes to ensure consistency with the regional demographic distribution in terms of age, gender, and education level. The server continuously adjusts prompts or sampling strategies through statistical comparisons, thereby achieving an approximate match to the population distribution during the generation phase. This feedback-based generation process helps improve the accuracy of subsequent simulation predictions.
[0231] IV. Management and Characterization of Candidate Entity Policy Information In the candidate entity management module, the server stores the input candidate entity name, region, policy text, etc., into the database. The server can perform domain classification and vectorization processing on the policy information, for example: The server uses a text embedding model to encode policy text into vectors, which can have hundreds to thousands of dimensions. These vectors represent the intensity distribution of policies across multiple areas such as "education," "economy," "healthcare," and "social security." The server can also perform keyword extraction and sentiment analysis on the policy text to obtain reinforced domain features.
[0232] Through the above feature processing, the server provides a similarity measurement basis in a high-dimensional feature space for subsequent behavioral simulation, which expands the matching between the personality object's focus and policy information from simple keyword comparison to vector similarity calculation, thereby improving the matching accuracy.
[0233] V. Select appropriate behavior simulation and matching algorithms 1. The server performs a comparison calculation between attributes and policies. The server reads a set of personality objects and a set of candidate subjects from the database. For each personality object, the server constructs a feature vector based on its attribute fields (age, occupation, education level, income level, list of issues of interest, etc.) and maps it to the same feature space as the policy vector of the candidate subject.
[0234] The server can employ the following multi-level matching mechanism: (1) Rule weight layer: The server assigns base weights to each domain based on rules learned from expert definitions or historical data. For example, for a personality who "focuses on education reform," the server automatically increases the weight of the education domain and decreases the weight of domains less relevant to that personality.
[0235] (2) Vector similarity layer: The server calculates the cosine similarity or other distance metric between the personality object's attention vector and the policy vectors of each candidate subject. The server weights the domain weights with the similarity values and expresses the comprehensive score of each candidate subject as a linear or non-linear combination.
[0236] (3) Decision-making level: The server sorts the scores of each candidate subject, and the one with the highest score becomes the voting target for that personality type in the simulation. The server then writes this result into the simulation voting result table.
[0237] Through this multi-layered matching method, the server implements a more refined and higher-dimensional matching process inside the computer than traditional simple rules or manual comparison, thereby improving the accuracy of the selected behavior simulation.
[0238] 2. The server utilizes a generative artificial intelligence model to assist in simulation judgment (optional implementation). In some implementations, the server can also construct more complex prompts for specific personality types, which are then used by a generative artificial intelligence model to comprehensively analyze the policy stances of candidate individuals and provide recommendations. For example: "You are now playing the role of a voter behavior analysis system."
[0239] Personality Partner: 35-year-old male, software engineer, bachelor's degree, middle-income, married, with one child, and very concerned about education reform and the quality of his child's education.
[0240] Candidate A's policy emphasizes 'promoting education reform', proposing to increase public school budgets and improve teachers' salaries.
[0241] Candidate B's policy focuses on 'tax cuts' and 'deregulation of business regulations,' with little mention of education reform.
[0242] Question: Which candidate is this personality type most likely to vote for? Please only output the candidate's name. The server merges contextual information and the question into a prompt statement input to the model, which then provides a bias judgment based on its internal language knowledge and semantic relationships. The server can fuse the model's judgment with the results of rule-vector matching, for example, by setting thresholds, using weighted averaging, or conflict resolution strategies, thereby introducing more complex automatic judgment logic into the simulation process.
[0243] In this way, the server not only simulates the manual analysis process, but also achieves complex situation judgments that are difficult for traditional rule systems to cover based on high-dimensional semantic space and contextual understanding capabilities, thereby further improving the robustness and accuracy of prediction results.
[0244] VI. Statistical Analysis and Explanatory Text Generation 1. The server performs statistical calculations. After the selected behavioral simulation is completed, the server uses the statistical analysis module to aggregate and calculate the simulation voting results. The server counts the number of votes for each candidate entity and calculates its vote share and related election prediction indicators. The server can calculate statistical parameters such as confidence intervals and fluctuation ranges based on the results of multiple rounds of simulation to provide a more comprehensive prediction and evaluation.
[0245] During this process, the server performs various aggregation operations and group statistics, such as statistical support distribution by region, age group, and occupation category. These statistical operations are completed directly by the server in the database or in-memory data structure, thereby avoiding the transmission of a large number of intermediate results between terminals, reducing network communication load and improving overall processing speed.
[0246] 2. The server generates chart data and explanatory text. The server constructs structured data suitable for visualization based on statistical results, such as sequences representing the vote ratios of different candidates or time-series change data. The server then passes this structured data to the visualization module, generating charts on the server side or directly returning the data to the terminal for graphical rendering.
[0247] In terms of explanatory text generation, the server constructs prompts for generative artificial intelligence models, requesting the generation of text suitable for news reports or analysis reports. For example: "Based on the following election simulation data, please generate a concise Chinese analysis report introducing which candidate is most likely to be elected and the main reasons:" - Selected area: X Selected area Candidate A: Received 12,500 simulated votes, representing 52% of the total votes. Candidate B: Received 9,800 simulated votes, representing 41% of the total. - Other candidates totaled: 1,700 votes, accounting for 7% Please analyze candidate A's areas of strength (such as education policy), candidate B's main supporters, and the limitations of this prediction. The server combines the statistical data with the aforementioned prompts and inputs them into the model. The model then generates a well-structured, logically coherent natural language explanatory text. The server outputs this text, along with the chart data, to the terminal, which displays it as a report page or exports it as a document.
[0248] VII. Technical Effects and Improvements in Computer Technology Through the above structure and process, the server has achieved the following technical improvements: 1. Improved processing efficiency resulting from the integration of data structure and workflow: The server is designed with a unified data structure for personality objects and policy information in the recording medium, which allows personality generation, simulation and statistical processing to be executed continuously on the same data model. This reduces format conversion and temporary file operations, thereby reducing storage access overhead and improving processing speed.
[0249] 2. High-dimensional feature modeling based on generative artificial intelligence models: When generating personality objects and explanatory text, the server uses a deep neural network model for high-dimensional semantic encoding and generation. Through pre-training and fine-tuning, this model has mastered semantic combination capabilities far exceeding those of manual rules. This allows the matching between personality object attribute combinations and policy text to go beyond explicit fields, and instead perform high-dimensional similarity calculations in vector space, significantly improving the accuracy of simulation predictions.
[0250] 3. The efficiency and controllability of model invocation brought about by phased prompt statements: The server uses a model control module to hierarchically and template the prompt statements, breaking down personality generation, simulation-assisted judgment, and report generation into multiple controllable sub-tasks. Compared to a single, large, and scattered calling method, this structured control reduces the uncertainty of model output, lowers post-processing costs, and improves overall computational efficiency.
[0251] 4. Automatically generate explanatory text to reduce the burden of human-computer interaction: After statistical analysis, the server automatically generates explanatory text adapted for media or information service platforms, achieving automatic conversion from data to descriptive language. Compared to manual writing, the server can generate a large number of reports with uniform structure and consistent content in a short time, reducing the manual processing pressure on the terminal side and enabling the generation of large-scale, multi-version reports that are difficult for humans to complete manually.
[0252] 5. Error reduction resulting from combinations of specific algorithms and rules: The server combines rule weights, vector similarity, and generative analysis results to achieve multi-source information fusion decision-making during simulation. By adjusting weights and thresholds through comparison of the error function with historical data, the server can gradually reduce prediction bias, achieving a lower error rate and more stable prediction performance.
[0253] In summary, by organically integrating generative artificial intelligence models, database structures, and statistical algorithms, the server has implemented a dedicated processing mechanism for election prediction scenarios within the computer. It not only automates traditional manual operations but also improves the technical performance of the computer system at multiple levels, including data representation, feature modeling, decision-making algorithms, and output structuring.
[0254] use Figure 13 The processing flow is explained.
[0255] Step 1: The user enters a prompt message and sets simulation conditions on the terminal. Users open the election prediction interface on the terminal, enter natural language prompts in the text input area to generate personality objects, and select parameters such as the target area, the number of personality objects to be generated, and the number of simulation rounds in the interface controls.
[0256] Input: The prompts entered by the user on the terminal (such as "Please generate a personality object based on the following conditions: 30 years old, female, primary school teacher, holds a master's degree in education, lives in a city, and is concerned about education reform and childcare policies.") and simulation condition parameters (region identifier, number of samples, etc.).
[0257] Output: The terminal generates a request data packet containing prompts and simulation parameters.
[0258] The terminal encapsulates the prompt statement and parameters into a structured request. After verifying locally whether the prompt statement is empty and whether the character length exceeds the limit, the terminal sends the request to the server through the network interface.
[0259] Step 2: The server parses the request and combines regional survey data with population attribute information. After receiving the request from the terminal, the server parses the prompt statement and simulation parameters in the interface management module. Based on the area identifier specified in the request, the server reads the corresponding regional unit survey results and group attribute information from the database.
[0260] Inputs include prompts from the terminal, area identifiers and simulation parameters, as well as survey results and population attribute data from the database.
[0261] Output: A combination of input text (expanded prompts) and internal feature vector representations used to input generative AI models.
[0262] The server encodes group attributes such as age distribution and education level distribution in a region into numerical feature vectors. Simultaneously, it appends explanatory text before and after the original prompt, such as "The percentage of female teachers aged 30-40 in this region is X%, and their support rate for education reform is Y%", thus creating an expanded prompt. Through this data processing, the server maps structured statistical data into natural language conditions that can be understood by generative artificial intelligence models, achieving the conversion from numerical features to textual conditions.
[0263] Step 3: The server inputs prompts into the generative artificial intelligence model and generates personality object data. The server calls the inference interface of a generative artificial intelligence model deployed locally or remotely. The server takes the expanded prompt as the input to the model, and the model encodes and decodes the prompt using a multi-layer self-attention network based on the Transformer structure.
[0264] Input: Expanded prompt text and optional control parameters (such as temperature, maximum generation length).
[0265] Output: Personality object description text or semi-structured output containing fields such as age, gender, occupation, education level, and issues of interest.
[0266] After receiving the model output, the server performs word segmentation, field recognition, and format correction on the output through the parsing module, mapping the natural language description into an internal personality object structure (such as a set of key-value pairs). In this process, the specific data computations performed by the server include: rule-based entity extraction (e.g., identifying the age corresponding to "30 years old"), pattern-matching-based field classification (e.g., identifying the occupation category corresponding to "teacher"), and standardized mapping of vocabulary related to relevant issues (e.g., mapping "education reform" to the predefined label "education_reform").
[0267] Step 4: The server performs batch generation and distribution correction of personality objects. Based on the number of samples specified in the simulation parameters, the server calls the generative artificial intelligence model multiple times in a loop or in parallel, each time passing in a finely tuned prompt (such as adjusting the age range or occupational distribution description) to generate multiple personality objects.
[0268] Input: Initial extended prompt statement, target sample size, and regional demographic distribution information.
[0269] Output: A set of personality objects that meet the target number.
[0270] After each generation, the server performs statistical analysis on the currently generated set of personality objects, comparing their age groups, gender ratios, and educational backgrounds with the population distribution of the target area. The server then adjusts the constraint descriptions in subsequent prompts (e.g., "Please increase the proportion of low- and low-education groups in this area") to achieve iterative data processing based on feedback. This makes the overall set of personality objects approximate the actual population distribution in terms of distribution, thereby reducing sample bias and improving the representativeness of the simulation input through this adaptive algorithm.
[0271] Step 5: The server stores the personality objects in the database. In the data management module, the server writes the generated personality objects one by one into the corresponding tables in the database.
[0272] Input: A collection of internal personality object data structures.
[0273] Output: Personality object records persistently stored in the database and their corresponding primary key identifiers.
[0274] The server performs field mapping, mapping simple fields such as age, gender, and occupation to the table's base columns, and mapping the list of topics of interest to related tables or JSON fields. Before performing an insert operation, the server validates the data type and value range (e.g., whether age is a positive integer, and whether gender is within a predefined enumeration). Through this structured storage, the server can leverage indexes to accelerate retrieval during subsequent simulated queries, thereby improving access efficiency at the data level.
[0275] Step 6: Users enter candidate entity and policy information on the terminal. Users can switch to the candidate management function in the terminal interface and enter the name of the candidate entity, its region, policy statement, and related descriptions through a form.
[0276] Input: Basic information and policy text of the candidate entity manually entered by the user, such as "Candidate A's main policies include: promoting education reform, increasing teachers' salaries, and expanding public school resources." Output: A request data packet for candidate subject information constructed by the terminal.
[0277] After validating the required fields, the terminal sends the candidate subject information to the server via the network.
[0278] Step 7: The server parses and characterizes the policy information of candidate entities. After receiving the candidate entity information, the server stores the information in the database in the candidate entity management module and extracts features from the policy text.
[0279] Input: Candidate entity name, region identifier, policy text, and other relevant metadata.
[0280] Output: Candidate entity records stored in the database, along with corresponding policy feature vectors and domain weight information.
[0281] The server uses a text vectorization model (e.g., an embedding model based on a pre-trained language model) to convert the policy text into a high-dimensional vector. The server further uses keyword matching and a classifier to label sentences or phrases in the text according to categories such as "education," "economy," and "healthcare," and calculates a strength score for each category (e.g., summation based on word frequency weights or aggregation based on attention weights), thus obtaining a multi-dimensional policy feature vector. Through this feature-based data processing, the server transforms the raw natural language policy text into a vector representation that is easy to compute numerically.
[0282] Step 8: The user initiates a request to simulate selected behaviors on the terminal. Users select the target area, set of personality objects, or simulation round on the terminal and click the "Start Simulation" button.
[0283] Input: User-selected region identifier, range of simulated personality objects (e.g., "Use all generated objects" or "Sample 10,000 objects"), whether to enable generative auxiliary judgment, etc.
[0284] Output: The simulation control request sent from the terminal to the server.
[0285] The terminal encodes the above options as parameters and sends them to the server along with the session identifier.
[0286] Step 9: The server selects the personality subjects and candidate subjects to participate in the simulation. Based on the simulation control request, the server queries the database for records of personality objects and candidate subjects under specified regions or conditions.
[0287] Input: region identifier, number of imitation samples, personality object table and candidate subject table in the database.
[0288] Output: A subset of personality objects and a list of candidate subjects for simulation.
[0289] The server can employ random sampling, stratified sampling, or full sampling, limiting the sample size as needed to avoid memory overflow. The server preloads selected personality objects and candidate subjects, caching their attributes and feature vectors in memory to reduce repeated database accesses during simulation and improve overall computational efficiency.
[0290] Step 10: The server performs matching calculations between the personality object and the candidate subject. In the selected behavior simulation module, the server calculates the matching degree between each personality object and each candidate subject.
[0291] Input: Personality object attributes (including concern issue tags and corresponding weights), candidate subject policy feature vectors, and domain basic weight configuration.
[0292] Output: A list of candidate subject matching scores for each personality object, and the final selected "voting object".
[0293] The server first calculates individual weights for each policy domain based on the personality's concerns (e.g., "education_reform" corresponds to a higher weight for the education domain). Then, it weights and sums the domain strength vectors of candidate entities according to these weights to obtain a score for each candidate entity for that personality. Additionally, the server can calculate the cosine similarity between the personality's concern vector and the candidate entity's policy vector, and assign scores and similarity scores using preset coefficient combinations. This process involves numerical vector operations, including dot product, normalization, and weighted summation. The server selects the candidate entity with the highest score for each personality as the voting result.
[0294] Step 11: The server may optionally invoke a generative artificial intelligence model to perform complex situation assessments. With generative assisted judgment enabled, the server constructs detailed contextual prompts for some or all personality objects, embedding personality attributes and candidate subject summaries into them.
[0295] Input: Detailed attributes of the personality object, main policy descriptions of the candidate subjects, and preset analysis question templates.
[0296] Output: Recommended candidate entity names or support tendency descriptions output by the generative artificial intelligence model.
[0297] For example, the server constructs the following prompt statement: "You are now playing the role of a voter behavior analysis system."
[0298] Personality Partner: 35-year-old male, software engineer, bachelor's degree, middle-income, married, with one child, and very concerned about education reform and the quality of his child's education.
[0299] Candidate A's policy emphasizes 'promoting education reform', proposing to increase public school budgets and improve teachers' salaries.
[0300] Candidate B's policy focuses on 'tax cuts' and 'deregulation of business regulations,' with little mention of education reform.
[0301] Question: Which candidate is this personality type most likely to vote for? Please only output the candidate's name. The server inputs the prompt into the generative artificial intelligence model, which then provides a result (such as "Candidate A") based on its internal semantic reasoning. The server then integrates this result with the numerical matching score. For example, if the two match, the confidence level of the result is increased; if they do not match, a final decision is made by setting a confidence threshold or using a weighted voting method, thereby achieving joint inference by rules and the model.
[0302] Step 12: The server records the simulated voting results. At the end of the simulation, the server writes the final votes for each personality object into the simulation voting results table.
[0303] Input: Personality object identifier, candidate subject identifier, simulation round identifier, matching score, and optional confidence level information.
[0304] Output: Persistent simulation voting results records in the database.
[0305] The server inputs data through batch insertion or transaction processing, ensuring data consistency and write performance in high-concurrency simulation scenarios. The server can also save the main decision factors that generated the result for each record (such as the highest domain score or generative judgment result) for subsequent interpretive analysis.
[0306] Step 13: The server performs statistical analysis on the simulation results. In the statistical analysis module, the server performs aggregate queries and statistical calculations on the simulation voting results table.
[0307] Input: Simulation round identifier, voting result record, candidate subject list, and optional grouping conditions (region, age group, etc.).
[0308] Output: Statistical results such as the number of votes, vote percentage, and election prediction indicators for each candidate.
[0309] The server uses aggregation functions to count the votes for each candidate and calculate their vote count. The server then compares the vote count with the total votes to calculate the vote percentage. Based on multi-round simulation results, the server can calculate the average vote percentage and variance, deriving the probability of election or ranking indicators. This process involves statistical data calculations such as group aggregation, proportion calculation, variance and confidence interval estimation.
[0310] Step 14: The server generates chart data and constructs a report, then generates prompt statements. The server generates sequence data for chart plotting based on the statistical results. At the same time, the server constructs descriptive prompts containing statistical summaries so that they can be fed into a generative artificial intelligence model to generate a natural language report.
[0311] Input: Number of votes received by the candidate, vote percentage and related statistical indicators, and a predefined report template.
[0312] Output: Chart data structure and prompts for report generation.
[0313] For example, the server can construct the following prompt statement: "Based on the following election simulation data, please generate a concise Chinese analysis report introducing which candidate is most likely to be elected and the main reasons:" - Selected area: X Selected area Candidate A: Received 12,500 simulated votes, representing 52% of the total votes. Candidate B: Received 9,800 simulated votes, representing 41% of the total. - Other candidates totaled: 1,700 votes, accounting for 7% Please analyze candidate A's areas of strength (such as education policy), candidate B's main supporters, and the limitations of this prediction. The server embeds the statistical data into the template to generate the final prompt statement.
[0314] Step 15: The server generates explanatory text using a generative artificial intelligence model and outputs it to the terminal. The server inputs the constructed prompt statement into the generative artificial intelligence model, requesting the generation of explanatory text that conforms to the style of news reports or analysis reports.
[0315] Input: A prompt statement containing statistical data and explanatory requirements.
[0316] Output: Complete natural language explanatory text, such as election prediction reports or analytical interpretations.
[0317] After receiving the model output, the server performs basic text validation (such as checking whether it contains necessary candidate subject names and whether it roughly meets length constraints), and then sends it to the terminal along with the chart data. Upon receiving the data, the terminal displays the charts and report text on its interface, which users can view, save, or export. Through this automated report generation process, the server achieves the conversion from numerical statistical results to readable explanatory text, reducing the burden of manual writing and, technically, improving the overall system's automation level and output consistency through a unified data flow and model control process.
[0318] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0319] In the existing field of election prediction and information recommendation, the commonly used technical solutions are mainly based on the following: On the one hand, servers mostly use fixed statistical models to process pre-questionnaires and demographic data, which can only build coarse-grained group profiles and are difficult to reflect in a timely manner the impact of segmented groups and emotional states on collective decision-making; on the other hand, when servers predict election results and generate information output (such as news content, propaganda content, or advertising content), they mostly rely on manually set rules or manually written logic code, lacking the ability to adaptively model complex multi-source data (including sentiment data), resulting in limited prediction accuracy and personalization.
[0320] Furthermore, in traditional systems, server-side processing such as "data preprocessing," "behavioral simulation," "result organization," and "information output optimization" are often implemented in a fragmented manner, making end-to-end control difficult through a unified generative artificial intelligence model interface. Specifically, this manifests as follows: (1) Servers typically perform only one-time statistical analysis on pre-questionnaires and demographic information, failing to refine the group structure by generating a large amount of virtual individual data, and thus cannot efficiently simulate large-scale survey behavior within the computer, thereby limiting the simulation accuracy and scalability. (2) When servers perform collective behavior simulations (such as election voting simulations), they generally use static mathematical models or simple rules. They cannot flexibly use generative artificial intelligence models to model complex behavior patterns, nor can they dynamically adjust simulation strategies and objectives through prompt statements. The computer system has poor programmability and scalability. (3) When generating report data or information output content, the server usually only formats the prediction results and lacks the calculation process of "dynamically estimating the responsiveness based on individual attributes and emotional state and comparing different output schemes", which results in the information output not being deeply coupled with the prediction model and the overall system processing flow being fragmented. (4) Existing systems mostly use emotional data only for simple label statistics, without embedding emotional states into virtual individual data in a structured way and participating in the entire process of subsequent election simulation and information output optimization. The computer's ability to model the "emotion-behavior-information response" link is insufficient. (5) Servers typically develop separate processing modules for each type of task, making it difficult to issue integrated instructions such as "data generation, behavior simulation, result analysis, and content optimization" to generative artificial intelligence models through a unified prompt statement interface. As a result, the advantages of generative models in automatic arrangement of complex processes and multi-step reasoning cannot be fully utilized, and the flexibility and maintainability of the system are limited.
[0321] Therefore, how can we provide a new system that enables servers to: – Based on small-scale pre-survey data and group statistics, a large amount of virtual individual data is automatically generated inside the computer; – By leveraging generative artificial intelligence models, a unified processing flow of “virtual group generation – behavior simulation – result analysis – information output optimization” is driven through prompts; – Associate emotional states obtained from text data with virtual individual data and use them as computational elements in behavioral simulation and information output optimization; – Automatically generate structured reports and various information output contents for different external information delivery media, and select or optimize them based on the estimated responsiveness; Improving the automation level of data processing links, the expressive power of behavioral simulation, and the personalization and refinement of information output at the computer level is a technical problem that urgently needs to be solved in this field.
[0322] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0323] In this invention, the server includes a unit for inputting prompt statements to a generative artificial intelligence model based on small-scale pre-survey results divided by region and group statistical information, indicating the generation of multiple virtual individual data; a unit for inputting prompt statements to the generative artificial intelligence model based on the virtual individual data and attribute information of multiple candidate objects, indicating the execution of group behavior simulation calculations and prediction of the candidate object selection results; a unit for inputting prompt statements to the generative artificial intelligence model based on the virtual individual data and individual attribute information and / or emotional state information, indicating the estimation of responsiveness to information output and comparison of the effects of multiple information output schemes; a unit for inputting prompt statements to the generative artificial intelligence model based on the selection results and the comparison results, indicating the generation of report data that can be provided to external output objects; and a unit for sending the report data to an external information transmission medium. This allows for programmable control of generative artificial intelligence models on the server side through a unified prompt statement interface. The computer can automatically complete an integrated processing flow, including virtual individual data generation, emotional state embedding, group behavior simulation, result analysis, and information output optimization. This improves the automation, computational efficiency, and prediction accuracy of election prediction and information output processing in the computer system, and enhances the adaptability of the information output content to different audiences and their emotional states.
[0324] A "system" refers to a combination of hardware and software consisting of at least one information processing device, a storage device, and a communication device, configured to perform data acquisition, data processing, model invocation, and result output.
[0325] A "processor" refers to an electronic computing unit in an information processing device that executes program instructions to perform calculations on input data, control the flow of data, and invoke external services, including but not limited to a central processing unit or a graphics processing unit.
[0326] "Information processing device" refers to a computing device equipped with a processor, memory and communication interface, used to execute application programs and receive, store, process and send data.
[0327] "Region" refers to a spatial unit defined by geographical location or administrative division, including but not limited to electoral districts, cities or other statistical areas.
[0328] "Small-scale preliminary survey results" refer to the survey data obtained and summarized from preliminary questionnaires or interviews conducted in a predetermined area with a small number of sample subjects.
[0329] "Population statistics" refers to statistical results data on the overall characteristics of a specific region or group, including but not limited to demographic data such as age distribution, gender ratio, occupational composition, and education level.
[0330] "Virtual individual data" refers to individual-level data represented in a computer data structure that does not actually exist but is generated based on statistical information and survey results. Each data point corresponds to a virtual individual with several attribute characteristics.
[0331] "Candidates" refers to the set of members of an object that is selected or evaluated in a group behavior simulation, including but not limited to candidates, solutions, products or services.
[0332] "Attribute information" refers to a set of parameters or labels used to characterize the features of candidate objects or individuals, including but not limited to information such as policy stance, functional characteristics, price range, or category identifiers.
[0333] "Group behavior simulation" refers to the process of simulating the selection behavior of multiple candidates based on virtual individual data in a computer to predict the overall selection distribution.
[0334] "Selection results" refers to the output data obtained through group behavior simulation or prediction model calculations, which shows the selection status or degree of support of multiple candidate objects.
[0335] "Individual attribute information" refers to characteristic data associated with a single virtual or real individual, including but not limited to age, gender, occupation, education level, interests and preferences.
[0336] "Emotional state information" refers to the quantitative or classification results obtained by performing sentiment analysis on text, voice or other data, used to characterize an individual's emotions or attitudes at a specific moment.
[0337] "Sentiment analysis technology" refers to computational methods and models that analyze data containing semantic content to identify and quantify the types and intensity of emotions expressed within it.
[0338] "Information output" refers to various types of content or data generated by the system and provided to users or external objects, including but not limited to reports, text descriptions, charts, recommended content, or advertising content.
[0339] "Reactivity" is a quantitative indicator of the degree of response that an individual or group may generate after receiving specific information output, such as attention, clicks, conversions, or attitude changes.
[0340] "Information output scheme" refers to multiple combinations or versions of information output that differ in content, form, wording, or presentation, targeting the same or similar objectives.
[0341] "Report data" refers to a structured or semi-structured data set that is generated based on internal calculations such as selection and comparison results and is intended to be provided to external parties. This includes, but is not limited to, numerical tables, charts, and text summaries.
[0342] "External output objects" refer to recipients of report data or information output content outside the system, including but not limited to information providers, service providers, or analysis users.
[0343] "External information delivery media" refers to external platforms or channels used to disseminate report data or information output to the end audience, including but not limited to publishing media, broadcasting media, and online media.
[0344] "Generative AI models" refer to AI models that automatically generate text, structured data, or other forms of output based on input data or prompts, including but not limited to deep learning-based language models or multimodal generative models.
[0345] "Prompt statements" refer to natural language or structured instruction texts that are constructed by the processor and input into the generative artificial intelligence model to instruct the model to perform specific generative or processing tasks.
[0346] In a preferred embodiment, the server serves as the core information processing device, the terminal as the data acquisition and result presentation device, and the user as the subject of condition setting and result utilization. These three components are connected via wired or wireless networks to form the overall system. The server includes a multi-core central processing unit, high-speed main memory, large-capacity non-volatile memory, and an optional graphics processing unit. The server runs an operating system, database management program, and multiple application modules to perform functions such as data management, feature calculation, generative artificial intelligence model invocation, group behavior simulation, and information output optimization. The terminal can be a mobile terminal or a fixed terminal, and includes a processor, display device, input device, and communication module, used to collect user data and display the server's output results.
[0347] In one implementation, the server utilizes a relational database management system to store the results of small-scale pre-surveys on regional divisions and group statistics. In another implementation, the server uses a distributed data storage system to store large volumes of virtual individual data and sentiment feature vectors. In all implementations, the server uses a general-purpose programming language as the control logic implementation tool and a data analysis library to process tabular and textual data.
[0348] In one implementation, the server constructs structured input feature vectors based on regional pre-survey results and population statistics. These vectors include age distribution histograms, gender ratio vectors, occupational category frequency vectors, education level distribution vectors, and the initial support ratios for each candidate in the pre-survey. The server combines these features into a high-dimensional numerical vector and stores it in memory as input to the virtual individual generation module.
[0349] In one implementation, the server defines virtual individual generation as a conditional data generation task performed by a generative artificial intelligence model. In this case, the server uses a generative artificial intelligence model with an encoder-decoder or autoregressive structure; the model can be a language model based on a transformer architecture or other sequence generation networks. The server inputs the prompts as natural language control signals, along with numerical feature vectors, into the generative artificial intelligence model. In one specific implementation, the server maps the numerical features to vectors through an embedding layer and concatenates them with the text embeddings obtained from the prompts to form the joint input representation of the model.
[0350] In order to generate virtual individual data, the server will construct the following type of prompt statement: "Generate 500,000 personality samples for constituency A based on the following statistics: Population distribution: 20–29 years old 18%, 30–39 years old 21%... Gender ratio: Male 48%, Female 52%. Occupation distribution: Students 12%, Company employees 45%, Self-employed 10%... Questionnaire results: Currently, candidate A has 30% support, candidate B has 40% support, and 30% are undecided. Please output a JSON-formatted array of personality samples, each containing: age, gender, occupation, education level, political interest level, currently supported candidate, and support strength (between 0 and 1)." In terms of internal model processing, the server encodes the prompt sequence into several embedding vectors, performs contextual modeling through a multi-layer attention network, and generates attribute values for virtual individuals field-by-field at the decoding end. During the training phase, the server constructs supervision signals using historical survey data and real demographic distributions, employs a cross-entropy loss function to constrain the attribute distribution of the model output to approximate the real distribution, and uses a stochastic gradient descent-like optimization algorithm to update the model weights. Through this training process, the server enables the generative artificial intelligence model to learn the mapping rules from macroscopic statistical features to microscopic individual samples, thereby allowing for the efficient generation of large-scale virtual individual data within the computer during the inference phase.
[0351] In various implementations, the server stores the generated virtual individual data in a structured format, with each record corresponding to a virtual individual and containing both numerical and enumeration fields. The server builds a composite index for this data, enabling it to quickly retrieve subgroups based on criteria such as age, occupation, region, and emotional state. Through this specific data structure and index design, the server can significantly reduce the amount of data scanned and improve query and aggregation speeds when performing group behavior simulations.
[0352] In one implementation, the server calculates sentiment state information using text data. The server receives user text from the terminal or obtains group text from an external data source. The server then invokes sentiment analysis technology to map the text to sentiment categories and corresponding intensities. In one implementation, the server calls an external sentiment analysis service; in another, it deploys an internal sentiment classification network. Regardless of the method, the server standardizes the sentiment output into a fixed-dimensional vector, such as the intensity of components like "anger," "sadness," "joy," "anxiety," and "indifference." The server merges this sentiment vector with the attribute vectors of the corresponding virtual or real individuals to form sentiment-enhanced individual features.
[0353] When constructing an emotion-enhancing personality object, the server will use the following types of prompts: "Based on the already generated personality objects in selection area A, please generate a batch of emotion-enhanced personality objects, with the emotion distribution set as follows: 40% anxiety about the economic outlook, 30% indifference to politics, and 30% optimism about reforms. Add the following fields to each object: emotion_state (enumeration: anxiety, indifference, optimism) and emotion_intensity (0–1). Please output the corresponding personality objects." The server uses high-dimensional vectors, including sentiment features, as input to simulate group behavior, constructing a classification model that outputs the selection probability for each candidate object. In one implementation, the server uses a multi-layer feedforward neural network, with its input layer receiving individual attributes and sentiment vectors; in another implementation, it uses a transformer-based structure to co-encode individuals and candidate objects to capture higher-order interactions. The server trains the network using a cross-entropy loss function and updates the weights using batch gradient descent. During training, the server randomly perturbs and downsamples the input features to augment the data and enhance the model's robustness to noise and bias.
[0354] Instead of using simple scaling rules, the server employs a pre-trained probability prediction model to calculate the matching degree between each virtual individual and each candidate when simulating voting or other selection behaviors. The server interprets these matching results as conditional probability distributions and simulates the specific choices of each virtual individual through random sampling, thereby generating a set of simulated selection results. Because the server uses an explicit probability model and sampling mechanism in the group behavior simulation phase, it can simulate complex behavioral distributions with greater precision, rather than simply linearly extrapolating pre-survey proportions.
[0355] In further processing, the server estimates the responsiveness to the information output based on virtual individual data and individual attribute or sentiment state information. The server constructs supervised samples using historical logs containing ad clicks, content dwell time, or interaction behavior, and inputs these samples, combining virtual individual features with information output features, into a rating network. The server uses a regression loss function or a ranking loss function to enable the network to learn and predict the continuous or ordinal metric of "responsiveness."
[0356] When generating advertising copy or news summaries, the server issues conditional generation tasks to the generative artificial intelligence model through prompts. For example, the server can construct the following prompts: "Target audience: College students in their 20s, whose main emotions are 'stress' and 'want to relax'. Based on this demographic data, simulate the click-through rate and favorability of the following ad copy, and provide suggestions for improvement. Ad copy: 'After finals week, take a relaxing beach trip. Students enjoy a 20% discount.' Output: Predicted click-through rate, predicted favorability, and three optimization suggestions." The server inputs the text or summary output by the generative AI model back into the reaction prediction network to obtain quantitative metrics. The server then compares multiple information output options and selects the one with the highest expected reaction rate. Through this non-human operation process of "generation-evaluation-selection," the server internally performs multiple rounds of automatic optimization, thereby improving the matching degree between the information output and the preferences of the target group. Due to this path-dependent unified prompt statement interface and shared virtual individual data structure, the server achieves a high degree of modularity and reusability in its program architecture.
[0357] When generating report data for external information delivery media, the server automatically selects the output structure based on the target media type. For example, for text media, the server generates reports with summaries, main indicators, and explanatory paragraphs; for visualization media, the server generates structured data with line charts, bar charts, and descriptions of filtering parameters. During this process, the server uses specific data structures to organize the prediction results, such as storing the main prediction results in key-value pairs ("Region ID—Candidate ID—Vote Rate—Confidence Interval") and storing explanatory information in a "Feature Contribution List." This data organization facilitates rapid querying and cross-media reuse.
[0358] In one implementation, the terminal displays the server's prediction results and information output scheme to the user. The terminal includes display components for rendering charts and text content. After receiving report data from the server, the terminal rearranges the content according to its screen size and interactivity. In some implementations, the terminal supports user filtering for specific regions or age groups to send more granular queries to the server.
[0359] In another implementation, the terminal also performs user data collection. With user authorization, the terminal reads user text input, browsing history overview, and self-reported attribute information, packages it into structured data, and sends it to the server via the network. In some implementations, the terminal can perform preliminary sentiment recognition locally and only upload sentiment tags, thereby reducing the amount of data uploaded and lowering the communication load.
[0360] In one implementation scenario, users define the target area, analysis objects, and output requirements via a terminal. Users can input content similar to the following: "Please predict the candidate with the highest probability of winning the next Tokyo gubernatorial election, taking into account the polls and social media sentiment over the past month." Upon receiving the request, the server initiates an internal data stream, loads the latest required data into memory, performs a simulation on the virtual individual set in the corresponding region, generates report data, and finally provides feedback to the user through the terminal.
[0361] In another implementation, users can directly drive other technical systems using optimized information output schemes from the server. For example, users can import advertising copy and delivery parameters generated by the server into an advertising platform. The advertising platform then controls network request routing and content distribution based on this scheme, thereby pushing personalized content to a large number of terminals in a real network environment. Therefore, this invention is not limited to abstract data processing flows, but rather achieves technical control over real information dissemination behavior through integration with external network systems and display devices.
[0362] In various embodiments of this invention, the server integrates the processes of virtual individual generation, emotion embedding, behavior simulation, result analysis, and information output optimization into a configurable processing pipeline by uniformly using a generative artificial intelligence model and prompt statement interface. Within this pipeline, the server utilizes high-dimensional vector representations and neural network structures to synthesize multiple features in a non-linear manner, thereby achieving higher prediction accuracy than traditional linear models under the same hardware resources. Because the server reduces multiple data format conversions and manual rule calls in the computation path, merging multiple traditionally independent steps into continuous vector operations, the server can improve overall throughput and reduce latency when processing large-scale virtual individual data.
[0363] In an alternative implementation, the server can employ a different generative artificial intelligence model, such as a variational autoencoder-based structure, to sample virtual individuals directly in the feature space without extensive text output. In this case, the prompts still serve as high-level control signal inputs, but the server maps them into numerical conditional vectors after parsing the prompts to drive latent space sampling.
[0364] In another alternative implementation, the server can eliminate the external sentiment analysis service and instead use a built-in convolutional network or transformer network as the sentiment parsing module to directly encode the user's text and output a sentiment vector. The server achieves end-to-end optimization by jointly training the sentiment parsing module and selecting the prediction module, making the sentiment representation more closely resemble the behavior prediction task, thereby further improving the overall prediction accuracy.
[0365] The server, through the technical configurations of the above implementation forms, enables the system to be implemented within the computer: The server performs detailed modeling of the group structure at the virtual individual level; The server models attribute features and sentiment features uniformly in a single model or a collaborative model; The server controls the generative artificial intelligence model to perform different stages of processing through prompt statements in a unified multi-step task. The server generates information output content optimized for different media types and audiences at the result level.
[0366] Thanks to the combined application of these specific data structures, neural network architectures, training methods, and prompting mechanisms, servers can achieve higher prediction accuracy, faster large-scale simulation speeds, and more reasonable data storage and access patterns with the same hardware resources. This results in significant improvements at the computer technology level, rather than simply replacing manual tasks such as simple questionnaire statistics or copywriting.
[0367] use Figure 14 The processing flow is explained.
[0368] Step 1: The user sets the analysis conditions and sends a request on the terminal. Users open the application interface on the terminal and input information such as the target region, analysis type, and a description of their needs. Input includes: region identifier, whether to perform election prediction, whether to simulate advertising effectiveness, whether to consider sentiment, and a natural language prompt (e.g., "Please predict the candidates to be elected in the next regional election, taking into account recent polls and social media sentiment."). The terminal encapsulates this input into a request data structure (e.g., containing fields such as region_id, task_type, use_emotion, and user_prompt) and sends it to the server as a network request via the communication module. The output is a structured request message sent to the server.
[0369] Step 2: The terminal collects and uploads user attribute data and text data. After obtaining user authorization, the terminal reads user attribute data (such as age, gender, occupation, and interest tags) and recent text data (such as social media posts, search keywords, and browsed topic summaries) from local storage or the interactive interface. Input consists of the user's historical activity on the terminal and information actively entered by the user. The terminal performs preliminary processing on this raw data, such as concatenating multiple texts into a single merged text and standardizing attribute fields into predefined key-value pairs. The terminal then performs simple data cleaning (removing empty strings and illegal characters) to form a structured user data object, which is then sent to the server over the network. Output consists of the user characteristics and text data uploaded to the server.
[0370] Step 3: The server loads basic data from the database and external data sources. After receiving request messages and user data from the terminal, the server retrieves the corresponding small-scale pre-survey results table and group statistical information table from its internal database based on the region identifier and task type in the request. The inputs are the region ID and task parameters. The server executes database queries to obtain data such as questionnaire sample entries, population age distribution, gender ratio, and occupational distribution. Simultaneously, the server can call external data interfaces to obtain the latest public opinion poll data, depending on task requirements. The server performs format standardization (field renaming, unit standardization) on data from different sources, loading it into an in-memory data table or feature matrix. The output is a set of raw multi-source data before cleaning.
[0371] Step 4: The server preprocesses the survey data and statistical information. The server takes the dataset output from step three as input and preprocesses the data using a data processing library. The server performs missing value imputation (e.g., using the mean or mode of similar samples), outlier detection and removal on the pre-survey results; it normalizes demographic data to convert it into proportional feature vectors later. The server also encodes categorical fields (such as occupational category and education level) into numerical vector representations. The goal of data processing is to transform heterogeneous raw data into a unified numerical feature space. The output is a structured feature dataset suitable for input into generative artificial intelligence models and predictive models.
[0372] Step 5: The server constructs prompts to generate virtual individuals and invokes generative artificial intelligence models. The server takes preprocessed statistical features as input, calculates the proportions of each age group, gender, occupation, etc., and inserts these proportions and regional descriptions into a natural language template to construct one or more prompt statements for generating personality / virtual individuals. The input consists of various proportion values and task scale parameters (e.g., the number of virtual individuals to be generated). An example of the server-generated prompt statement is: “Generate 500,000 personality objects for selection area A based on the following statistical information. Population distribution: 20–29 years old 18%, 30–39 years old 21%... Gender ratio: Male 48%, Female 52%. Occupation distribution: Students 12%, Company employees 45%, Self-employed 10%... Questionnaire results: Current support rate for each candidate... Please output an array of personality objects, each containing age, gender, occupation, education level, interest level, current supported object, and support strength (0–1).” The server takes this prompt statement as text input, along with numerical conditions (such as the number of generated individuals), and sends it to the generative AI model interface, triggering the model to generate the virtual individuals. The output is the original text or structured description of the virtual individuals returned by the model.
[0373] Step 6: The server parses the output of the generative artificial intelligence model and stores the virtual individual data. The server takes text or semi-structured results returned by a generative AI model as input. It uses text parsing and JSON parsing tools to perform syntax checking and field extraction on the model output. For each virtual individual entry, the server extracts fields such as age, gender, occupation, education level, political interests, current support, and support strength, converting them into an internally defined data structure (e.g., database records or vector rows). The server discards or repairs data that fails to parse or has incomplete fields. Data processing is accomplished through field mapping and type conversion steps. The output is a large collection of virtual individual data stored in a database or memory.
[0374] Step 7: The server calculates sentiment states based on user text or group text and associates them with virtual individuals. The server takes user text uploaded by the terminal and / or group text obtained from external sources as input. It calls the sentiment analysis module or sentiment analysis interface to segment, extract sentiment features, and classify each text segment. The server outputs a multi-dimensional sentiment vector, representing the intensity values of sentiment components such as "joy," "sadness," "anger," "anxiety," and "indifference." The server associates this sentiment vector with the corresponding individual identifier to generate "sentiment state information." For group-level sentiment, the server can randomly or according to rules assign it to some virtual individuals based on regional and population characteristics, thus forming "sentiment-enhanced virtual individuals." This processing involves converting the original text into numerical sentiment features and embedding them into the virtual individual feature records. The output is a dataset of virtual individuals with sentiment feature fields.
[0375] Step 8: The server constructs emotion-enhancing prompts to supplement or adjust the emotional distribution of virtual individuals. When the server needs to adjust the overall sentiment distribution, it uses the current virtual individual sentiment distribution statistics and the target sentiment ratio as input to construct prompts for regenerating or correcting sentiment fields. For example: "Based on the already generated personality objects, please generate a batch of sentiment-enhanced personality objects, with the sentiment distribution set as follows: 40% anxiety about the economic outlook, 30% political indifference, and 30% optimism about reform. Add the following fields to each object: `emotion_state` (anxiety, indifference, optimism) and `emotion_intensity` (0–1)." The server sends these prompts to the generative AI model, which outputs new or updated sentiment field descriptions. The server parses these descriptions and updates the original virtual individual records. The output is a dataset of virtual individuals that meets the target sentiment distribution.
[0376] Step 9: The server constructs behavioral simulation input features and loads a prediction model. The server takes as input virtual individual data with sentiment enhancement fields and attribute information from multiple candidate objects. For each virtual individual, the server constructs a feature vector, including numerical encodings of age, gender, occupation, education level, interest tags, and sentiment state. Simultaneously, the server constructs attribute vectors for each candidate object (such as a candidate or proposal), including numerical features such as key issues, policy tags, and historical performance indicators. The server loads the virtual individual feature matrix and candidate object feature matrix into memory, along with a pre-trained behavior prediction model (e.g., a multi-layer feedforward neural network or an attention-based model). The output is the feature data and model instance ready for inference.
[0377] Step 10: The server performs a group behavior simulation and generates a prediction of the selection outcome. The server takes the feature matrix and behavior prediction model from step nine as input and performs forward inference operations on each virtual individual and each candidate object combination. The server performs matrix multiplication, activation function calculation, and normalization operations in the neural network to obtain the selection probability distribution of each virtual individual for each candidate object. Based on these probabilities, the server determines the simulated selection result for each virtual individual by random sampling or selecting the option with the highest probability. Subsequently, the server aggregates the simulated selections of all virtual individuals, counts the total number of votes or selections for each candidate object, and calculates the proportion and confidence interval. Data computation is essentially batch neural network inference and probability sampling of high-dimensional features. The output is the group selection prediction result for a specific task.
[0378] Step 11: The server evaluates the responsiveness of different information output schemes. The server takes virtual individual characteristics, emotional state, and several candidate information output schemes (such as multiple advertising copy or news summary) as input. It concatenates the content features (text embedding, length, tone, etc.) of each scheme with the virtual individual characteristics to form a joint feature vector. The server then inputs these joint features into a pre-trained responsiveness prediction model (such as a regression network or ranking network) to calculate the expected click-through rate, dwell time, or favorability of each scheme within the target audience. This data computation includes multiple forward inferences and result aggregations. The output is a list of responsiveness metrics corresponding to each information output scheme.
[0379] Step 12: The server selects or optimizes the information output content based on responsiveness. The server takes the response index obtained in step eleven as input. It first compares the index values of each candidate solution and selects one or more solutions with the best response index as candidate outputs. When further optimization is needed, the server can construct new prompts to request the generative AI model to adjust the copy. For example: "Target audience: College students in their 20s, currently experiencing stress and wanting to relax. Please generate 5 different styles of travel advertisement copy, each no more than 50 characters, with styles of soothing, encouraging, humorous, data-driven, and story-based." The server repeatedly evaluates the response index of the newly generated copy and ultimately determines the optimal content. The output is the optimized information to be presented to external media or users.
[0380] Step 13: The server generates report data adapted to external information transmission media. The server takes the group behavior prediction results and information output scheme optimization results as input, and generates report data with corresponding structure based on the target external media type (e.g., text media, image media, network interface). The server encapsulates key indicators (e.g., prediction support rates for each candidate), explanatory information (e.g., characteristics of the most influential virtual individuals), and recommendation information output schemes into structured data. The server also generates data points and annotation information needed for visualization, allowing clients or other systems to create charts. Data processing includes field reorganization, unit standardization, and format conversion. The output is a report data package that can be consumed by external systems.
[0381] Step 14: The server sends the results to the terminal and external information transmission medium. The server takes the report data generated in step thirteen as input and sends it to the terminal and pre-configured external information delivery media (such as news publishing systems and advertising delivery systems) via the network communication module. The server can compress and encrypt the data before sending to reduce communication load and improve security. The output is the result data that has been transmitted to each recipient over the network.
[0382] Step 15: The terminal receives and displays the prediction results and information output. The terminal takes report data from the server as input, reads the structured data through a parsing module, and combines it into text blocks, charts, and lists suitable for screen display. The terminal displays election prediction results, explanations of key influencing factors, and recommended advertisements or information content on the interface. The terminal can dynamically request more granular data from the server based on user actions (zooming, filtering). The output is a visualized results interface presented on the terminal's display device.
[0383] Step 16: Users view the results and provide feedback on preferences or ratings. Users input the results displayed on the terminal, browse the prediction results and recommended information, and can select feedback options on the interface, such as rating the prediction accuracy and marking advertisements as "useful," "indifferent," or "disliked." The terminal organizes the user's feedback into structured feedback data and sends it to the server. The output is the uploaded feedback data.
[0384] Step 17: The server updates or retrains relevant models based on user feedback. The server takes feedback data uploaded from the terminal as input, associating and storing the feedback with corresponding virtual individual characteristics, emotional states, and information output scheme features. The server can use this feedback data as new supervisory signals in offline or online phases to fine-tune or retrain the behavior prediction model and responsiveness prediction model. Data computation includes sample construction, loss function calculation (e.g., mean squared error, cross-entropy, or ranking loss), and weight updates. The output is updated model parameters and improved prediction performance, thus providing more accurate and user-preferred results in subsequent request processing.
[0385] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0386] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0387] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0388] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0389] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0390] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0391] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0392] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0393] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0394] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0395] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0396] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0397] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0398] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0399] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0400] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0401] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0402] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0403] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0404] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0405] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0406] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0407] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0408] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0409] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0410] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0411] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0412] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0413] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0414] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0415] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0416] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0417] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0418] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0419] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0420] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0421] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0422] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0423] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0424] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0425] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0426] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0427] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0428] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0429] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0430] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0431] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0432] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0433] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0434] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0435] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0436] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0437] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0438] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0439] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0440] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0441] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0442] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0443] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0444] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0445] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0446] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0447] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0448] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0449] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0450] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0451] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0452] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0453] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0454] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0455] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0456] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0457] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0458] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0459] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0460] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0461] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0462] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0463] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0464] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0465] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0466] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0467] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0468] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0469] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0470] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0471] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0472] In addition, the following notes are provided in response to the above explanation.
[0473] Example 1 (Note 1) An information processing system, characterized in that it comprises: An apparatus for obtaining survey information and population attribute information recorded according to election units from an information storage device, and performing preprocessing on the survey information and population attribute information, including missing value completion processing and outlier removal processing, to generate preprocessed population attribute distribution information and preprocessed support tendency information. An apparatus for generating prompt statements based on the preprocessed population attribute distribution information and the preprocessed support tendency information, instructing a generative artificial intelligence model to generate multiple individual information objects including age attribute, gender attribute, occupation attribute, education attribute and support candidate attribute, inputting the prompt statements into the generative artificial intelligence model, and obtaining the multiple individual information objects from the generative artificial intelligence model. An apparatus for performing simulated voting processing using statistical processing or machine learning processing based on the supporting candidate attributes and population attribute information contained in the plurality of individual information objects, in order to generate prediction result information including the estimated votes of each candidate and the estimated results of the elected candidates. An apparatus for converting the prediction result information into tabular data or visual image data, and updating the prediction result information according to the user's correction instructions, thereby generating updated prediction result information and outputting it as data for information provision. A means for transmitting the information providing data to various types of information providing devices via a communication network.
[0474] (Note 2) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating prompt statements to instruct a generative artificial intelligence model to generate multiple individual information objects is configured to calculate the occurrence ratio of each population attribute category from the preprocessed population attribute distribution information, automatically generate prompt statements containing constraints reflecting the occurrence ratios, and input the prompt statements into the generative artificial intelligence model so that the generative artificial intelligence model generates the multiple individual information objects having a population composition corresponding to the occurrence ratios.
[0475] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for processing the prediction result information is configured to receive media type information as input, and based on the media type information, generate a prompt statement to instruct a generative artificial intelligence model to perform format conversion processing, input the prompt statement into the generative artificial intelligence model, so that the generative artificial intelligence model performs format conversion processing on the prediction result information, thereby obtaining media-specific provision data including at least one of text data for reporting, summary data for broadcasting, or structured data for electronic distribution, and provide the media-specific provision data as information provision data to the information provision apparatus.
[0476] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: An apparatus for enabling a computing device located in an information processing device to acquire survey information and population composition information corresponding to multiple geographic units, and to perform missing value completion, outlier removal, classification value normalization and summary processing on the survey information based on the population composition information, thereby generating population composition distribution data corresponding to each geographic unit. A device for enabling the computing device to perform random number generation and weighted sampling processing based on the population composition distribution data, so as to generate a large number of virtual individual data for each regional unit, and to attach attribute information and behavioral tendency information to the virtual individual data, while clustering the virtual individual data through statistical learning processing and assigning identification information to each cluster. An apparatus for enabling the computing device to extract statistical characteristics of a population that meets the target conditions from the virtual individual data and its clusters based on target conditions from the terminal device and target information related to information presentation, and to generate a prompt statement that describes the statistical characteristics in natural language and includes prompt information, and to input the prompt statement into a generative artificial intelligence model to obtain generated text data containing a description of the characteristics of the virtual receiver group corresponding to the target conditions and multiple information presentation content schemes. An apparatus for enabling the computing device to parse the generated text data, extract evaluation indicators for each of the information presentation content schemes, convert the information presentation content schemes into a comparable form using the evaluation indicators, and send the comparison results and the information presentation content schemes to the terminal device for display. A device for enabling the computing device to generate predictions about election results or advertising effectiveness based on the information presentation content scheme, and for converting the predictions into a distribution format for various information delivery media.
[0477] (Note 2) The information processing system according to Appendix 1 is characterized in that, An apparatus for enabling the computing device to convert attribute information and behavioral tendency information attached to the virtual individual data into numerical vectors, performing clustering processing on the numerical vectors to divide the virtual individual data into multiple groups, statistically calculating representative demographic characteristics and behavioral characteristics for each group, generating explanatory text containing the representative characteristics, and including the explanatory text in the prompt statement to automatically constitute the content input to the generative artificial intelligence model.
[0478] (Note 3) The information processing system according to Appendix 1 is characterized in that, An apparatus for enabling the computing device to record, in association with the selection results and edited content of the information presentation content scheme from the terminal device and the generated text data, and to update the generation rules of the prompt statement or the calculation conditions of the evaluation index based on the recorded content, thereby iteratively optimizing the control of the generative artificial intelligence model and the evaluation processing of the information presentation content scheme.
[0479] Example 2 (Note 1) An information processing system, characterized in that it comprises: A unit used in an information processing device to input prompt statements into a generative artificial intelligence model, based on survey results from regional units and group attribute information, to instruct the generation of personality objects representing virtual characters. A unit used in an information processing device to compare the attribute information of a generated personality object with the policy information of a candidate subject, thereby simulating the selection of the most suitable candidate subject for each personality object and calculating the probability of the candidate subject being elected based on the simulation selection result, and inputting prompt statements to a generative artificial intelligence model to instruct the execution of the selection behavior simulation process. A unit for setting a data structure in a storage area on a recording medium device in order to store policy information of the personality object and the candidate subject in an information processing device, and for registering, updating and obtaining policy information of the personality object and the candidate subject according to the data structure. A unit used in an information processing device to summarize and statistically process the voting result data obtained as the result of the selected behavior simulation processing, so as to calculate the number of votes, the vote rate and the election prediction index of each candidate subject. A unit for generating chart data and explanatory text data based on the statistical processing results in an information processing device, and inputting prompt statements to a generative artificial intelligence model to instruct the output of the chart data and explanatory text data in a form suitable for information dissemination platforms. This unit is used in an information processing device to control the generative artificial intelligence model, enabling it to automatically generate personality objects, simulate selected behaviors, and automatically generate explanatory text based on the statistical processing results, and to input various prompt statements in stages to specify the processing content and output format.
[0480] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to sequentially generate a predetermined number of personality objects by means of prompts input to the generative artificial intelligence model before performing the selected behavior simulation processing using group attribute information and candidate subject policy information, and to store the generated personality objects in the storage area of the recording medium device, and to use the entire group of stored personality objects as the input object for the selected behavior simulation processing.
[0481] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to input the statistical processing results and the condition information of the selected behavior simulation processing as a data set for analysis into the generative artificial intelligence model, so that the generative artificial intelligence model automatically generates explanatory text data including at least one of news report text, analysis report or explanatory text, and provides the information provider with prompt statements to instruct the generation and output of the explanatory text data.
[0482] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A unit for execution by a processor in an information processing device, wherein the processor is configured to input prompt statements for instructing the generation of multiple virtual individual data into a generative artificial intelligence model based on the results of a small-scale pre-survey divided by region and population statistics. A unit for the processor to input prompt statements into the generative artificial intelligence model based on the virtual individual data and attribute information of multiple candidate objects, indicating the execution of a group behavior simulation operation and predicting the selection result of the candidate objects; A unit for inputting prompt statements into the generative artificial intelligence model by the processor based on the virtual individual data and individual attribute information and / or emotional state information, for indicating the estimated responsiveness to information output and comparing the effects of multiple information output schemes; A unit for the processor to analyze the selection results and comparison results, and input prompt statements to the generative artificial intelligence model to instruct the generation of report data that can be provided to external output objects; A unit for sending the report data to an external information transmission medium by the processor.
[0483] (Note 2) The information processing system according to Appendix 1 is characterized in that, The processor is also configured to use text data obtained from individuals to calculate emotional states through sentiment analysis technology, and to input prompt statements into the generative artificial intelligence model to instruct the association of the emotional states with the virtual individual data and to weight the selected behavior in the group behavior simulation operation.
[0484] (Note 3) The information processing system according to Appendix 1 is characterized in that, The processor is also configured to input prompt statements into the generative artificial intelligence model based on the virtual individual data and the emotional state, for instructing the generation of various information output contents, calculating the responsiveness index for each information output content, and selecting and / or optimizing the information output content based on the responsiveness index.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured as follows: The prompts for generating personality objects based on the results of small-scale pre-questionnaires and demographic information in each constituency are input into the generative artificial intelligence model, so that the generative artificial intelligence model generates multiple personality objects; A prompt for instructing the generative artificial intelligence model to perform a predictive simulation of elected candidates in an election using the generated personality object is input into the generative artificial intelligence model, so that the generative artificial intelligence model outputs simulation results related to the candidate election prediction; A prompt indicating that the prediction results should be analyzed and output in a form suitable for sale to various information media is input into the generative artificial intelligence model, so that the generative artificial intelligence model outputs prediction result data for sales.
2. The information processing system according to claim 1, characterized in that, The processor is configured to input a prompt instructing the generation of personality objects based on the demographic data of each constituency into the generative artificial intelligence model, so that the generative artificial intelligence model generates personality objects corresponding to the population structure of each constituency based on the demographic data.
3. The information processing system according to claim 1, characterized in that, The processor is configured to input prompts to the generative artificial intelligence model, instructing the analysis of the prediction results and their provision in a form suitable for national newspapers, local newspapers, and television media, so that the generative artificial intelligence model outputs prediction result data for the media.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A