system
The system addresses regional environmental prediction inaccuracies by integrating observation and interview data with AI models to provide accurate future predictions for resource allocation and business planning.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-24
- Publication Date
- 2026-07-06
AI Technical Summary
Conventional methods for regional environmental prediction rely on limited statistical information, neglecting non-quantitative local insights, leading to inaccurate guidelines for investment planning and misallocation of resources, hindering regional revitalization and quality of life improvement.
A system integrating observation and interview information, utilizing a server with AI models to generate predictive information for resource allocation policies, incorporating environmental and stakeholder opinions, and generating indicator information for business plans.
Accurately predicts future regional environmental characteristics, enabling optimal resource allocation and business planning by comprehensively utilizing observational and interview data.
Smart Images

Figure 2026112135000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventionally, future predictions regarding the regional environment were made relying on limited publicly available statistical information and some indicators, so only superficial elements were considered and there was a problem that the actual situation in the region could not be accurately reflected. Also, non-quantitative information from local residents and stakeholders was not fully utilized, and it was difficult to provide accurate guidelines in medium- and long-term investment planning and strategy formulation. As a result, there were problems of misallocation of resources and business expansion, hindering the realization of regional revitalization and improvement of the quality of life of residents.
Means for Solving the Problems
[0005] In this invention, the server includes: [observation information acquisition means for acquiring observation information consisting of environmental information including structures, moving objects, and brightness states in multiple spatial areas]; [listening information acquisition means for acquiring listening information consisting of opinion information provided by local stakeholders]; [instruction statement generation means and prediction information acquisition means for generating instruction statements for predicting regional environmental characteristics multiple periods in the future based on the observation information and the listening information, and inputting them into a generation AI model]; and [indicator information output means for outputting indicator information applicable to resource allocation policies or business plan formulation based on the prediction information]. This makes it possible to accurately grasp the actual state of the regional environment and derive optimal strategies and investment plans from a medium- to long-term perspective.
[0006] "Observational information" refers to environmental information including structures, moving objects, and brightness conditions in multiple spatial areas, which objectively and visually represent the actual conditions of the target area. "Information gathered through interviews" refers to opinion information provided by local stakeholders, which includes subjective and experiential knowledge about the local environment. "Instruction statement" refers to a text that represents a command or request input into the AI model based on the observation information and the interview information, and includes conditions and indicators for predicting future regional environmental characteristics from the observation information and the interview information. A "generative AI model" is a trained information processing means that generates information for predicting future regional environmental characteristics in response to input instructions, and is a model that operates using artificial intelligence technologies, including natural language processing. "Predictive information" refers to information indicating future regional environmental characteristics, output by a generation AI model in response to the aforementioned instructions, and provides guidance that contributes to the formulation of resource allocation policies and business plans. "Indicator information" refers to information that provides quantitative or qualitative criteria for use in making decisions regarding resource allocation policies or business plan formulations based on predictive information. [Brief explanation of the drawing]
[0007] [Figure 1] This is a sequence diagram showing the processing flow of the system in this embodiment. [Modes for carrying out the invention]
[0008] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0009] <System Configuration>
[0010] An example of an embodiment for carrying out the present invention is shown below. Note that the configuration shown below is merely an example and can be modified without departing from the spirit of the present invention.
[0011] In this embodiment, the present invention is realized as a system that predicts future characteristics of the local environment based on observational and interview information and generates indicator information for resource allocation policies or business plan formulation. This system is implemented with a configuration including at least a "server" and a "terminal".
[0012] The server is an information processing device and can utilize general computer equipment as a program execution environment. For example, the server preferably has a configuration that includes multiple processing units (CPU, GPU, etc.), main memory, auxiliary memory, and a network interface. The server may have a database for receiving and storing observation information (environmental information including structures, moving objects, and brightness states in multiple spatial domains) and interview information (opinion information obtained from local stakeholders) provided from terminals. The server has the function of generating instruction sentences based on the observation and interview information and inputting these instruction sentences into a generation AI model to obtain predictive information regarding future regional environmental characteristics. The server also generates indicator information applicable to resource allocation policies or business plan formulation based on the predictive information and transmits it to the terminal described later. The server includes software modules for executing this series of processes. For example, natural language processing means, image analysis means, time series analysis means, and interface modules for utilizing the generation AI model are implemented to analyze the observation and interview information.
[0013] The terminal is an information processing device capable of communicating with a server via a network. The terminal can be equipped with multiple means for acquiring observational information, such as a camera, microphone, brightness sensor, motion sensor, and temperature / humidity sensor. The terminal executes programs (e.g., image analysis libraries, speech recognition engines, sensor value formatting scripts) to preprocess the information obtained from these acquisition means. In addition to acquiring observational information, the terminal also acquires auditory information by directly contacting local stakeholders and recording and transcribing their statements. A general-purpose speech recognition engine can be used for transcription. The terminal transmits the acquired observational and auditory information to the server using protocols such as HTTP. The terminal also receives indicator information transmitted from the server and visualizes it on a map display screen or dashboard-style UI for presentation to the user. The terminal can be implemented in configurations such as a mobile computing device, tablet device, personal digital assistant, or a stationary information processing device with an attached display.
[0014] The communication method connecting the server and terminal can be wired or wireless, and may include an internet connection or a local network environment. The server stores the observation and interview information transmitted from the terminal in its memory and performs preprocessing on this information (missing value imputation, keyword extraction from text, object detection in images, time axis analysis, etc.). The server then generates prompt sentences (instructions) using the preprocessing results and inputs them into a generating AI model to obtain predictive information that forecasts regional environmental characteristics with a view to the next 2-3 years. The server further generates indicator information from this predictive information, such as regional investment priorities, infrastructure development needs, and estimated investment recovery periods. When the terminal receives this indicator information, it visualizes it on the user interface, enabling users (regional government officials, private businesses, infrastructure project stakeholders, etc.) to formulate effective resource allocation policies and business plans based on this information.
[0015] With this configuration, the system of this embodiment can accurately predict future regional environmental characteristics by comprehensively utilizing observational and interview information, and can provide users with decision-making information based on those predictions.
[0016] <Specific examples of system configuration>
[0017] The server and terminal configured in this embodiment will be described in more detail below.
[0018] (Server configuration) The server is a core information processing device for integrating diverse information about the local environment and predicting future characteristics of the local environment. The server can be configured as a general-purpose information processing device, for example, comprising multiple processing units (CPU, GPU), main memory (DRAM), secondary storage (HDD, SSD), and a network interface. An operating system (UNIX®-based or other appropriate operating system) runs on the server, and multiple software modules are executed on top of it.
[0019] The server is equipped with analytical means for processing observational and interview information received from terminals. The observational information is environmental information including structures (buildings, facilities, etc.), moving objects (vehicles, bicycles, pedestrians, etc.), and brightness conditions (nighttime illuminance, horizontal light source density, etc.) in the spatial domain of the target area. The server has a database for storing this observational information, and time-series data, location information, category information, etc. are associated within this database. The interview information is information including subjective opinions and experiential knowledge obtained from local stakeholders, and the server stores this as text data in the database. The server integrates and processes this observational and interview information, and uses natural language processing means, image analysis means, and time-series analysis means as preprocessing for making the desired future predictions.
[0020] As a natural language processing means, the server performs keyword extraction, sentiment analysis, and theme classification on text data (transcription information). For example, if local residents make a statement such as "There are more stores operating at night", keywords such as "night operation", "increase", and "store" are extracted from the text and treated as elements suggesting changes in the local environment. Also, as an image analysis means, the server analyzes the pedestrian flow density in the street, building deterioration degree, vehicle type distribution, etc. using the image / video data included in the observation information. The time series analysis means is used to perform variation analysis of these data on the time axis, and for example, changes such as "the late-night outings of the young generation have been increasing over the past six months" can be extracted. These preprocessing results are used as materials for generating the instruction text that the server inputs into the generated AI model.
[0021] Based on these analysis results, the server generates an instruction text for predicting the characteristics of the local environment two to three years in the future. The instruction text is the input information to the generated AI model and includes the target area, time period, and factors to be considered (such as the behavior patterns of the young generation, infrastructure development status, utilization rate of commercial facilities, etc.). The server sends this instruction text to the generated AI model and obtains prediction information as a response. The generated AI model is an advanced pre-trained model based on natural language processing and makes future predictions based on the input instruction text.
[0022] The server further reprocesses the obtained prediction information and generates index information applicable to resource allocation policies or business plan formulation. Examples of this index information include the priority of infrastructure investment in a specific area, the effectiveness of commercial facility development, and the location plan for new services targeting the young generation. The server sends this index information to the terminal to enable visualization and presentation on the terminal.
[0023] (Embodiment of the Terminal) The terminal is an information processing device that acquires observational and auditory information in the local environment and transmits it to the server. The terminal may have a configuration that allows connection of multiple external sensors, such as a camera, microphone, illuminance sensor, temperature and humidity sensor, and motion detection sensor. For example, the terminal can be a portable small computer device and can be placed in the observation area, whether outdoors or indoors. The camera acquires images or videos of the target location or streetscape, and the microphone collects street sounds and human voices. The illuminance sensor quantifies the brightness at night, and the motion detection sensor can measure the number of pedestrians and vehicles passing by.
[0024] The terminal also features an interface for conducting interviews with local residents and stakeholders. For example, it can record the voices of interviewees using a microphone connected to the terminal and transcribe them into text using speech recognition software. The transcribed interview information undergoes basic formatting (conversion to data formats such as JSON) within the terminal and is then sent to the server. The terminal communicates with the server via a communication line (wireless LAN, mobile communication network, wired line, etc.) and uploads observation and interview information to the server periodically or at predetermined triggers (e.g., after a certain period of time, detection of a specific event).
[0025] The terminal may have a function to visualize the indicator information received from the server on the user interface. For example, it may display a regional map on the terminal's screen and color-code the investment priority and the degree of infrastructure development necessity for each region. It can also provide an interface that makes it easy for users to formulate strategies based on the obtained indicator information using text-based reports and graph displays. The terminal can be configured as a tablet, mobile device, notebook PC, or stationary workstation, and can be selected according to the user's operability and screen size.
[0026] As described above, by organically coordinating the terminal and the server, this embodiment makes it possible to effectively utilize observational and interview information that reflects the actual local environment to make future predictions, and to provide the results to the user as information useful for resource allocation policies and business plan formulation.
[0027] <System Operation>
[0028] The server executes the system's programs using Python 3.8 running on a Linux environment. The server stores observation and interview information in a structured format using the MySQL database management system and functions as a web server receiving HTTP requests from terminals using Nginx. The server analyzes observation and interview information using software such as spaCy for natural language processing, statsmodels for time series analysis, and OpenCV for image analysis, and generates prompts for the generative AI model using the analysis results. The server also obtains future predictions using generative AI models such as the OpenAI GPT-4 API, and derives indicator information from this prediction information that can be used to formulate resource allocation policies or business plans.
[0029] The terminal runs Python scripts on small computers such as Raspberry Pi® or tablet devices, and acquires observation information using Logitech® cameras, microphones, light sensors, temperature and humidity sensors, etc. The terminal uses a speech recognition engine such as the Whisper® model to transcribe the listening content into text and sends the observation and listening information to the server via HTTP requests. The terminal receives the indicator information returned from the server in JSON format and presents it to the user via map display tools such as Leaflet® or Mapbox®, or a user interface using HTML / CSS / JavaScript®.
[0030] Users can visually refer to indicator information such as regional investment priorities, infrastructure development needs, and payback periods on their device screens. Using this information, users can formulate resource allocation policies and investment plans.
[0031] (Example of a prompt message) "The following is observational data from a certain local area over the past two years (traffic volume increase rate, changes in nighttime brightness index, increase in late-night activity among young people) and opinions from local residents and businesses that 'areas where young people gather late at night are becoming more active.' Based on this data, comprehensively predict and summarize the potential for commercial area expansion, infrastructure development needs, and changes in the living environment in this area in 2-3 years."
[0032] <Specific example of system operation (Figure 1)>
[0033] Step 1: The terminal acquires observational information (video data, audio data, brightness values, weather values) and auditory information (audio opinions from local stakeholders) using a camera, microphone, light sensor, and temperature / humidity sensor. The terminal receives signals from the above sensor devices (camera images, recorded audio, sensor values) as input, processes them using a Python script with OpenCV for image analysis, pyaudio® for audio input / output, and I2C communication for environmental value acquisition, and generates formatted observational information and auditory information (including transcribed audio information) as output. This output data is converted to JSON format for subsequent processing and sent to the server.
[0034] Step 2: The server receives observation and interview information (input) sent from the terminal via HTTP requests using Nginx as the frontend and stores it in a MySQL database. The server receives the input data in JSON format and performs data processing and calculations such as formatting the time series data using pandas®, extracting text using a natural language processing tool (spaCy), and associating it with information analyzed with OpenCV, to obtain a reformatted, analyzable dataset as output.
[0035] Step 3: The server analyzes the changes from the past to the present using time series analysis (statsmodels®) on a formatted dataset (input), extracting fluctuation trends and features. The server performs data calculations on this input data, such as calculating the rate of change in traffic volume, analyzing the frequency of times when young people are present, and averaging brightness values over time, and as output, obtains an extended dataset with a set of features necessary for future prediction (such as the rate of increase in traffic and the nighttime brightness trend index).
[0036] Step 4: The server automatically generates prompts for the generating AI model using the expanded dataset (input). Here, the server performs data calculations that comprehensively consider keywords extracted from transcribed interview information (such as "increase in nighttime business operations" and "late-night use by young people") and trends in urban areas obtained through image analysis, and as output, creates instructional sentences (prompts) that can be input into the generating AI model. An example of such a prompt is: "The following is observational data from the past two years in a certain local area (traffic volume increase rate, changes in nighttime brightness index, increase in late-night activity by young people) and opinions obtained from local residents and businesses. Based on this data, comprehensively predict and summarize the possibility of commercial area expansion, infrastructure development needs, and changes in the living environment in this area in 2-3 years."
[0037] Step 5: The server sends the above prompt message (input) to the generating AI model and receives future prediction information (output) as a response from the generating AI model. The server accesses the generating AI model via the OpenAI GPT-4 API and receives text information indicating the regional environmental characteristics 2-3 years in the future (for example, "Area A has a high priority for commercial investment, and Area B will require infrastructure improvements in the medium term") as a response based on the input prompt message.
[0038] Step 6: The server analyzes the received future forecast information (input) and extracts and generates indicator information useful for resource allocation policies or business plans. The server performs text mining and rule-based analysis on this input information and obtains indicator information in JSON format as output, including specific investment plan indicators (for example, "Area A: High priority for commercial investment (5-year payback period)", "Area B: Medium-term need for infrastructure strengthening (7-year payback expected)").
[0039] Step 7: The terminal visualizes the indicator information (input) received from the server using HTML / CSS / JavaScript and map display tools such as Leaflet, and presents it to the user. The terminal analyzes the input indicator information, processes the data by displaying color-coded areas on a map, or displaying summaries in graph or text format, and generates a screen display that the user can directly check as output. Based on this display result, the user can make decisions regarding the allocation of regional resources and investment strategies.
Claims
1. An observation information acquisition means for acquiring observation information consisting of environmental information including structures, moving objects, and brightness states in multiple spatial domains, A means for acquiring interview information consisting of opinion information provided by local stakeholders, Instructions generation means for generating instructions for predicting regional environmental characteristics multiple periods in the future based on the observation information and the interview information, A prediction information acquisition means inputs the instruction sentences generated by the instruction sentence generation means into a generation AI model and obtains prediction information from the generation AI model that shows the regional environmental characteristics after the multiple periods, An indicator information output means generates and outputs indicator information applicable to the formulation of resource allocation policies or business plans based on the aforementioned forecast information. A system that includes this.
2. The system according to claim 1, further comprising means for the instruction sentence generation means to analyze the observation information and the interview information for each period to extract regional environmental change trends, and dynamically change the prompt sentence to be input to the generated AI model based on the results of the extraction of said regional environmental change trends.
3. The system according to claim 1, wherein the predictive information acquisition means is a means for organizing the predictive information obtained from the generated AI model according to regional environmental characteristics, and for successively generating a plurality of prompt sentences for formulating resource allocation policies or business plans according to the results of organizing the predictive information.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A