Interview server and interview video generation system
Patent Information
- Application Number
- JP2025031185
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
Smart Images

Figure 2026144087000001_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to an interview server and an interview video generation system that generate videos from interview information based on highly reliable resources. BACKGROUND ART
[0002] Conventionally, techniques as described in Patent Document 1 are known. In the information processing system according to Patent Document 1, when a processing unit classifies a plurality of pieces of information from one or more information sources, and performs preprocessing to extract information with a reliability higher than a specific value from the classified two or more pieces of information, the information processing system performs a specific process of creating an article based on the extracted information. Specific examples thereof are shown on pages 16 and 17, where stepwise evaluation is performed to determine whether to adopt a resource or not. PRIOR ART DOCUMENTS PATENT DOCUMENTS
[0003] Patent Document 1 Japanese Unexamined Patent Publication No. 2025-471 SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION
[0004] However, the above conventional technology has a problem that the reliability of the generated article becomes low because resources are based on low-reliability content such as personal SNS posts. The present invention has been made to solve such a problem. MEANS FOR SOLVING THE PROBLEM
[0005] The interview server of the present invention supplies interview resources, including text, audio, or images, to a video generation server, causing the video generation server to generate an interview video. The server includes an interface means for the user and a question management means for setting questions for the interview subject, displaying the questions on the interface means in a predetermined order, and acquiring text, audio, or images. Preferably, the question management means issues image capture instructions through the interface means. Furthermore, it is preferable to transmit the interview resources acquired based on the questions to the video server.
[0006] Furthermore, the interview video generation system comprises the aforementioned interview server and a video generation server connected via the network that generates interview videos based on interview resources obtained through the aforementioned questions. [Brief explanation of the drawing]
[0007] [Figure 1] This is a diagram illustrating the configuration of a video recording system. [Figure 2] Figure 1 is a block diagram showing the interview video generation system. [Figure 3] This is a block diagram illustrating the online functionality. [Figure 4] This is a diagram showing a video reporting system according to Embodiment 2 of the present invention. [Modes for carrying out the invention]
[0008] Figure 1 is a configuration diagram showing a news video generation system according to an embodiment of the present invention. Figure 2 is a block diagram of the news video generation system shown in Figure 1. This news video generation system 100 consists of a news computer 1 with a predetermined program installed, a portable information terminal 2 connected to the news computer 1, and a video computer 3 with a video generation program installed. Note that the news computer 1 and the video computer 3 may be composed of multiple computers.
[0009] The following elements are comprised of the hardware and software of the above-mentioned computer 1 used for news gathering. As shown in Figure 2, the news gathering video generation system 100 consists of a news gathering server 101, a video generation server 102, and a user interface unit 103. The news gathering server 101 and the video generation server 102 are connected by a network, but they may also be composed of the same hardware.
[0010] The interview server 101 includes a question management unit 111 for setting questions such as question items (including the subject and purpose of the questions; the same applies hereinafter), a question guidance unit 112 for guiding the user to the next question during the questioning process, and a check unit 113 for checking for omissions in the question content and the quality of the questions.
[0011] The Question Management Unit 111 allows the administrator to set the questions. The questions are composed of perspectives from various fields, such as those from experts like certified small and medium-sized enterprise consultants and tax accountants, from business people like managers, bankers, and MBA holders, and from technical perspectives like engineers and patent attorneys. There are no specific limitations on the content, but a balanced approach is preferable to ensure the reliability and quality of the interview. The Question Management Unit 111 organizes and stores the questions by field and extracts them so that questions are distributed evenly across each field. For example, the Question Management Unit 111 stores multiple questions for each field, randomly selects two questions from each field, and provides them to the administrator in a predetermined order. The administrator then sees the questions displayed on the mobile information terminal 2, which is the Interface Unit 103, and asks these questions to the interviewee (the person being interviewed).
[0012] Furthermore, the system can classify and store questions by field, and extract questions from each field to ensure that only questions from that field are asked. If the subject of the interview is a technical product, the question management unit 111 extracts multiple technical questions and displays them on the mobile information terminal 2, which is the interface unit 103. When the question management unit 111 organizes questions by field and stores them in a database, it can further classify each field in detail and select questions from within. In the case of technical fields, questions can be classified into mechanical, electrical, chemical, etc., and questions can be extracted from within these categories.
[0013] The question management unit 111 recognizes terms contained in the interviewee's words from the microphone of the interface unit 103 and automatically extracts questions related to these terms. These relationships are pre-configured and stored in the question management unit 111. Alternatively, AI may be used to automatically generate and associate these relationships. For example, if the term "genes" comes up in the middle of a discussion about business management, the system will extract questions related to the technology of "genes." In this case, extracting questions about the technology of "genes" after the business questions have been completed avoids interrupting the flow of the conversation. In this case, the question, "You mentioned 'genes' earlier in the discussion about business management; how do you perform analysis of 'genes'?" would be displayed on the interface unit 103 when a technical question is asked. In other words, when generating a set of multiple questions, this can be done during the interview using speech recognition, and these generated questions are extracted and displayed on the interface unit 103 when questions in the relevant field are asked.
[0014] The Question Management Unit 111 sets the photos and videos necessary for the interview. These photos and videos (hereinafter referred to as "images") are categorized and set according to the above-mentioned fields. The Question Management Unit 111 provides instructions for taking the necessary images. Replacing the questions with the subjects to be photographed, the Question Management Unit 111 performs the same settings and selections as above. For example, in the field of business management, the shooting instructions would be a photograph of the president of the company being interviewed and images of the company's offices; in the field of technology, images of the company's products themselves; and in the field of taxation, photographs of the accounting staff and tax advisors. The Question Management Unit 111 displays these instructions on the Interface Unit 103.
[0015] The question guidance unit 112 displays the next question that is preferable to ask to the interviewer (the administrator of the interview server 101, or a person commissioned by the administrator, etc.) via the interface unit 103. The question guidance unit 112 extracts relevant terms using speech recognition to understand the content of the current question, and processes the current question to display the next question that should be asked on the interface unit 103. This ensures that the interviewer can complete all the questions extracted by the question management unit 111. The question guidance unit 112 can either display the next questions to be asked on the interface unit 103 in order, or select the next question that is preferable to ask from the question content acquired by speech recognition and display it on the interface unit 103. In other words, since the flow of questions is important, the question guidance unit 112 selects the next question from the speech analysis so that it does not jump to a jump in the conversation.
[0016] Images are transmitted in real time from the camera of the mobile information terminal 2 in the interface unit 103, and the question guidance unit 112 analyzes the images and determines whether or not they are images related to the photography instruction. If there are insufficient images, it issues a photography instruction to the interface unit 103. In addition, unlike in the case of questions, the opportunity to take pictures on the spot is important, so instructions for the necessary images are given immediately. For example, when taking pictures during a factory tour, it is a waste of time to move to another location and then return to the original location, so an instruction is immediately given to the interface unit 103 to instruct on the next images. In the case of images, it is important not to miss the opportunity to take pictures, so instructions for taking images are given separately from the flow of questions described above. For example, in the case of questions, even while asking the interviewee president questions about management, instructions for taking pictures are given when passing in front of the products.
[0017] The checking unit 113 checks for missing questions and the quality of the interview. The checking unit 113 obtains the content of the questions from voice analysis and text keywords, compares it with the interview content pre-set by the question management unit 111, and checks for omissions. This result is displayed on the interface unit 103. The checking unit 113 checks the quality of the interview. It performs text recognition on the audio data of the entire interview content, compares the content of this text with the questions in relation to them, and confirms whether the necessary questions were asked. If the questions are as set by the question management unit 111 and there are no omissions, it is determined that the quality of the interview is maintained. This ensures the reliability of the interview.
[0018] Further, the content of images is grasped through image recognition, and it is confirmed whether appropriate images have been acquired. Images are stored in association with answers to questions. In the interview server 101, the acquired interview information is stored in association with each other, thereby improving accuracy as a resource for video generation. An answer to a management question is associated with a photograph of a company building or a video of a company president as related items. For image information obtained through analysis from the content of a photograph and character information obtained through speech recognition, items with high relevance are associated with each other according to a predetermined rule. For example, when a photograph of the president shows him wearing a suit, it is stored in association with an answer text to a management question. A photograph of a product is stored in association with an answer to a question in the technical field. The association may be in a spider-web shape, or may be in pairs.
[0019] The interview server 101 starts from setting and selecting interview questions, and finally checks the interview quality. This serves as a primary resource and determines the quality of the video to be generated next. The interview server 101 converts audio data into character data in a predetermined format. Alternatively, an interviewer or an administrator listens to the audio data and inputs the data through the input means of the interface unit 103. The resource is composed of images, sound, and character data. When the expression format of character data is automatically generated, AI may be instructed by a predetermined prompt to generate a text. For example, an instruction may be given to generate a text in an interactive format. The interview server 101 stores the information interviewed as described above as a resource in a memory (the storage location is not limited to the inside of the server, and may be a portable information terminal or any other storage), and transmits the information to the video generation server 102.
[0020] The video generation server 102 is connected to the interview server 101 via a network. Alternatively, the video generation server 102 may be integrated with the interview server 101. The server may be composed of a plurality of video computers 3. The video generation server 102 is composed of the video computer 3 and a predetermined program. The video generation server 102 acquires the interviewed resource from the interview server 101. This resource includes images, sound, and character data.
[0021] The video generation server 102 is a server equipped with artificial intelligence that newly generates videos from text and images. It may also be an online service provided by a third party. The video generation server 102 generates a video using the text data and image data generated by the coverage server 101. Instructions to the video generation server 102 are specifically given. Examples thereof include generating a video using only the text and images generated by the coverage server 101, and not changing the structure of the text.
[0022] The video generation server 102 may apply blockchain technology to the generated video. For example, blockchain transaction information may be stored in a metadata storage area of the video format. This makes it possible to determine the authenticity of the video.
[0023] DRM or watermarking may be employed as an unauthorized copy prevention technology.
[0024] The generated video is used for advertising purposes as a corporate coverage video. Since this coverage video is generated based on appropriate resources, its credibility and quality are high.
[0025] Text information and audio information of the generated video are translated into languages of various countries by automatic translation. The translation means uses a service provided by an external server. Since the generated text is for coverage and has a fixed pattern, it is suitable for self-learning and suitable for neural machine translation.
[0026] The generated video data is returned to the coverage server 101 and stored therein.
[0027] The translated video data can be supplemented with basic transaction information by the information-adding unit 114. When overseas companies view the video data, transaction-related information such as regulatory information and tariff information necessary for exports and service provision from Japan is displayed all at once. The interview server 101 adds this information. The video data with transaction-related information added is extracted as transaction-related information for countries where that language is the native or second language, depending on the language to which it is translated, and displayed on the interface unit 103. This makes it easier for overseas companies that view the video to determine the possibility of doing business with the subject of the interview in the video.
[0028] The external information acquisition unit 115 of the news gathering server 101 can acquire information from third parties. Since the credibility and quality of information from third parties vary considerably, it can be carefully reviewed and selected.
[0029] The external information acquisition unit 115 of the reporting server 101 has a program that implements crawler functionality. This crawler searches for related articles and images of the same type. Images on the same theme are collected using an image analysis program. The date and time of the collected articles, images, and other external information are obtained, and the earliest one is extracted as a resource. Third parties providing these extracted resources are considered highly reliable, and related information is also collected. The external information acquisition unit 115 saves the URL of the extracted resource.
[0030] The resource generation unit 116 accesses the information source based on the URL obtained by the external information acquisition unit 115, analyzes the content, and understands the context. This process is not displayed to the administrator on the interface unit 103 and is not disclosed to them. In this state, the resource generation unit 116 independently generates text based on the topics extracted from the understood context. This generated text is then displayed to the administrator for the first time. This generated text data can be used as a resource when generating videos.
[0031] For images, the subject of the extracted image is extracted through image analysis, a prompt corresponding to this subject is generated, and image generation based on this prompt is performed by AI. In other words, the system identifies what a specific product is from an image of that product and generates another image of the same product. This prevents the crawler from reproducing copyrighted images it has collected. The materials used for generation can be generated by the AI, or they can be materials distributed by the crawler under the condition that they are free to reproduce and can be used commercially. Processing is faster if these images are collected by the crawler in advance.
[0032] The external information acquisition unit 115 extracts the structure of the external information. It determines whether it is in a dialogue format, what the topic is, and whether the balance between questions and answers is appropriate, and extracts external information with a structure that meets pre-set criteria, and determines that it is suitable as a resource. Using this resource, it extracts the subject according to the above prompt and generates a text independently. By analyzing the structure of such external information, it determines the information quality, such as the credibility of the external information.
[0033] Next, when conducting interviews online, the interviewee will hold the personal information terminal 2 and interact with the interviewer, but the interviewing method in this case is almost the same as when conducting interviews in person as described above. Since it is difficult to take photos freely in an online setting, the interviewer will be asked to take photos using the personal information terminal 2. Switching to shooting mode is done online by the interview server 101 operating the personal information terminal 2. A remote control program for interviews is installed on the personal information terminal 2.
[0034] The program works in conjunction with the hardware of the mobile information terminal 2 to configure the following functions. Figure 3 shows a block diagram of the online function. It includes an audio recording unit 121, an image capture unit 122, a video recording unit 123, and an online communication unit 124. The online communication unit 124 allows the interviewer to be displayed in real time via the camera and engage in dialogue. In this process, the content instructed by the question management unit 111, the question guidance unit 112, and the check unit 113 is not displayed on the interviewer's mobile information terminal 2, but is only displayed on the interviewer's screen.
[0035] On the screen of the subject's mobile device 2, the interviewer's questions are recognized by voice and displayed as text. Shooting instructions are given remotely by voice through the speaker on the mobile device 2. Captured images are sent to the interview server 101.
[0036] In online interviews, a third-party staff member can conduct the interview with the interviewer. In this case, the same screen as the interviewer is shared on the staff member's mobile device 2. Since the staff member is unfamiliar with the process, a chat window appears to receive instructions from the interviewer.
[0037] To indicate the source of the image, reporters and dispatched staff can be instructed to ensure that name tags or armbands bearing the "Kanagawa Keizai Shimbun" text or logo are included in the image. A QR code is attached next to the text on these name tags or armbands, and if this QR code is not included in the image, a notification is sent to the mobile information terminal 2 at the time of shooting. This allows the original reporter to be identified even if the image is reproduced without permission.
[0038] (Embodiment 2) Figure 4 is a configuration diagram showing a news video generation system according to Embodiment 2 of the present invention. In this news video generation system 200, a news robot 201 is built into the system, and this news robot 201 autonomously conducts interviews with the subject of the interview. The system in which the news robot 201 is built is the system described in Embodiment 1 above.
[0039] The interview robot 201 is configured by installing a predetermined program on the interview server 101 or the mobile information terminal 2. The interview robot 201 has the parts described in Embodiment 1 and a function to correct answers. The correction unit 221 analyzes the answers of the interviewee via voice and makes a determination based on whether the content contains terms that would be commonly used as answers to questions. Alternatively, the AI may be used to make judgments from multiple perspectives.
[0040] The process of revising a question involves asking a follow-up question. This follow-up question is pre-set and rephrased as the same question. Since asking the same question repeatedly can be unpleasant for the interviewee, the question is structured from a different angle to ensure it includes the desired answer.
[0041] The biometric information detection unit 222 wirelessly connects the camera or smartwatch of the mobile information terminal 2 to the interview robot 201 to acquire biometric information of the interviewee, such as body temperature and heart rate. During the interview, the interviewee may become nervous, causing fluctuations in body temperature and heart rate. The interview robot 201 acquires this information and uses it to schedule breaks or change the questions. In other words, the interview robot 201 changes the content and timing of questions based on biometric information. The interview robot 201 also uses biometric information to determine the credibility of the answers to the questions.
[0042] The camera on the portable information terminal 2 measures body temperature by detecting infrared radiation, but a separate infrared camera could also be installed. Alternatively, the camera could be used to detect changes in blood flow by having the subject place their finger on it. This would allow the interview to be conducted according to the interviewer's condition, thereby increasing the reliability of the interview.
[0043] The interview robot 201 asks the pre-set questions in order. The interface unit 103 displays an image of the robot's character. The interface unit 103 also displays the operation screen along with the character. For example, a button to start the interview is displayed. When the interviewee presses this button, a question is asked, specifically, the question is spoken aloud from the speaker, and the interviewee answers. The answer is recorded. The answer is also analyzed by voice analysis, converted into text information, and saved. When the interview robot determines that one answer has been completed, it prompts for the next question in the same way, speaking aloud from the speaker. This process is repeated until the interview is complete.
[0044] Furthermore, during the questioning, interview robot 201 prompts the interviewer to take a picture with its camera. The interviewer takes a picture of the subject according to the instructions of interview robot 201. The captured image is sent to and stored on interview server 101. Interview server 101 also analyzes the captured image to determine whether or not an appropriate image has been taken.
[0045] In summary, according to the interview video generation system 200 of this invention, realistic interviews can be conducted even without staff present, and the videos generated from this interview information are highly credible. [Explanation of symbols]
[0046] 100 Interview Video Generation System 101 Interview Server 102 Video Generation Server 103 Interface section 111 Question Management Department 112 Question Guidance Section 113 Check section
Claims
1. This involves supplying interview resources, including text, audio, or images, to a video generation server and having the server generate the interview video. User interface means, A question management means sets questions for the subject of the interview, displays the questions in a predetermined order on the interface means, and acquires either text, audio, or images. A news gathering server.
2. The interview server according to claim 1, wherein the question management means gives instructions for taking images through the interface means.
3. Furthermore, the interview server according to claim 1, which transmits the interview resources obtained based on the above questions to the video server.
4. A video reporting system comprising: an interview server according to any one of claims 1 to 3; and a video reporting server connected to the network and which generates interview videos based on interview resources obtained based on the aforementioned questions.
Citation Information
Patent Citations
Information processing system
JP2025000471A