Content evaluation device, content evaluation method, program, and storage medium

The content evaluation device addresses the lack of feedback in push-type content output by quantitatively assessing the effectiveness of content through voice recognition of keywords and exclamations in passenger speech, enhancing user reaction evaluation in vehicle environments.

JP7751066B2Active Publication Date: 2025-10-07PIONEER IP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024503302
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-28
Filing Date
2023-02-28
Publication Date
2025-10-07
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

Existing push-type content output technologies cannot evaluate the effectiveness of the content output to users, particularly in vehicle environments, as they lack feedback mechanisms to assess user reactions.

Method used

A content evaluation device that includes a content acquisition unit, an output unit, a voice recognition unit, and an evaluation unit to recognize keywords and exclamations in passenger speech after content output, quantifying the effectiveness based on the number of recognitions and passenger count.

Benefits of technology

Enables quantitative evaluation of the content's emotional impact on vehicle occupants by recognizing keywords and exclamations, providing a score that reflects the content's effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007751066000001
    Figure 0007751066000001
  • Figure 0007751066000002
    Figure 0007751066000002
  • Figure 0007751066000003
    Figure 0007751066000003
Patent Text Reader

Abstract

This content evaluation device comprises a content acquisition unit, an output unit, a speech recognition unit, and an evaluation unit. The content acquisition unit acquires speech content to be output to a passenger of a vehicle. The output unit outputs the speech content. The speech recognition unit performs a speech recognition process for recognizing predetermined words included in a passenger's utterance after the speech content is output. The evaluation unit evaluates the effectiveness of the speech content output to the passenger, on the basis of the result of the speech recognition process.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique that can be used in the evaluation of push-type content. [Background technology]

[0002] 2. Description of the Related Art There is a known push-type content output technology that outputs content corresponding to various information obtained through sensors or the like to a user without a request from the user, based on the information.

[0003] Specifically, for example, Patent Document 1 discloses a technology for outputting a greeting voice when passengers get on and off a vehicle based on information obtained through a vibration sensor that detects the opening and closing of the vehicle door. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-237453 Summary of the Invention [Problem to be solved by the invention]

[0005] However, with push-type content output, it is not possible to obtain feedback indicating the user's reaction, and therefore it is not possible to evaluate the effectiveness of the content output to the user.

[0006] Therefore, for example, when push-type content output is applied while driving a vehicle, a problem may arise in that it is not possible to evaluate the effectiveness of the content output to the passengers of the vehicle.

[0007] However, Patent Document 1 does not particularly disclose a method capable of solving the above-mentioned problems. Therefore, according to the configuration disclosed in Patent Document 1, problems corresponding to the above-mentioned problems still exist.

[0008] The present invention has been made to solve the above-mentioned problems, and its main object is to provide a content evaluation device that can evaluate the effectiveness of content output to vehicle occupants in push-type content output. [Means for solving the problem]

[0009] The claimed invention is a content evaluation device that outputs information to a vehicle occupant. a content acquisition unit that acquires audio content for the audio recording; and an output unit that outputs the audio content. a sound output unit, and a sound output unit for outputting the sound included in the speech of the passenger after the audio content is output. Keywords in voice content and an exclamation expressing admiration for the audio content. Recognize a speech recognition unit that performs speech recognition processing for The keyword and and / or obtain a score according to the number of times the exclamation is recognized, and Based on this, the degree of influence that the audio content had on the emotions of the passengers is quantitatively estimated. and evaluate the effectiveness of the audio content. and an evaluation unit. The claimed invention is a content evaluation device, which outputs a rating to a vehicle occupant. a content acquisition unit for acquiring audio content to be input; and an output unit for outputting the audio content. an output unit for outputting the audio content, and Keywords in the audio content and exclamations expressing admiration for the audio content a voice recognition unit that performs a voice recognition process for recognizing the voice content; As an index for determining whether or not the number of passengers is a passenger, the keywords and and / or an evaluation unit that acquires a score according to the number of times the exclamation is recognized, The evaluation unit is configured to evaluate whether the passengers are a plurality of passengers and the keywords and / or the exclamations are If the recognition is performed a first number of times, a first score corresponding to the first number of times is obtained; If there is one passenger and the keyword and / or the exclamation is recognized the first number of times, If so, a second score greater than the first score is acquired.

[0010] The claimed invention is a content evaluation method executed by a computer. and acquiring audio content to be output to the passengers of the vehicle, and and outputting the audio content before the audio content is included in the passenger's speech after the audio content is output. Keywords in the audio content and an exclamation expressing admiration for the audio content. Recognize It performs voice recognition processing to identify the The keyword and / or A score is obtained according to the number of times the exclamation is recognized, and based on the obtained score, Quantitatively estimating the degree of influence that the audio content has had on the emotions of the passengers. The effectiveness of the audio content is also evaluated. The claimed invention is a content evaluation method executed by a computer. and acquiring audio content to be output to the passengers of the vehicle, and and outputting the audio content before the audio content is included in the passenger's speech after the audio content is output. Recognizing keywords in the audio content and exclamations that express admiration for the audio content The speech recognition process is performed to identify the speech and use it as an index for evaluating the effectiveness of the speech content. Then, the number of passengers and the keyword and / or the feeling are recognized by the voice recognition processing. A score is obtained according to the number of times the exclamation is recognized, and when the passengers are multiple and If the keyword and / or the exclamation is recognized a first number of times, obtain a corresponding first score, and If the exclamation is recognized the first number of times, a second score greater than the first score is given. Get the score.

[0011] The claimed invention is also implemented by a content evaluation device equipped with a computer. a program for acquiring audio content to be output to a vehicle occupant, a content acquisition unit for acquiring the audio content; an output unit for outputting the audio content; Keywords in the voice content included in the passenger's speech after and before Exclamation to express admiration for the audio content Speech recognition is a process for recognizing speech. Knowledge section, and the keyword and / or the exclamation are recognized by the speech recognition process. and obtain a score according to the number of times the audio content has been played, and The degree of influence that the voice control has had on the emotions of the passengers is quantitatively estimated, and Evaluate the effectiveness of the content The computer functions as an evaluation unit. The claimed invention is also implemented by a content evaluation device equipped with a computer. a program for acquiring audio content to be output to a vehicle occupant, a content acquisition unit for acquiring the audio content; an output unit for outputting the audio content; Keywords in the voice content included in the passenger's speech after the A speech recognition system that performs speech recognition processing to recognize exclamations that express appreciation for the audio content. and a recognition unit for the passengers as an index for evaluating the effectiveness of the audio content. the number of people and the number of times the keyword and / or the exclamation were recognized by the speech recognition process. and causing the computer to function as an evaluation unit that acquires a score according to the number of is a number in which the passengers are plural and the keyword and / or the exclamation is / are a first number of times. If the passenger is recognized, a first score corresponding to the first number of times is obtained, and the passenger is recognized as a first If the person is a person and the keyword and / or the exclamation is recognized the first number of times, obtains a second score that is greater than the first score. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an audio output system according to an embodiment. [Figure 2] FIG. 1 is a block diagram showing a schematic configuration of an audio output device. [Figure 3] FIG. 2 is a diagram showing an example of a schematic configuration of a server device. [Figure 4] 10 is a flowchart illustrating processing performed in the server device. DETAILED DESCRIPTION OF THE INVENTION

[0013] In one preferred embodiment of the present invention, a content evaluation device includes a content acquisition unit that acquires audio content to be output to a vehicle occupant, an output unit that outputs the audio content, a voice recognition unit that performs voice recognition processing to recognize predetermined words included in the occupant's speech after the audio content has been output, and an evaluation unit that evaluates the effectiveness of the audio content output to the occupant based on the results of the voice recognition processing.

[0014] The content evaluation device includes a content acquisition unit, an output unit, a voice recognition unit, and an evaluation unit. The content acquisition unit acquires voice content to be output to a vehicle occupant. The output unit outputs the voice content. The voice recognition unit performs voice recognition processing to recognize predetermined words included in the occupant's speech after the voice content is output. The evaluation unit evaluates the effectiveness of the voice content output to the occupant based on the result of the voice recognition processing. This makes it possible to evaluate the effectiveness of the content output to the vehicle occupant in push-type content output.

[0015] In one aspect of the above content evaluation device, the evaluation unit obtains a score corresponding to the number of times the specified phrase is recognized by the voice recognition process as an index for evaluating the effectiveness of the audio content.

[0016] In one aspect of the content evaluation device, the voice recognition unit performs the voice recognition process from immediately after the audio content is output until a predetermined time has elapsed.

[0017] In one aspect of the content evaluation device, the voice recognition unit stops the voice recognition process when a predetermined time has elapsed since the score was last obtained after the voice content was output.

[0018] In one aspect of the content evaluation device, the speech recognition unit recognizes, as the predetermined phrase, at least one of a phrase expressing admiration for the audio content and a keyword in the audio content.

[0019] In another embodiment of the present invention, a content evaluation method includes acquiring audio content to be output to a vehicle occupant, outputting the audio content, performing a speech recognition process to recognize predetermined words included in an utterance of the occupant after the audio content has been output, and evaluating the effectiveness of the audio content output to the occupant based on the results of the speech recognition process. This makes it possible to evaluate the effectiveness of the content output to the vehicle occupant in a push-type content output.

[0020] In yet another embodiment of the present invention, a program executed by a content evaluation device including a computer causes the computer to function as a content acquisition unit that acquires audio content to be output to a vehicle occupant, an output unit that outputs the audio content, a voice recognition unit that performs voice recognition processing to recognize predetermined words included in the occupant's speech after the audio content has been output, and an evaluation unit that evaluates the effectiveness of the audio content output to the occupant based on the results of the voice recognition processing. The content evaluation device described above can be realized by executing this program on a computer. This program can be stored in a storage medium and used. [Example]

[0021] Preferred embodiments of the present invention will now be described with reference to the drawings.

[0022] [System Configuration] (Overall composition) 1 is a diagram illustrating an example of the configuration of an audio output system according to an embodiment. The audio output system 1 according to the embodiment includes an audio output device 100 and a server device 200. The audio output device 100 is mounted on a vehicle Ve. The server device 200 communicates with a plurality of audio output devices 100 mounted on a plurality of vehicles Ve.

[0023] The audio output device 100 basically performs route search processing, route guidance processing, and the like for a user who is a passenger in a vehicle Ve. For example, when a destination or the like is input by the user, the audio output device 100 transmits an upload signal S1 including location information of the vehicle Ve and information about the specified destination to the server device 200. The server device 200 calculates a route to the destination by referring to map data, and transmits a control signal S2 indicating the route to the destination to the audio output device 100. The audio output device 100 provides route guidance to the user by audio output based on the received control signal S2.

[0024] Furthermore, the audio output device 100 provides various types of information to the user through dialogue with the user. For example, when the user makes an information request, the audio output device 100 supplies the server device 200 with an upload signal S1 including information indicating the content or type of the information request and information regarding the running state of the vehicle Ve. The server device 200 acquires and generates the information requested by the user and transmits it to the audio output device 100 as a control signal S2. The audio output device 100 provides the received information to the user by audio output.

[0025] (Audio output device) The audio output device 100 travels with the vehicle Ve and provides route guidance primarily through audio so that the vehicle Ve travels along the guidance route. Note that "route guidance primarily through audio" refers to route guidance that allows the user to understand information necessary for driving the vehicle Ve along the guidance route at least through audio alone, and does not exclude the audio output device 100 supplementarily displaying a map or the like around the current location. In this embodiment, the audio output device 100 outputs various driving-related information by audio, such as at least points on the route where guidance is required (also referred to as "guidance points"). Herein, guidance points include, for example, intersections where the vehicle Ve must turn right or left, and other important passing points for the vehicle Ve to travel along the guidance route. The audio output device 100 provides audio guidance regarding guidance points, such as the distance from the vehicle Ve to the next guidance point and the direction of travel at that guidance point. Hereinafter, audio guidance regarding the guidance route will also be referred to as "route audio guidance."

[0026] The audio output device 100 is attached, for example, to the top of the windshield or on the dashboard of the vehicle Ve. The audio output device 100 may also be incorporated into the vehicle Ve.

[0027] 2 is a block diagram showing a schematic configuration of the audio output device 100. The audio output device 100 mainly includes a communication unit 111, a storage unit 112, an input unit 113, a control unit 114, a sensor group 115, a display unit 116, a microphone 117, a speaker 118, an exterior camera 119, and an interior camera 120. The elements within the audio output device 100 are connected to each other via a bus line 110.

[0028] The communication unit 111 performs data communication with the server device 200 under the control of the control unit 114. The communication unit 111 may receive, for example, map data for updating a map DB (DataBase) 4 (described later) from the server device 200.

[0029] The storage unit 112 is configured with various types of memory such as a RAM (Random Access Memory), a ROM (Read Only Memory), and a non-volatile memory (including a hard disk drive, a flash memory, etc.). The storage unit 112 stores programs for the audio output device 100 to execute predetermined processes. The above-mentioned programs may include an application program for providing route guidance by voice, an application program for playing music, an application program for outputting content other than music (such as television), etc. The storage unit 112 is also used as a working memory for the control unit 114. The programs executed by the audio output device 100 may be stored in a storage medium other than the storage unit 112.

[0030] The storage unit 112 also stores a map database (hereinafter, database will be referred to as "DB") 4. The map DB 4 stores various data necessary for route guidance. The map DB 4 stores, for example, road data that represents a road network using a combination of nodes and links, and facility data that indicates facilities that are candidates for destinations, stop-off points, or landmarks. The map DB 4 may be updated based on map information that the communication unit 111 receives from a map management server under the control of the control unit 114.

[0031] The input unit 113 is a button, a touch panel, a remote controller, or the like that is operated by the user. The display unit 116 is a display or the like that displays information under the control of the control unit 114. The microphone 117 collects sounds inside the vehicle Ve, particularly sounds uttered by the driver. The speaker 118 outputs audio for route guidance to the driver or the like.

[0032] The sensor group 115 includes an external sensor 121 and an internal sensor 122. The external sensor 121 is one or more sensors for recognizing the surrounding environment of the vehicle Ve, such as a lidar, radar, ultrasonic sensor, infrared sensor, or sonar. The internal sensor 122 is a sensor for measuring the position of the vehicle Ve, such as a Global Navigation Satellite System (GNSS) receiver, a gyro sensor, an Inertial Measurement Unit (IMU), a vehicle speed sensor, or a combination thereof. Note that the sensor group 115 may include any sensor that allows the control unit 114 to directly or indirectly (i.e., by performing estimation processing) derive the position of the vehicle Ve from the output of the sensor group 115.

[0033] Exterior camera 119 is a camera that captures the outside of vehicle Ve. Exterior camera 119 may be only a front camera that captures the view in front of the vehicle, or may include a rear camera that captures the view behind the vehicle in addition to the front camera, or may be an omnidirectional camera that can capture the entire periphery of vehicle Ve. On the other hand, interior camera 120 is a camera that captures the interior of vehicle Ve, and is installed in a position that can capture at least the area around the driver's seat.

[0034] The control unit 114 includes a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc., and controls the entire audio output device 100. For example, the control unit 114 estimates the position (including the direction of travel) of the vehicle Ve based on the output of one or more sensors in the sensor group 115. When a destination is specified by the input unit 113 or the microphone 117, the control unit 114 generates route information indicating a guidance route that is a route to the destination, and provides route guidance based on the route information, the estimated position information of the vehicle Ve, and the map DB 4. In this case, the control unit 114 outputs route voice guidance from the speaker 118. The control unit 114 also controls the display unit 116 to display information on the music being played, video content, a map of the area around the current location, etc.

[0035] The processing performed by the control unit 114 is not limited to being realized by software programs, but may be realized by any combination of hardware, firmware, and software. The processing performed by the control unit 114 may also be realized by a user-programmable integrated circuit, such as an FPGA (field-programmable gate array) or a microcomputer. In this case, the program executed by the control unit 114 in this embodiment may be realized by using this integrated circuit. In this way, the control unit 114 may be realized by hardware other than a processor.

[0036] The configuration of the audio output device 100 shown in FIG. 2 is an example, and various modifications may be made to the configuration shown in FIG. 2. For example, instead of storing the map DB4 in the memory unit 112, the control unit 114 may receive information necessary for route guidance from the server device 200 via the communication unit 111. In another example, instead of including the speaker 118, the audio output device 100 may be connected electrically or via a known communication means to an audio output unit configured separately from the audio output device 100, thereby outputting audio from the audio output unit. In this case, the audio output unit may be a speaker provided in the vehicle Ve. In yet another example, the audio output device 100 may not include the display unit 116. In this case, the audio output device 100 may not perform any control related to the display, and may be electrically connected to a display unit provided in the vehicle Ve, etc., via a wired or wireless connection, thereby causing the display unit to execute a predetermined display. Similarly, instead of including the sensor group 115, the audio output device 100 may acquire information output by sensors provided in the vehicle Ve from the vehicle Ve based on a communication protocol such as CAN (Controller Area Network).

[0037] (Server device) The server device 200 generates route information indicating a guide route along which the vehicle Ve should travel, based on an upload signal S1 including a destination and the like received from the audio output device 100. The server device 200 then generates a control signal S2 related to information output in response to the user's information request, based on the user's information request and the traveling state of the vehicle Ve indicated in the upload signal S1 transmitted thereafter by the audio output device 100. The server device 200 then transmits the generated control signal S2 to the audio output device 100.

[0038] Furthermore, the server device 200 generates content for providing information to the user of the vehicle Ve and for dialogue with the user, and transmits the content to the audio output device 100. The provision of information to the user is mainly push-type information provision that is initiated from the server device 200 side when triggered by the vehicle Ve entering a predetermined driving situation. Furthermore, the dialogue with the user is basically pull-type dialogue that is initiated by a question or inquiry from the user. However, the dialogue with the user may also be initiated by push-type content provision.

[0039] 3 is a diagram showing an example of a schematic configuration of the server device 200. The server device 200 mainly includes a communication unit 211, a storage unit 212, and a control unit 214. The elements within the server device 200 are connected to each other via a bus line 210.

[0040] The communication unit 211 performs data communication with external devices such as the audio output device 100 under the control of the control unit 214. The storage unit 212 is configured with various types of memory such as RAM, ROM, and non-volatile memory (including a hard disk drive, flash memory, etc.). The storage unit 212 stores programs for the server device 200 to execute predetermined processes. The storage unit 212 also includes a map DB4.

[0041] The control unit 214 includes a CPU, a GPU, and the like, and controls the entire server device 200. The control unit 214 also executes programs stored in the storage unit 212 to operate together with the audio output device 100 and perform processes such as route guidance and information provision for the user. For example, the control unit 214 generates route information indicating a guidance route or a control signal S2 related to information output in response to an information request from the user, based on an upload signal S1 received from the audio output device 100 via the communication unit 211. The control unit 214 then transmits the generated control signal S2 to the audio output device 100 via the communication unit 211.

[0042] The control unit 214 has a voice recognition engine 214a for recognizing the content of speech by a passenger of the vehicle Ve based on the speech included in the driving condition information received from the voice output device 100 via the communication unit 211. The control unit 214 also performs a voice recognition process for recognizing a predetermined phrase included in the speech of the passenger of the vehicle Ve by activating the voice recognition engine 214a in a situation described below. The control unit 214 also acquires a score according to the number of times the predetermined phrase is recognized.

[0043] [Push-type content provision] Next, push-type content provision will be described. Push-type content provision refers to the audio output device 100 audibly outputting content related to a driving situation of the vehicle Ve to the user when the vehicle Ve is in a predetermined driving situation. Specifically, the audio output device 100 acquires driving situation information indicating the driving situation of the vehicle Ve based on the output of the sensor group 115 as described above, and transmits the information to the server device 200. The server device 200 stores table data for performing push-type content provision in the storage unit 212. The server device 200 refers to the table data, and when the driving situation information received from the audio output device 100 installed in the vehicle Ve matches a trigger condition defined in the table data, the server device 200 generates output content using a script corresponding to the trigger condition and transmits the generated content to the audio output device 100. The audio output device 100 audibly outputs the output content received from the server device 200. In this way, content corresponding to the driving situation of the vehicle Ve is audibly output to the user.

[0044] The driving situation information may include at least one piece of information that can be acquired based on the functions of each unit of the audio output device 100, such as the position of the vehicle Ve, the direction of the vehicle, traffic information around the position of the vehicle Ve (including speed limits and congestion information), the current time, the destination, etc. The driving situation information may also include any of audio acquired by the microphone 117, an image captured by the external camera 119, and an image captured by the internal camera 120. The driving situation information may also include information received from the server device 200 via the communication unit 111.

[0045] [Processing related to evaluation of push-type content] Next, a process related to evaluation of push-type content will be described.

[0046] (Example) The server device 200 acquires audio content VC to be output to the passengers of the vehicle Ve based on the driving status information of the vehicle Ve received from the audio output device 100, and outputs (transmits) the acquired audio content VC to the audio output device 100.

[0047] The audio content VC includes trigger content VCT, dynamic content VCD, and static content VCS.

[0048] The trigger content VCT is configured as content linked to a trigger condition such as the current location of the vehicle Ve. Specifically, the trigger content VCT is configured as a script SCT such as, for example, "The vehicle has entered Kawajima-cho, Hiki-gun from Kawagoe-shi."

[0049] The dynamic content VCD is configured as content including a variable portion that changes according to the driving situation of the vehicle Ve. Specifically, the dynamic content VCD is configured as a script SCD such as, for example, "The driving time in Kawagoe City was X minutes." The "X minutes" included in the script SCD corresponds to the variable portion that changes according to the driving time of the vehicle Ve.

[0050] The static content VCS is configured as content including at least one keyword linked to the trigger content VCT. Specifically, the static content VCS is configured as a script SCS such as, for example, "Strawberries are a specialty of Kawajima Town, Hiki District." The "strawberry" included in the script SCS corresponds to the keyword linked to "Kawajima Town, Hiki District" included in the script SCT. Note that in this embodiment, for example, if multiple keywords are linked to one trigger content VCT, at least one keyword to be incorporated into the static content VCS may be selected from the multiple keywords. Furthermore, in this embodiment, for example, by setting the portion of the script SCS other than "strawberry" as a fixed phrase and incorporating a keyword other than "strawberry" into the fixed phrase, a script different from the script SCS can be generated.

[0051] Hereinafter, an example will be described in which audio content VC including scripts SCT, SCD, and SCS is output to a passenger in a vehicle Ve.

[0052] The server device 200 activates the voice recognition engine 214a during the period from immediately after outputting (transmitting) the audio content VC to the audio output device 100 until a predetermined time TP has elapsed, thereby performing a voice recognition process for recognizing predetermined words included in the speech of the passenger of the vehicle Ve after the audio content VC has been output. Specifically, the server device 200 uses the voice recognition engine 214a to perform a voice recognition process for recognizing at least one of a word indicating an exclamation in response to the audio content VC and a keyword in the audio content VC as the predetermined words. Note that, in this embodiment, the speech content of the passenger of the vehicle Ve may be identified based on the voice included in the driving situation information of the vehicle Ve. Also, in this embodiment, the predetermined time TP may be set to, for example, 30 seconds. Furthermore, the predetermined words may be words that indicate the reaction of the passenger of the vehicle Ve to the audio content VC. Furthermore, hereinafter, unless otherwise specified, the description will be given assuming that both a word indicating an exclamation in response to the audio content VC and a keyword in the audio content VC are recognized as the predetermined words.

[0053] According to the above-mentioned voice recognition process, it is possible to recognize exclamations such as "Wow" and "Hmm" as expressions expressing admiration for the voice content VC. Furthermore, according to the above-mentioned voice recognition process, it is possible to recognize "strawberry" as a keyword in the voice content VC.

[0054] The server device 200 acquires a score SR corresponding to the number of times a predetermined phrase is recognized through speech recognition processing using the speech recognition engine 214a. Specifically, for example, when two passengers in a vehicle Ve have a conversation such as "Wow, they grow a lot of strawberries around here," "I see. Shall we buy some strawberries as a souvenir?", and "Yeah, that's right. Then, let's stop by if there's a strawberry farmer's market!", the server device 200 acquires 5 points as the score SR corresponding to the number of times (5 times) the predetermined phrase is recognized. Also, for example, when two passengers in a vehicle Ve have a conversation such as "They grow a lot of strawberries around here," and "That seems so," the server device 200 acquires 1 point as the score SR corresponding to the number of times (1 time) the predetermined phrase is recognized.

[0055] According to this embodiment, the server device 200 may change the score SR depending on the number of passengers in the vehicle Ve. Specifically, for example, when the predetermined phrase is recognized Y times, the server device 200 may acquire Y points as the score SR if there are two or more passengers in the vehicle Ve, and may acquire Z points, which is more than Y points, if there is only one passenger in the vehicle Ve.

[0056] The server device 200 stops the speech recognition process by the speech recognition engine 214a when a predetermined time TP has elapsed since the last time the score SR was acquired after the audio contents VC were output. In other words, if the server device 200 is able to acquire the score SR during the predetermined time TP from immediately after outputting (transmitting) the audio contents VC to the audio output device 100, the server device 200 continues the speech recognition process from the last time the score SR was acquired until the predetermined time TP has elapsed again. Note that if the server device 200 is unable to acquire the score SR during the predetermined time TP from immediately after outputting (transmitting) the audio contents VC to the audio output device 100, the server device 200 stops the speech recognition process at that time.

[0057] The server device 200 evaluates the effectiveness of the audio content VC output to the passenger of the vehicle Ve based on the score SR obtained from the start to the stop of the audio recognition process using the audio recognition engine 214a. Specifically, for example, if the score SR is relatively low, the server device 200 evaluates that the audio content VC is not effective for the passenger of the vehicle Ve. On the other hand, for example, if the score SR is relatively high, the server device 200 evaluates that the audio content VC is highly effective for the passenger of the vehicle Ve.

[0058] According to the above-described process, the number of times a predetermined phrase is recognized by the speech recognition process corresponds to the score SR obtained from the start to the end of the speech recognition process. According to the above-described process, the number of times a predetermined phrase is recognized by the speech recognition process can be rephrased as the result of the speech recognition process. Therefore, the server device 200 of this embodiment can evaluate the effectiveness of the audio content VC output to an occupant of the vehicle Ve based on the result of the speech recognition process for recognizing the predetermined phrase contained in the speech of the occupant. According to the above-described process, the server device 200 can obtain a score corresponding to the number of times a predetermined phrase is recognized by the speech recognition process as an index for evaluating the effectiveness of the audio content VC output to the occupant of the vehicle Ve. According to the above-described process, for example, the degree to which the audio content VC affected the emotions of the occupant of the vehicle Ve can be quantitatively estimated based on the score SR obtained in response to the output of the audio content VC. Furthermore, according to the processing described above, only specific words contained in the speech of the occupant of the vehicle Ve are recognized by the voice recognition processing, and the voice recognition processing is performed within a limited period corresponding to the specified time TP, thereby protecting the privacy of the occupant.

[0059] (Processing flow) Next, a description will be given of the processing performed in the server device 200. Fig. 4 is a flowchart for explaining the processing performed in the server device.

[0060] First, the control unit 214 of the server device 200 acquires the driving situation information of the vehicle Ve received from the audio output device 100 (step S11).

[0061] Next, the control unit 214 acquires audio content VC according to the driving situation information acquired in step S11, and outputs (transmits) the acquired audio content VC to the audio output device 100 (step S12).

[0062] Immediately after step S12, the control unit 214 starts a voice recognition process for recognizing predetermined words included in the speech of the occupant of the vehicle Ve that can be identified based on the driving situation information of the vehicle Ve. The control unit 214 also performs the voice recognition process (step S13) from immediately after step S12 or immediately after step S15 described below until a predetermined time TP has elapsed.

[0063] The control unit 214 determines whether or not a predetermined phrase included in the speech of the passenger of the vehicle Ve has been recognized through the speech recognition process of step S13 (step S14).

[0064] If the control unit 214 recognizes the predetermined phrase included in the speech of the passenger of the vehicle Ve (step S14: YES), the control unit 214 acquires the score SR (step S15). Then, the control unit 214 returns to step S13 and performs the speech recognition process from immediately after step S15 until the predetermined time TP has elapsed.

[0065] If the control unit 214 fails to recognize a predetermined phrase included in the speech of the passenger of the vehicle Ve through the speech recognition process in step S13 (step S14: NO), the control unit 214 stops the speech recognition process (step S16).

[0066] The control unit 214 evaluates the validity of the audio content VC output to the passenger of the vehicle Ve based on the score SR obtained immediately after step S12 and immediately before step S16 (step S17).

[0067] According to this embodiment, the control unit 214 has functions as a content acquisition unit, a voice recognition unit, and an evaluation unit. Also, according to this embodiment, the communication unit 211 has a function as an output unit.

[0068] As described above, according to this embodiment, the effectiveness of the audio content VC can be evaluated based on the result of recognizing a predetermined phrase included in the speech of the occupant of the vehicle Ve after the audio content VC is output. In other words, according to this embodiment, in push-type content output, the effectiveness of the content output to the occupant of the vehicle can be evaluated.

[0069] According to this embodiment, for example, when the communication unit 111 or the control unit 114 has a function as a content acquisition unit, the control unit 114 has a function as a voice recognition unit and an evaluation unit, and the speaker 118 has a function as an output unit, a process substantially similar to the series of processes in FIG. 4 can be performed in the audio output device 100.

[0070] In the above-described embodiments, the program can be stored using various types of non-transitory computer-readable media and supplied to a control unit, such as a computer. The non-transitory computer-readable media includes various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)).

[0071] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications within the scope of the present invention that would be understood by those skilled in the art can be made to the configuration and details of the present invention. In other words, the present invention naturally includes various modifications and alterations that would be possible for those skilled in the art based on the entire disclosure, including the claims, and the technical ideas. Furthermore, the disclosures of the above-cited patent documents and other documents are incorporated herein by reference. [Explanation of symbols]

[0072] 100 Audio output device 200 Server device 111, 211 Communications Department 112, 212 Storage section 113 Input section 114, 214 Control unit 115 Sensor Group 116 Display section 117 Mike 118 Speaker 119 Exterior Camera 120 In-car camera

Claims

1. a content acquisition unit that acquires audio content to be output to a vehicle occupant; an output unit that outputs the audio content; The voice content included in the passenger's speech after the voice content is output to recognize keywords in the audio content and exclamations indicating appreciation for the audio content; a speech recognition unit that performs speech recognition processing; Depending on the number of times the keyword and / or the exclamation is recognized by the speech recognition process, and based on the score obtained, the audio content is The degree of impact on the emotions of the users is quantitatively estimated, and the effectiveness of the audio content is evaluated. an evaluation unit that evaluates the sex; A content evaluation device having:

2. A content acquisition unit that acquires audio content to be output to a vehicle occupant; an output unit that outputs the audio content; The voice content included in the passenger's speech after the voice content is output to recognize keywords in the audio content and exclamations indicating appreciation for the audio content; a speech recognition unit that performs speech recognition processing; As an index for evaluating the effectiveness of the audio content, the number of passengers and and the number of times the keyword and / or the exclamation is recognized by the speech recognition process. an evaluation unit for obtaining a score; and The evaluation unit is configured to evaluate whether the passengers are a plurality of passengers and whether the keywords and / or the exclamations are If the character is recognized a first number of times, a first score corresponding to the first number of times is obtained, and the previous The passenger is one and the keyword and / or the exclamation is recognized the first number of times. If the content is evaluated, the content evaluation device obtains a second score that is greater than the first score. Place.

3. The voice recognition unit continues to output the voice content until a predetermined time has elapsed from immediately after the voice content is output.

3. The content evaluation device according to claim 1, wherein the speech recognition process is performed between 。

4. The speech recognition unit may select a score that is the last score obtained after the speech content is output. and stopping the speech recognition process when a predetermined time has elapsed since the timing of the acquisition.

3. The content evaluation device according to claim 1 or 2.

5. 1. A computer-implemented method for rating content, comprising: Obtaining audio content to be output to a vehicle occupant; outputting the audio content; The voice content included in the passenger's speech after the voice content is output to recognize keywords in the audio content and exclamations indicating appreciation for the audio content; Performs voice recognition processing, Depending on the number of times the keyword and / or the exclamation is recognized by the speech recognition process, and based on the score obtained, the audio content is The degree of impact on the emotions of the users is quantitatively estimated, and the effectiveness of the audio content is evaluated. A content evaluation method for assessing sexuality.

6. A computer-implemented content evaluation method, comprising: Obtaining audio content to be output to a vehicle occupant; outputting the audio content; The voice content included in the passenger's speech after the voice content is output to recognize keywords in the audio content and exclamations indicating appreciation for the audio content; Performs voice recognition processing, As an index for evaluating the effectiveness of the audio content, the number of passengers and and the number of times the keyword and / or the exclamation is recognized by the speech recognition process. Get a score The passengers are a plurality of passengers and the keyword and / or the exclamation are recognized a first number of times. If the passenger is recognized, a first score corresponding to the first number of times is acquired, and and if the keyword and / or the exclamation is recognized the first number of times, and obtaining a second score that is greater than the first score.

7. A program executed by a content evaluation device having a computer, a content acquisition unit that acquires audio content to be output to a vehicle occupant; an output unit that outputs the audio content; The voice content included in the passenger's speech after the voice content is output to recognize keywords in the audio content and exclamations indicating appreciation for the audio content; A speech recognition unit that performs speech recognition processing; Depending on the number of times the keyword and / or the exclamation is recognized by the speech recognition process, and based on the score obtained, the audio content is The degree of impact on the emotions of the users is quantitatively estimated, and the effectiveness of the audio content is evaluated. A program that causes the computer to function as an evaluation unit that evaluates the quality of the product.

8. A program executed by a content evaluation device having a computer, a content acquisition unit that acquires audio content to be output to a vehicle occupant; an output unit that outputs the audio content; The voice content included in the passenger's speech after the voice content is output to recognize keywords in the audio content and exclamations indicating appreciation for the audio content; A speech recognition unit that performs speech recognition processing; As an index for evaluating the effectiveness of the audio content, the number of passengers and and the number of times the keyword and / or the exclamation is recognized by the speech recognition process. causing the computer to function as an evaluation unit that acquires a score; The evaluation unit is configured to evaluate whether the passengers are a plurality of passengers and whether the keywords and / or the exclamations are If the character is recognized a first number of times, a first score corresponding to the first number of times is obtained, and the previous The passenger is one and the keyword and / or the exclamation is recognized the first number of times. If so, the program obtains a second score that is greater than the first score.

9. A storage medium storing the program described in claim 7 or 8.

Citation Information

Patent Citations

  • Communication system for vehicle

    JP2003237453A

  • Image display system, digital photo-frame, information processing system, program, and information storage medium

    JP2010224715A

  • Electronic apparatus

    JP2010237761A

  • Sightseeing guide system

    JP2017194898A

  • Information processing device, information processing method, and program

    WO2018142686A1