Analysis system, information processing device, analysis method, and program
By recording and analyzing user audio feedback during virtual experiences, the system accurately determines user preferences for virtual environment objects, addressing inaccuracies in existing methods.
Patent Information
- Application Number
- JP2022017993
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-08
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-02-08
AI Technical Summary
Existing analysis systems inaccurately determine user preferences based on abstract attributes in a virtual space, leading to discrepancies between perceived and actual user interests.
An analysis system that records and analyzes audio information uttered by users while they view videos, using this information to accurately assess their preferences for objects in the virtual environment.
The system provides analysis results that align with actual user preferences by correlating audio feedback with visual content, enhancing the accuracy of preference determination.
Smart Images

Figure 0007809998000001 
Figure 0007809998000002 
Figure 0007809998000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an analysis system, an information processing device, an analysis method, and a program. [Background technology]
[0002] In recent years, AR (Augmented Reality) services and VR (Virtual Reality) services that provide various experiences in virtual spaces have been provided. AR services and VR services provide users with images that represent virtual worlds, allowing users to experience various virtual spaces.
[0003] Patent Document 1 discloses the configuration of an information analysis system that ascertains a user's preferences or interests based on the user's viewpoint in a virtual space. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2021-43819 Summary of the Invention [Problem to be solved by the invention]
[0005] In the information analysis system disclosed in Patent Document 1, a score is assigned to an abstract attribute associated with a marker that appears in a user's field of view, and the higher the score, the higher the user's interest in or preference for that abstract attribute. However, an abstract attribute that appears in a user's field of view by chance does not necessarily indicate a high level of interest or preference for the user. Therefore, there is a problem in that analyzing preferences based on a user's perspective may reveal trends that differ from the user's actual preferences.
[0006] One object of the present disclosure is to provide an analysis system, an information processing device, an analysis method, and a program that can output analysis results that are in line with the actual preference trends of a user. [Means for solving the problem]
[0007] The analysis system according to the first aspect of the present disclosure includes a display means for displaying video, a recording means for recording audio information uttered by a user watching the video, and an analysis means for using the audio information to analyze the user's preferences for objects included in the video that was displayed when the audio information was recorded.
[0008] An information processing device according to a second aspect of the present disclosure includes an acquisition unit that acquires audio information uttered by a user watching video displayed on a video device, and an analysis unit that uses the audio information to analyze the user's preferences for objects included in the video that was displayed when the audio information was recorded.
[0009] An analysis method according to a third aspect of the present disclosure acquires audio information uttered by a user watching video displayed on a video device, and uses the audio information to analyze the user's preferences for objects included in the video that was displayed when the audio information was recorded.
[0010] A program according to a fourth aspect of the present disclosure is a program that causes a computer to acquire audio information uttered by a user watching video displayed on a video device, and use the audio information to analyze the user's preferences for objects included in the video that was displayed when the audio information was recorded. [Effects of the Invention]
[0011] The present disclosure makes it possible to provide an analysis system, an information processing device, an analysis method, and a program that can output analysis results that are in line with the actual preference trends of a user. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a configuration diagram of an analysis system according to a first embodiment. [Figure 2] 1 is a configuration diagram of an information processing device according to a first embodiment. [Figure 3] FIG. 2 is a diagram showing a process flow of the analysis method according to the first embodiment. [Figure 4] FIG. 10 is a configuration diagram of an analysis system according to a second embodiment. [Figure 5] FIG. 10 is a configuration diagram of an HMD-equipped device according to a second embodiment. [Figure 6] FIG. 10 is a configuration diagram of an analysis server according to a second embodiment. [Figure 7] FIG. 10 is a diagram showing a flow of data collection processing in an HMD-equipped device according to the second embodiment. [Figure 8] FIG. 10 is a diagram showing the flow of analysis processing in the analysis server according to the second embodiment. [Figure 9] FIG. 10 is a diagram showing a detailed processing flow of an analysis process according to the second embodiment. [Figure 10] FIG. 10 is a diagram showing a detailed processing flow of an analysis process according to the second embodiment. [Figure 11] FIG. 10 is a diagram illustrating calculation of a position score according to the second embodiment. [Figure 12] FIG. 10 is a diagram showing a detailed process flow of negative / positive evaluation according to the second embodiment. [Figure 13] FIG. 13 is a diagram illustrating calculation of a position score according to the fourth embodiment. [Figure 14] FIG. 13 is a diagram illustrating calculation of a position score according to the fifth embodiment. [Figure 15] 3A and 3B are configuration diagrams of an HMD-equipped device and an analysis server according to each embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] (Embodiment 1) Hereinafter, embodiments of the present invention will be described with reference to the drawings. An example of the configuration of an analysis system according to the first embodiment will be described with reference to FIG. 1. The analysis system of FIG. 1 has a display means 11, a recording means 12, and an analysis means 13. The display means 11, the recording means 12, and the analysis means 13 may be software or modules whose processes are performed by a processor executing a program stored in a memory. Alternatively, the display means 11, the recording means 12, and the analysis means 13 may be hardware such as a circuit or a chip.
[0014] The display means 11, the recording means 12, and the analysis means 13 are provided in a computer device. The display means 11, the recording means 12, and the analysis means 13 may be provided in different computer devices, or may be provided in the same computer device. Alternatively, two or more elements of the display means 11, the recording means 12, and the analysis means 13 may be provided in a single computer device. A computer device is a device that operates when a processor executes a program stored in a memory.
[0015] The display means 11 displays images. The display means 11 may be, for example, a display device. Specifically, the display means 11 may be an HMD (Head Mounted Display) that provides an AR service or a VR service. The images displayed by the display means 11 may include actual scenery, game images, CG (Computer Graphics), etc.
[0016] The recording means 12 records audio information when a user utters a sound while watching a video. The recording means 12 may record the user's audio information via a microphone, for example. The recording means 12 may be a memory built into the computer device, or may be a memory externally attached to the computer device.
[0017] The analysis means 13 uses the audio information to analyze the user's preference for objects included in the video that was displayed when the audio information was recorded. The object included in the video may be, for example, a building, natural object, human, animal, vehicle, or other object with a name included in the video. Analyzing the user's preference for objects may be analyzing the user's preference trends for objects. Analyzing the user's preference for objects may be analyzing the user's feelings toward the object, for example, determining whether the user has a positive or negative impression of the object. Alternatively, analyzing the user's preference for objects may be analyzing the user's level of interest in the object, for example, determining whether the user has a high or low interest in the object. The analysis means 13 may analyze the user's preference for objects based on, for example, the content of the user's utterance.
[0018] 2, an information processing device 20, which is a computer device, may have an acquisition means 21 and an analysis means 13. The analysis means 13 included in the information processing device 20 is similar to the analysis means 13 in FIG. 1. The acquisition means 21 may acquire user voice information recorded in another device from the other device via a network or the like. The analysis means 13 in the information processing device 20 may analyze user preferences using the voice information acquired by the acquisition means 21.
[0019] Next, the flow of the analysis process according to the first embodiment will be described with reference to Fig. 3. Here, the flow of the analysis process executed by the information processing device 20 shown in Fig. 2 will be described. First, the acquisition means 21 acquires voice information when a user utters a voice while watching a video displayed on a video device (S11). The acquisition means 21 may acquire the voice information from another device via a network, or may acquire the voice information via a microphone provided in the information processing device 20.
[0020] Next, the analysis means 13 uses the audio information to analyze the user's preferences for objects included in the video that was displayed when the audio information was recorded (S12).
[0021] As described above, the analysis system of FIG. 1 or the information processing device 20 of FIG. 2 analyzes a user's preferences for objects included in a video based on the user's voice information. The user's voice information includes utterances that express the user's emotions. Thus, by using the content of the user's utterances regarding objects included in a video, the user's preferences for the objects can be accurately analyzed.
[0022] (Embodiment 2) Next, a configuration example of the analysis system according to the second embodiment will be described with reference to FIG. 4. The analysis system of FIG. 4 includes an HMD-equipped device 30 and an analysis server 40. The HMD-equipped device 30 and the analysis server 40 may be computer devices that operate when a processor executes a program stored in a memory. When using an AR service or a VR service, a user wears the HMD-equipped device 30 and watches an image output by the HMD-equipped device 30. The analysis server 40 analyzes the user's preferences using information received from the HMD-equipped device 30 via a network. The HMD-equipped device 30 may be connected to the analysis server 40 via a wireless communication line or a fixed communication line. The analysis server 40 may be a cloud server.
[0023] Next, a configuration example of the HMD-equipped device 30 according to the second embodiment will be described with reference to Fig. 5. The HMD-equipped device 30 has a display unit 31, an audio information recording unit 32, a visual field information recording unit 33, a control unit 34, and a communication unit 35. Each component of the HMD-equipped device 30 may be software or a module that performs processing by a processor executing a program stored in a memory. Alternatively, each component of the HMD-equipped device 30 may be hardware such as a circuit or a chip.
[0024] The display unit 31 may be a display device that displays or outputs video. Video used in AR or VR services is displayed on the display unit 31. The user experiences a virtual space by viewing the video displayed on the display unit 31.
[0025] The audio information recording unit 32 records as audio information what the user says while viewing the video displayed on the display unit 31 and experiencing the virtual space. For example, the audio information recording unit 32 records the audio information via an input interface such as a microphone. Recording audio information may be rephrased as recording audio. The audio information recording unit 32 may execute the recording process for the audio information while the video is displayed on the display unit 31. In other words, the audio information recording unit 32 may continue recording while the video is displayed on the display unit 31. In other words, the audio information recording unit 32 may record as audio information what the user says while the video is displayed on the display unit 31, including silent periods. The audio information recording unit 32 outputs the audio information to the control unit 34.
[0026] The visual field information recording unit 33 records visual field information in the virtual space experienced by the user. The visual field information may be an image viewed by the user. The visual field information includes at least one or more objects shown in the image. For example, the image displayed on the display unit 31 may change each time the user's viewpoint moves. Specifically, depending on the direction the user faces, the image of the virtual space displayed on the display unit 31 is changed to an image of the virtual space in the direction the user is facing. The user's viewpoint may be assumed to be at the center of the image displayed on the display unit 31. As the image displayed on the display unit 31 changes, the objects included in the image also change. The visual field information recording unit 33 outputs the visual field information to the control unit 34.
[0027] The control unit 34 extracts user speech information from the audio information acquired from the audio information recording unit 32 and generates one data or one audio file. One audio file contains speech information from when the user starts speaking to when the speech ends. Furthermore, if a silent period is shorter than a predetermined period, for example, if the silent period is less than one second, the speech may not be considered to have ended but may be considered to be continuing. In other words, one audio file may contain a silent period shorter than a predetermined period. Furthermore, the control unit 34 may associate the audio file with view field information corresponding to the timing of the speech included in the audio file. In other words, the control unit 34 may associate the speech information with view field information including the video the user was watching while speaking.
[0028] The communication unit 35 transmits the audio file and the visual field information generated by the control unit 34 to the analysis server 40.
[0029] Next, a configuration example of the analysis server 40 according to the second embodiment will be described with reference to Fig. 6. The analysis server 40 has an analysis unit 41, an output unit 42, and a communication unit 43. The analysis unit 41, the output unit 42, and the communication unit 43 may be software or modules that are executed by a processor executing a program stored in a memory. Alternatively, the analysis unit 41, the output unit 42, and the communication unit 43 may be hardware such as a circuit or a chip.
[0030] The communication unit 43 receives the audio file and visual field information associated with the audio file transmitted from the HMD-equipped device 30. The analysis unit 41 receives the audio file and visual field information received by the communication unit 43 and analyzes the user's preferences. The output unit 42 outputs the analysis results of the analysis unit 41 to a display or the like of the analysis server 40. The analysis process of the analysis unit 41 will be described later.
[0031] Next, the flow of data collection processing in the HMD-equipped device 30 according to the second embodiment will be described with reference to FIG. 7. First, the control unit 34 detects the end of the user's virtual space experience (S21). For example, the control unit 34 may determine that the user's virtual space experience has ended when the playback of the video used in the virtual space experience has ended, or when the control unit 34 receives a signal from the user instructing the video of the virtual space experience to stop. While the user is experiencing the virtual space, the audio information recording unit 32 continues to record audio information, and the visual field information recording unit 33 continues to record visual field information.
[0032] Next, the control unit 34 acquires audio information from the audio information recording unit 32 (S22). Furthermore, the control unit 34 acquires visual field information from the visual field information recording unit 33 (S23). The control unit 34 manages the audio information acquired in step S22 and the visual field information acquired in step S23 in association with each other.
[0033] Next, the control unit 34 obtains the start and end timing of all utterances from the audio information. The user utters various words while experiencing the virtual space. Furthermore, the user does not always utter words; after uttering a word, the user may utter another word after a period of time, such as a few seconds. In other words, the audio information includes utterance information and silence information. The utterance information is information about the period from when the user utters a word until a silence state occurs. Here, if the silence state is shorter than a predetermined period, it may be considered that the user's utterance is continuous. The utterance information corresponds to an audio file.
[0034] Next, the control unit 34 acquires the nth piece of utterance information from all of the utterance information acquired in step S24 (S25). n is an integer equal to or greater than 1, and the control unit 34 first acquires the first piece of utterance information. For example, the utterance information may be arranged in order of the timing at which the utterance was made, and the first piece of utterance information may be the utterance information made at the oldest timing among the audio information. Alternatively, the first piece of utterance information may be the utterance information made at the most recent timing among the audio information.
[0035] Next, the control unit 34 acquires visual field information associated with the n-th piece of utterance information (S26). Specifically, the control unit 34 acquires visual field information at the same timing as the n-th piece of utterance information. The visual field information at the same timing as the n-th piece of utterance information is visual field information from the timing when the n-th piece of utterance information starts to the timing when it ends.
[0036] Next, the control unit 34 stores the n-th piece of utterance information acquired in step S25 and the visual field information acquired in step S26 in a data set n (S27). The data set n is data including the n-th piece of utterance information and the visual field information associated with the n-th piece of utterance information.
[0037] Next, the control unit 34 determines whether all of the speech information included in the audio information has been saved in a data set n (S28). The data set n may be saved in, for example, a memory built into or external to the HMD-equipped device 30. If the control unit 34 determines that all of the speech information included in the audio information has been saved in the data set n, it transmits all of the data sets to the analysis server 40 (S29).
[0038] If the control unit 34 determines that not all of the speech information contained in the audio information has been saved in the data set n, that is, that there is speech information that has not been acquired, it sets n=n+1 and executes the processing from step S25 onwards.
[0039] Next, the flow of the analysis process in the analysis server 40 according to the second embodiment will be described with reference to Fig. 8. First, the communication unit 43 receives a data set from the HMD-equipped device 30 (S31). Next, the analysis unit 41 executes a process of analyzing the user's preferences using the data set (S32). Next, the output unit 42 outputs the analysis result of the analysis unit 41 to a display device or the like (S33).
[0040] Next, the analysis process in step S32 in Fig. 8 will be described in detail with reference to Fig. 9. The analysis unit 41 receives a dataset n (S41). n is an integer equal to or greater than 1, and the analysis unit 41 first receives dataset 1 where n=1. Alternatively, the analysis unit 41 may select and extract dataset n from all datasets stored in the memory or the like of the analysis server 40. For example, the analysis unit 41 may extract datasets in order starting from dataset 1.
[0041] Next, the analysis unit 41 plays back the video of the visual field information included in the dataset n (S42). Next, the analysis unit 41 detects objects appearing in the video (S43). For example, the analysis unit 41 detects objects included in the video using image recognition technology. For example, the analysis unit 41 may generate a learning model that identifies the names of objects by performing machine learning in advance. The analysis unit 41 may identify the names of objects by inputting detected objects into the learning model. The name of the object may be a name that indicates the attributes of the object, such as a building, a person, or a dog. The analysis unit 41 may also detect objects using a learning model generated by performing machine learning. Alternatively, the analysis unit 41 may detect only predetermined objects. For example, if a specific building, person, or object is specified, the analysis unit 41 may detect objects included in the video that have feature values that differ from the feature values of the predetermined building, etc. within a predetermined range. Alternatively, the analysis unit 41 may detect a specific building, etc. from the video using a learning model generated by machine learning a specific building, etc.
[0042] Next, the analysis unit 41 calculates a position score for the detected object (S44). Here, the calculation of the position score will be described with reference to FIG. 11. The area enclosed by the solid-line rectangle in FIG. 11 represents the image of the reconstructed field of view information. The dotted line represents the center line that bisects the area enclosed by the solid line. The ellipses indicated by A11 and A12 represent objects. FIG. 11 shows an image of the field of view information including objects A11 and A12. The position score assigned to an object increases as the object approaches the center line. For example, a position score of 100 is assigned to an object on the center line, and a position score of 0 is assigned to a position farthest from the center line. Specifically, the distance from the center line to the vertical edge may be divided into 100 equal parts, and a position score ranging from 0 to 100 may be assigned to each position. FIG. 11 shows that object A12 is closer to the center line than object A11. In such a case, for example, the position score of object A12 may be set to 50, and the position score of object A11 may be set to 20.
[0043] Returning to FIG. 9, the analysis unit 41 waits for a predetermined period of time (S45). For example, the analysis unit 41 may wait for 1 second or 0.1 seconds. Next, the analysis unit 41 determines whether the playback of the video of the visual field information has finished (S46). If the analysis unit 41 determines that the playback of the video of the visual field information has not finished, it repeats the processes from step S43 onward and calculates the position score of the object. Here, in step S44, for objects whose position scores have already been calculated, the newly calculated position score is added to the already calculated position score. In other words, if the position score of an object is calculated as 50 and then 1 second later the position score of that object is calculated as 20, the position score of that object will be 70.
[0044] The shorter the predetermined period in step S45, the more frequently the object is detected in step S43, and therefore the more accurate the object position score becomes.
[0045] In step S46, if the analysis unit 41 determines that the playback of the video of the visual field information has ended, it performs a negative / positive evaluation of the content of the user's utterance (S47). Negative / positive refers to negative and positive. For example, the analysis unit 41 determines whether the content of the utterance information included in the data set n received in step S41 is negative or positive. If the analysis unit 41 determines that the content of the utterance information is negative (S48), it multiplies the position score of each object by "-1" (S49). If the analysis unit 41 determines that the content of the utterance information is positive (S48), it executes the processing from step S50 onwards using the value of the position score of each object.
[0046] Next, the analysis unit 41 sets the position score calculated in step S48 or S49 as the attention score (S50). If the content of the utterance information is determined to be positive in step S48, the attention score is a positive number, and if the content of the utterance information is determined to be negative in step S48, the attention score is a negative number.
[0047] Next, the analysis unit 41 determines whether or not attention scores have been calculated for all data sets (S51). If the analysis unit 41 determines that attention scores have not been calculated for all data sets, it sets n=n+1 (S53) and repeats the processes from step S41 onwards. If the analysis unit 41 determines that attention scores have been calculated for all data sets, it sums the attention scores for each object to calculate an object score for each object (S52). The object score for each object is the sum of the attention scores for each object calculated for all data sets.
[0048] Next, the negative / positive evaluation performed in step S47 of Fig. 9 will be described in detail with reference to Fig. 12. First, the analysis unit 41 performs morphological analysis on the utterance information (S61). Specifically, the analysis unit 41 divides the words included in the utterance information into morphemes. A morpheme is intended to be the smallest unit in which a term has meaning.
[0049] Next, the analysis unit 41 obtains an average score for each divided term using a polarity dictionary (S62). The polarity dictionary may be, for example, the Japanese Evaluation Polarity Dictionary from the Inui-Suzuki Laboratory at Tohoku University or the Word Sentiment Polarity Correspondence Table from the Okumura-Takamura Laboratory at the Institute of Future Interdisciplinary Research in Science and Technology, Tokyo Institute of Technology. The analysis unit 41 calculates a score for each divided term using the polarity dictionary and calculates an average score for all terms. Alternatively, the analysis unit 41 may calculate a score for each divided term using the polarity dictionary and calculate a total value by adding up the scores for all terms.
[0050] The analysis unit 41 determines whether the average score of the utterance information is equal to or greater than 0 (S63). If the average score of the utterance information is equal to or greater than 0, the analysis unit 41 determines that the content of the utterance information is positive (S64), and if the average score is less than 0, the analysis unit 41 determines that the content of the utterance information is negative (S65).
[0051] Here, the output process of the analysis results in step S33 of Fig. 8 will be described. For example, the analysis unit 41 may create data in which the calculated object scores are arranged in descending order of score, and create a ranking of objects that the user is most interested in. The output unit 42 may output the ranking of objects created by the analysis unit 41 to a display device or the like. The display device may be a device used integrally with the analysis server 40, or may be a terminal device such as a smartphone held by a user who utilizes the analysis results. The analysis unit 41 may transmit information indicating the ranking to the terminal device via the communication unit 43.
[0052] Alternatively, the analysis unit 41 may display the object score in the virtual space. For example, the analysis unit 41 may transmit the object score to the HMD-equipped device 30. At this time, the HMD-equipped device 30 may display the received object score superimposed on the object in the virtual space.
[0053] Furthermore, the analysis unit 41 may display the scores of the objects using different colors. By displaying the object scores in various formats in this way, the analysis unit 41 allows the user of the virtual space to more visually grasp the user's interest or concern for each object.
[0054] As described above, the analysis server 40 according to the second embodiment can identify the negative or positive emotion a user has toward each object included in the visual field information. This allows the analysis server 40 to combine the user's emotion with the position score to output a quantitative result of the user's interest or concern. Furthermore, the analysis server 40 can estimate whether the user's gaze is based on a negative emotion or a positive emotion by correcting the position score calculated using the visual field information with the negative / positive evaluation result of the speech information.
[0055] (Embodiment 3) Next, the analysis process of the analysis unit 41 according to the third embodiment will be described. The analysis unit 41 may calculate the result of the negative / positive evaluation of the utterance information as a score having a specific range of values, instead of a binary value of negative or positive. The score indicating the result of the negative / positive evaluation indicates, for example, the user's level of interest in the object.
[0056] Furthermore, the analysis unit 41 may acquire the pitch or volume of the user's voice when he / she speaks as the voice information, and may correct the score calculated as a result of the negative / positive evaluation using the pitch or volume of the voice. For example, the analysis unit 41 calculates the average value of the pitch (unit: hertz) and volume (unit: decibel) of the voice for all the speech information for which the negative / positive scores have been calculated.
[0057] If the average pitch value of each piece of utterance information is higher than the average pitch value of all pieces of utterance information and the score of the negative / positive evaluation is a positive number, the analysis unit 41 may multiply the score by 1.1. Alternatively, if the average pitch value of each piece of utterance information is higher than the average pitch value of all pieces of utterance information and the score of the negative / positive evaluation is a negative number, the analysis unit 41 may multiply the score by 0.9.
[0058] If the average pitch value of each piece of utterance information is lower than the average pitch value of all pieces of utterance information and the score of the negative / positive evaluation is a positive number, the analysis unit 41 may multiply the score by 0.9. Alternatively, if the average pitch value of each piece of utterance information is lower than the average pitch value of all pieces of utterance information and the score of the negative / positive evaluation is a negative number, the analysis unit 41 may multiply the score by 1.1.
[0059] The analysis unit 41 may also correct the score value based on the loudness of the voice, in the same way as for the pitch of the voice.
[0060] Alternatively, the analysis unit 41 may perform a negative / positive evaluation using at least one of the volume and pitch of the voice without considering the content of the utterance. In other words, the analysis unit 41 may evaluate the content of the utterance information as positive if the voice is loud or high, and may evaluate the content of the utterance information as negative if the voice is quiet or low.
[0061] As described above, by correcting the score calculated as a result of the negative / positive evaluation using the pitch or volume of the voice, the user's emotions can be more accurately reflected in the score.
[0062] (Fourth embodiment) Next, an example of calculating a position score will be described with reference to Fig. 13. The area enclosed by the solid-line rectangle in Fig. 13 indicates the image of the reproduced field of view information. The dotted circle is a circle whose center is the intersection of the diagonals of the area enclosed by the solid-line rectangle and is a circle that circumscribes the solid-line rectangle. The ellipse indicated by A21 indicates an object. Fig. 13 shows the image of the field of view information including object A21.
[0063] The position score assigned to an object is 100 for the center of the circle and 0 for the point on the circle. Furthermore, the distance from the center of the circle to the point on the circle may be divided into 100 equal parts, and a position score from 0 to 100 may be assigned to each position. However, even if an object is within the circle, no position score is assigned to an object that is outside the image of the field of view information, that is, outside the area enclosed by the solid-line rectangle. In the example shown in FIG. 13, object A21 is located within the area enclosed by the solid-line rectangle and is also located in the middle of the distance from the center of the circle to the point on the circle, so its position score is 50.
[0064] Alternatively, the distance from the center of the circle to the solid square may be divided into 100 equal parts, and each position may be assigned a position score from 0 to 100. In this case, the position score of the center of the circle may be set to 100, and the position score of each side of the solid square may be set to 0.
[0065] As described above, by calculating the position score using two-dimensional information in the X-axis and Y-axis directions of the image of the visual field information, it is possible to calculate a more accurate position score taking the user's viewpoint into consideration. In Fig. 13, the X-axis direction is the direction parallel to the long side, and the Y-axis direction is the direction parallel to the short side.
[0066] (Embodiment 5) Next, a position score correction process according to the fifth embodiment will be described. The analysis unit 41 may correct the position score so that an object that is closer to the center receives a higher score than the previous position score measurement. For example, if the corrected position score is a score obtained by applying a correction to the position score, the corrected position score may be calculated using the following formula: Corrected Position Score = Position Score + {0.001 × (difference in approach from the position score at the previous measurement + number of consecutive approaches)}. FIG. 14 shows an image of the visual field information, similar to FIG. 11, indicating the presence of an object 31. For example, as shown in FIG. 14, a case will be described in which the position score of the object 31 fluctuates from (first measurement) 20 → (second measurement) 50 → (third measurement) 90. In this case, the corrected position score for the first measurement is 20 + 0.001 × {0 + 0} = 20. Similarly, the corrected position scores for the second and third times are 50+0.001×{(50−20)+1}=50.031 and 90+0.001×{(90−50)+2}=90.042, respectively.
[0067] As explained above, it is possible to assign a higher position score to an object that is initially located away from the center and then moves closer to the center. Note that there is no specific value for the coefficient (0.001) used to calculate the corrected position score. The lower this value is, the less impact the action of the object moving closer to the center will have on the corrected position score, and the higher this value is, the greater the impact on the corrected position score.
[0068] FIG. 15 is a block diagram showing an example configuration of the HMD-equipped device 30 and analysis server 40 (hereinafter referred to as the HMD-equipped device 30, etc.) described in the above-mentioned embodiment. Referring to FIG. 15, the HMD-equipped device 30, etc. includes a network interface 1201, a processor 1202, and a memory 1203. The network interface 1201 may be used to communicate with a network node. The network interface 1201 may include, for example, a network interface card (NIC) that complies with the IEEE 802.3 series. IEEE stands for Institute of Electrical and Electronics Engineers.
[0069] The processor 1202 reads and executes software (computer programs) from the memory 1203 to perform the processing of the HMD-equipped device 30 and the like described using flowcharts in the above-described embodiments. The processor 1202 may be, for example, a microprocessor, an MPU, or a CPU. The processor 1202 may include multiple processors.
[0070] The memory 1203 is configured by a combination of volatile memory and non-volatile memory. The memory 1203 may include storage located remotely from the processor 1202. In this case, the processor 1202 may access the memory 1203 via an I / O (Input / Output) interface (not shown).
[0071] 15, the memory 1203 is used to store software modules. The processor 1202 reads and executes these software modules from the memory 1203, thereby performing the processing of the HMD-equipped device 30 and the like described in the above-described embodiment.
[0072] As explained using Figure 15, each of the processors possessed by the HMD-equipped device 30, etc. in the above-mentioned embodiments executes one or more programs including a group of instructions for causing a computer to perform the algorithm explained using the drawings.
[0073] In the above examples, the program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored on a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
[0074] The present invention is not limited to the above-described embodiment, and can be modified as appropriate within the scope of the invention.
[0075] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix 1) a display means for displaying an image; a recording means for recording voice information when a user utters a voice while watching the video; and an analysis means for using the audio information to analyze the user's preferences for objects included in the video that was being displayed when the audio information was recorded. (Appendix 2) The analysis means An analysis system as described in Appendix 1, which performs a negative / positive evaluation of the audio information and uses the negative / positive evaluation to analyze the user's preferences for the object. (Appendix 3) The analysis means The analysis system of Appendix 2 calculates a score indicating the degree of interest in the object using the negative / positive evaluation, and analyzes the user's preferences for the object based on the score. (Appendix 4) The recording means Further recording gaze information indicating the user's gaze, The analysis means 4. The analysis system according to claim 1, wherein the voice information and the gaze information are used to analyze the user's preferences for the object. (Appendix 5) The analysis means The analysis system described in Appendix 4, wherein the closer an object is to the user's line of sight, the higher its position score is, the object score of the object is calculated by correcting the position score using a negative / positive evaluation of the audio information, and the user's preferences for the object are analyzed based on the object score. (Appendix 6) The analysis means The analysis system of claim 5, wherein the position score is corrected using a negative / positive evaluation of the audio information and at least one of the volume and pitch of the voice contained in the audio information. (Appendix 7) The analysis means 7. The analysis system of claim 5, wherein the position score is corrected based on a change in the difference in position between the object and the user's line of sight. (Appendix 8) an acquisition unit that acquires voice information when a user utters a voice while watching a video displayed on a video device; an analysis unit that uses the audio information to analyze the user's preferences for objects included in a video that was being displayed when the audio information was recorded. (Appendix 9) The analysis unit 9. The information processing device according to claim 8, wherein a negative / positive evaluation of the audio information is performed, and the user's preference for the object is analyzed using the negative / positive evaluation. (Appendix 10) Acquires voice information when a user utters a voice while watching a video displayed on a video device; An analysis method that uses the audio information to analyze the user's preferences for objects included in a video that was being displayed when the audio information was recorded. (Appendix 11) Acquires voice information when a user utters a voice while watching a video displayed on a video device; A program that causes a computer to use the audio information to analyze the user's preferences for objects included in the video that was being displayed when the audio information was recorded. [Explanation of symbols]
[0076] 11 Display means 12 Recording means 13 Analysis tools 20 Information processing equipment 21 Acquisition method 30 HMD-equipped device 31 Display section 32 Audio information recording unit 33 Visual field information recording unit 34 Control Unit 35 Communications Department 40 Analysis Server 41 Analysis Department 42 Output section 43 Communications Department
Claims
1. a display means for displaying an image; a recording means for recording voice information when a user utters a voice while watching the video; an analysis means for analyzing, using the audio information, the user's preferences for objects included in the video that was being displayed when the audio information was recorded; The recording means Further recording gaze information indicating the user's gaze, The analysis means An analysis system that uses the audio information and the gaze information to assign a higher position score to objects closer to the user's gaze, calculates an object score for the object by correcting the position score using a negative / positive evaluation of the audio information, and analyzes the user's preferences for the object based on the object score.
2. The analysis means The analysis system according to claim 1 , wherein a negative / positive evaluation of the audio information is performed, and the user's preference for the object is analyzed using the negative / positive evaluation.
3. The analysis means The analysis system according to claim 2 , further comprising: calculating a score indicating a degree of interest in the object using the negative / positive evaluation; and analyzing the user's preferences for the object based on the score.
4. The analysis means The analysis system according to claim 1 , wherein the position score is corrected using a negative / positive evaluation of the audio information and at least one of the volume and pitch of the voice included in the audio information.
5. The analysis means The analysis system according to claim 1 , wherein the position score is corrected based on a variation in a difference between the position of the object and the line of sight of the user.
6. an acquisition unit that acquires voice information when a user utters a voice while watching a video displayed on a video device and gaze information that indicates a gaze of the user; an analysis unit that uses the audio information and the gaze information to assign a higher position score to objects closer to the user's gaze, calculates an object score for the object by correcting the position score using a negative / positive evaluation of the audio information, and analyzes the user's preferences for objects included in the video that was displayed when the audio information was recorded based on the object score.
7. Acquires voice information when a user utters a voice while viewing a video displayed on a video device and gaze information indicating a gaze of the user; An analysis method that uses the audio information and the gaze information to assign a higher position score to objects closer to the user's gaze, calculates an object score for the object by correcting the position score using a negative / positive evaluation of the audio information, and analyzes the user's preferences for objects included in the video that was displayed when the audio information was recorded based on the object score.
8. Acquires voice information when a user utters a voice while viewing a video displayed on a video device and gaze information indicating a gaze of the user; A program that causes a computer to execute the following steps: use the audio information and the gaze information to assign a higher position score to objects closer to the user's gaze, calculate an object score for the object by correcting the position score using a negative / positive evaluation of the audio information, and analyze the user's preferences for objects included in the video that was displayed when the audio information was recorded based on the object score.
Citation Information
Patent Citations
Method executed by computer for providing information via head mount device, program for causing computer to execute the same, and information processing device
JP2019101457A
Information presentation system, information analysis system, information presentation method, information analysis method, and program
JP2021043819A
Method and system for rendering multimedia content based on interest level of user in real-time
US20190050486A1
Speech providing system, server, client machine, information providing management server, and voice providing method
WO2003085511A1