Control method, and arithmetic unit
The control method dynamically manages the level of detail in virtual spaces by using user context and action scripts to calculate the likelihood of script options, effectively addressing the challenges of data complexity and processing load in CAD and BIM applications.
Patent Information
- Application Number
- JP2023210863
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for displaying design data in virtual spaces, such as those used in CAD and BIM applications, face challenges in efficiently managing the level of detail (LOD) to balance data complexity and processing load.
A control method that uses an arithmetic device to acquire user context and action scripts, calculates the likelihood of script options using a language model, and adjusts the display in the virtual space based on these calculations to dynamically manage the level of detail.
This approach allows for dynamic control of the display in virtual spaces based on user actions, optimizing the level of detail to reduce processing load while ensuring sufficient detail for design discussions.
Smart Images

Figure 2025095073000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a control method and an arithmetic unit.
Background Art
[0002] There are known applications that display an object in a 3D space and use the space for virtual experience, education and training, and design discussions. In these applications, it is also possible to freely move within the virtual space and discuss while operating the object. In some cases, the data used in CAD (Computer-aided design) at the design time is converted and used for the data constructing this virtual space. Particularly, in fields such as BIM, such activities have come to be actively used. When displaying design data such as CAD in a virtual space, adjustment of the level of detail is required.
[0003] This degree of fineness is called LOD (Level of Detail), and is related not only to BIM but also to detailed display of 3D design and CG. In the world of 3D animation, in many cases, the model depicted near the user is displayed in detail, and what appears in the distance is depicted with a reduced level of detail to reduce the processing burden. Since all the data of the BIM model or CAD model in a certain scene is enormous, if an attempt is made to display all the data, problems such as an excessive amount of data to be read and an excessive processing load will occur.
[0004] This problem can be solved by using a simplified model instead of using detailed data. However, when it is desired to discuss design procedures, quality problems, etc. using the virtual space, detailed data is required for the object to be discussed. For example, in CAD data, information such as screws for fixing parts is unnecessary in a general display, but is required for confirmation and visualization of the assembly procedure.
[0005] Patent Document 1 discloses a model display control method in which a computer executes a process for controlling the display of a model of an observation target. The method includes: specifying, from among texture models of the observation target stored in a storage unit, the model of a target member constituting the observation target and the model of an adjacent member adjacent to the target member; specifying, for the specified model of the target member and the model of the adjacent member, the level of detail required for the accuracy of observation corresponding to each preset model; and selecting, as display targets, the models with the specified levels of detail for the model of the target member and the model of the adjacent member, respectively.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] In the invention described in Patent Document 1, there is room for improvement in the display method in the virtual space.
Means for Solving the Problems
[0008] The control method according to the first aspect of the present invention is a control method in which an arithmetic device controls the display in a virtual space, and includes: a context acquisition process for acquiring a user context, which is data indicating the intention of a user; a script acquisition process for acquiring an action script indicating the operation of a component in the virtual space; a likelihood calculation process for inputting the state of the virtual space, the user context, and options of a plurality of the action scripts into a language model to obtain the likelihoods of the options of the plurality of action scripts; and a display change process for changing the display in the virtual space based on the likelihoods. The arithmetic unit according to the second aspect of the present invention includes at least one processor, and the processor acquires a user context which is data indicating the intention of the user, acquires an action script indicating the operation of a component in a virtual space, inputs the state of the virtual space, the user context, and options of a plurality of the action scripts into a language model to acquire the likelihood of the options of the plurality of action scripts, and changes the display in the virtual space based on the likelihood.
Effect of the Invention
[0009] According to the present invention, the display in the virtual space can be controlled based on the actions of the user.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Mode for Carrying Out the Invention
[0011] - First Embodiment - Hereinafter, with reference to FIGS. 1 to 15, a first embodiment of the control method will be described.
[0012] FIG. 1 is an overall configuration diagram of a virtual space display system. The virtual space display system includes a server 1, a client terminal 2, and a database 3. The server 1, the client terminal 2, and the database 3 are interconnected via a network 4. A user 5 uses the client terminal 2. The client terminal 2 is provided with an input device 104 and a display device 105. The user 5 performs an operation input on the virtual space using the input device 104. The client terminal 2 receives the operation input of the user 5 and transmits it to the server 1. Further, the client terminal 2 generates a video of the virtual space using the data received from the server 1 and outputs it to the display device 105. The video of the virtual space includes, for example, a virtual object 6.
[0013] Data related to the virtual space is stored in the database 3. The server 1 refers to the database 3 as necessary and performs calculations related to the virtual space. The database 3 includes a component data model 801, component-specific work data 802, a list of components of the same material 803, employee work data 804, an action script collection 805, and a component operation collection 806. The network 4 is a communication network. The network 4 is, for example, a wide area network such as the Internet, a dedicated line, a local area network, or a wireless communication network provided by a communication carrier. Also, the network 4 may be a combination of different types of networks, for example, a combination of a dedicated line and a wireless communication network.
[0014] Figure 2 is a hardware configuration diagram of the client terminal 2. The client terminal 2 includes a bus interface (hereinafter, the interface is also referred to as "IF") 100, a main memory device 101, a CPU 102, a video generation unit 103, an input device 104, a display device 105, a camera 106, an external memory device 107, and a network IF 109. However, the client terminal 2 does not necessarily have to be integrally configured. For example, a device in which the display device 105 and the camera 106 are integrated may be connected to a device including other configurations by wire or wirelessly.
[0015] The bus IF 100 is an interface of a communication bus built into the client terminal 2. The main memory device 101 is, for example, a DRAM. The CPU 102 is a central processing unit and is also called a "processor". The CPU 102 may include a plurality of arithmetic cores, or may include a plurality of physical CPUs each having one or more arithmetic cores. That is, the CPU 102 can be called "one or more processors".
[0016] The video generation unit 103 generates an image of the virtual space viewed from the viewpoint based on the data of the objects and viewpoints arranged in the virtual space. The video generation unit 103 is, for example, a GPU (Graphics Processing Unit). However, the video generation unit 103 may be built into the CPU 102. The input device 104 is a device for the user 5 to input some data to the client terminal 2. The input device 104 is, for example, a mouse or a keyboard. The input device 104 may include a microphone for picking up the speech of the user 5. The display device 105 displays the image of the virtual space generated by the video generation unit 103. The display device 105 is, for example, a liquid crystal display.
[0017] The camera 106 is a camera that photographs at least one of the hand of user 5, the eyes of user 5, and the surroundings of user 5. When the camera 106 photographs the hand of user 5, the CPU 102 recognizes the gesture of user 5 from the image obtained by the camera 106. When the camera 106 photographs the eyes of user 5, the CPU 102 detects the line of sight of user 5 from the image obtained by the camera 106. When the camera 106 photographs the surroundings of user 5, the CPU 102 detects the movement of user 5, that is, changes in movement and posture, from the image obtained by the camera 106. The external storage device 107 is a non-volatile storage device. A program for realizing the functions described later is stored in the external storage device 107, and the CPU 102 expands and executes this program in the main storage device 101. The network IF 109 is a network interface card. Note that the hardware configuration of server 1 may be the same as that of client terminal 2, or may not include the input device 104, the display device 105, and the camera 106.
[0018] Figure 3 is a diagram showing an example of the client terminal 2. The client terminal 2 shown in Figure 2 is a head-mounted display, and the display device 105 and the camera 106 are shown. The display device 105 is worn on the head of user 5 and is arranged at a position covering the field of view of user 5. The camera 106 photographs the outside of the display device 105, specifically, the hand and fingers of user 5. The processor 203 generates an image and outputs it to the display device 105. Also, the processor 203 recognizes the movement of the fingers and hand of user 5 photographed by the camera 106 as a gesture, and transmits the identifier of the recognized gesture to the server 1.
[0019] FIG. 4 is a diagram showing an example of a component data model 801. In the component data model 801, data of each component included in the virtual space is stored. The component data model 801 is composed of a plurality of records, and each record has fields of component ID 80101, 3D data ID 80102, material ID 80103, position 80104, rotation information 80105, component name 80106, tag list 80107, component description 80108, child element 80109, isomorphic device 80110, related similarity 80111, material list 80112, parent component group 80113, fixed parent component 80114, parent component fixture 80115, fixed child component 80116, and child component fixture 80117.
[0020] However, hereinafter, for the sake of brevity, the data input to each field will be described using the field names. For example, strictly speaking, the identifier of the component is stored in the field of component ID 80101. However, hereinafter, for the sake of brevity, it will be described as "the component ID 80101 is the identifier of the component". The component ID 80101 is a unique value in the component data model 801 and has different values in any record. Although only two records are shown in FIG. 4, actually, there are a large number of records.
[0021] The 3D data ID 80102 is the identifier of the 3D model used for the component, for example, the file name of the 3D model. For example, when a vehicle is equipped with a plurality of wheels, the component ID 80101 of each wheel must have different values, but the 3D data ID 80102 may have the same value. The material ID 80103 is data indicating the appearance of the component and is the file name of a pattern or texture image to be pasted on the surface of the 3D model.
[0022] The position 80104 indicates the position where the component is arranged. The rotation information 80105 indicates the posture of the component, for example, the change from a preset standard posture for each component. The component name 80106 indicates the name of the component in natural language. The tag list 80107 is a list of tags related to the component.
[0023] The component description 80108 is a natural language description of the component. The child element 80109 is a list of other components that make up the component, in other words, a list of sub-components. The same-type device 80110 is a list of component IDs of components having the same model number. The related similar 80111 is a list of component IDs of components that are not of the same model number but are similar. The material list 80112 is a list of identifiers of the materials that make up the component. The parent component group 80113 is a list of identifiers of the groups of parent components. The fixed parent component 80114 is the identifier of the parent component that fixes the said component. The parent component fixture 80115 is the identifier of the component that fixes the parent component of the said component. The fixed sub-component 80116 is the identifier of the sub-component that fixes the said component. The sub-component fixture 80117 is the identifier of the component that fixes the sub-component of the said component.
[0024] Figure 5 is a diagram showing an example of component-specific work data 802, the same-material component list 803, and the employee work data 804. The component-specific work data 802 is composed of a plurality of records, and each record has fields of the target component 8021, the use lot ID 8022, the responsible worker 8023, and the target product number 8024. The target component 8021 is the identifier of the component corresponding to the component ID 80101 of the component data model 801. The use lot ID 8022 is an identifier indicating the manufacturing lot of the target component 8021. The responsible worker 8023 is the identifier of the person in charge who worked using the target component 8021. The target product number 8024 is the identifier of the product created using the target component 8021, and in this embodiment, it can also be said to be the vehicle ID.
[0025] The same-material component list 803 is composed of a plurality of records, and each record has fields of the component material ID 8031 and the corresponding component 8032. The component material ID 8031 is the identifier of the material that makes up the component, and it corresponds to the material list 80112 of the component data model 801. The corresponding component 8032 is a list of identifiers of components having the material indicated by the component material ID 8031, and it corresponds to the component ID 80101 of the component data model 801.
[0026] The worker operation data 804 is composed of a plurality of records, and each record has fields of a worker ID 8041, a working time 8042, parts related to the target operation 8043, a target product number 8044, and work record data 8045. The worker ID 8041 is an identifier of a worker and corresponds to the responsible worker 8023 of the part-specific operation data 802. The working time 8042 is the time period during which the work was engaged in. The parts related to the target operation 8043 are identifiers of the parts used in the work and correspond to the part ID 80101 of the part data model 801 and the target part 8021 of the part-specific operation data 802. The target product number 8044 is an identifier of the product including the parts to be worked on and corresponds to the target product number 8024 of the part-specific operation data 802. The work record data 8045 is an identifier of the data recording the work, for example, the file name of a video recording the state of the work.
[0027] Figure 6 is a diagram showing an example of the action script set 805. The action script set 805 is composed of a plurality of records, and each record has fields of an action ID 8051, an action name 8052, a required detail level 8053, a target group 8054, a parts set 8055, parts operation information 8056, and a work description 8057. A plurality of action scripts are stored in the action script set 805, and each record is one action script. The action ID 8051 is an identifier of the action script.
[0028] The action name 8052 is the name of the action. The required detail level 8053 is the detail level of the parts required to execute the action. The target group 8054 is an identifier of the group of parts targeted by the action script. The parts set 8055 is a list of identifiers of the parts targeted by the action script. The parts operation information 8056 is a list of identifiers of parts operations in which operations on the parts are described. A specific example of the parts operation will be described with reference to Figure 7. The work description 8057 is the content of the action script described in natural language. In other words, the work description 8057 is an explanation of the process of connecting each of the parts operations described in the parts operation information 8056.
[0029] Note that the action script shown in FIG. 6 is only an example, and other processes may be included. For example, an action script for displaying specific items of a component in a virtual space may be included. The items to be displayed may be the description 80108 of the component in the component data model 801, or may be two or more items.
[0030] FIG. 7 is a diagram showing an example of the component operation set 806. The component operation set 806 is composed of a plurality of records, and each record has fields of an operation ID 8061, a target component ID 8062, a relative movement distance 8063, relative rotation coordinates 8064, a material change 8065, an operation time 8066, and a work description 8067. The operation ID 8061 is an identifier of the component operation.
[0031] The target component ID 8062 is an identifier of the component to be operated in the component operation. The relative movement distance 8063 is the relative distance to move in the component operation. The relative rotation coordinates 8064 are the relative coordinates for the component to rotate in the component operation. The material change 8065 is the appearance of the component before and after the operation in the component operation, and corresponds to the 3D data ID 80102 of the component data model 801. The operation time 8066 is the start time and end time of the operation when the start time of the action script is set to 0 seconds. The work description 8067 is a description of the component operation described in natural language.
[0032] FIG. 8 is a functional block diagram showing the functions provided by the server 1 and the client terminal 2 as functional blocks. The server 1 includes a language model 11, a data acquisition unit 12, and a relevance determination unit 13. The client terminal 2 includes a speech recognition unit 21, a context transmission unit 22, a component data acquisition unit 23, a script execution unit 24, a drawing unit 25, and a data storage unit 26.
[0033] The language model 11 is a large language model. When a question is input into the language model 11 in natural language, an answer in natural language can be obtained. The language model 11 also has the function of vectorizing character strings. However, it is not an essential configuration that the language model 11 is stored inside the server 1, and the language model 11 existing outside the server 1 may be used via communication. The data acquisition unit 12 acquires various data used by the relevance determination unit 13 and the data requested from the client terminal 2 from the database 3. However, the description regarding the data acquisition unit 12 will be omitted below, and it will be described as if the relevance determination unit 13 directly acquires data from the database 3, for example. The relevance determination unit 13 performs estimation of the likelihood of action scripts and calculation of the importance of components, etc.
[0034] The speech recognition unit 21 converts the speech of the user 5 into a character string. The context transmission unit 22 collects the user context indicating the intention of the user 5 and transmits it to the server 1. The user context is, for example, the speech of the user 5 converted into a character string, the characters input by the user 5, the object pointed to by the user 5 with the mouse pointer, the object indicated by the user 5 in the virtual space, etc. The component data acquisition unit 23 determines the necessity of acquiring component data and downloads the necessary component data from the server 1. The script execution unit 24 executes the action script designated from the server 1. The script execution unit 24 also acquires the component data necessary for the execution of the action script from the server 1. The rendering unit 25 renders an image of the virtual space that reflects the movement of the user 5 in the virtual space and the change in the environment due to the action script executed by the script execution unit 24. The data storage unit 26 stores the data transmitted from the server 1.
[0035] FIG. 9 is a diagram showing an example of a virtual object 6 drawn in a virtual space. FIG. 6(a) shows a schematic representation of the virtual object 6, and FIG. 6(b) shows a detailed representation of the virtual object 6. In the detailed representation of the virtual object 6 shown in FIG. 6(b), internal components 6a and attachment components 6b that are not included in FIG. 6(a) are also shown. The internal component 6a is a component that exists inside the virtual object 6 and cannot be visually recognized from the outside. The attachment component 6b is a screw or the like.
[0036] Although the database 3 stores detailed data of the virtual object 6, it is not desirable to always display detailed data as it would result in excessive use of resources. Specifically, there is a problem that the internal components 6a and attachment components 6b stored in the database 3 need to be acquired via the server 1, and furthermore, the processing load on the drawing unit 25 increases. Therefore, as will be described later, the server 1 sets an importance level W for the components, and the client terminal 2 preferentially draws the components with a high importance level W.
[0037] FIG. 10 is a transition diagram showing the overall processing of the virtual space display system. The left side of the figure shows the processing of the client terminal 2, and the right side of the figure shows the processing of the server 1. First, the context transmission unit 22 of the client terminal 2 transmits user context such as the speech of the user 5 to the server 1, and further, the drawing unit 25 transmits the actions of the user 5 in the virtual space. The server 1 estimates the likelihood of the action script using the received data. Specifically, a prompt is added to a large language model to give an instruction process, estimates which action script the user 5 hopes to execute, and calculates the probability as the likelihood. The details of these processes will be described later with reference to FIGS. 11 and 12.
[0038] When the calculated likelihood exceeds the threshold, Server 1 instructs Client Terminal 2 to execute the action script. The script execution unit 24 of Client Terminal 2 executes the action script instructed by Server 1. Server 1 weights the context target and calculates the importance of the components. Then Server 1 notifies Client Terminal 2 of the importance of each component. Client Terminal 2 starts and monitors the download of component data according to the notified importance of the components. The rendering unit 25 of the client performs rendering using the component data for which the download has been completed.
[0039] Figure 11 is a flowchart showing the processing of the context transmission unit 22. Note that the processing of the speech recognition unit 21 is also included in Figure 11. First, in step S301, the speech recognition unit 21 of Client Terminal 2 performs a sound collection process from the microphone included in the input device 104. In the subsequent step S302, the speech recognition unit 21 performs a speech recognition process on the sound collection data collected in step S301 and converts the speech of User 5 into a character string (hereinafter referred to as "utterance text").
[0040] In the subsequent step S303, the context transmission unit 22 receives the pointer input of User 5. The pointer input is the position of the mouse pointer when User 5 operates the mouse which is the input device 104. When a head-mounted display is used as Client Terminal 2, the direction of User 5's line of sight is used as the pointer input. In the subsequent step S304, the context transmission unit 22 discriminates the designated object of User 5. Specifically, the context transmission unit 22 designates the object pointed to by the pointer or the object existing at the tip of User 5's line of sight as the designated object. For example, the context transmission unit 22 specifies the coordinate position of the object designated by User 5 in the virtual space.
[0041] In the subsequent step S305, the context transmission unit 22 identifies the target specified by the user 5 as a component model. The context transmission unit 22 can identify the component model specified by the user 5 by comparing the coordinate position specified in step S304 with the position information 80104 of the component data model 801. In the subsequent step S306, the context transmission unit 22 transmits the uttered text obtained in step S304 and the identifier of the component model identified in step S305 to the server 1.
[0042] In the subsequent step S307, the context transmission unit 22 updates the position and orientation of the user 5 in the virtual space based on the input of the user 5. In the subsequent step S308, the context transmission unit 22 generates an image of the virtual space as seen from the viewpoint of the user 5 in the virtual space based on the position and orientation of the user 5 updated in step S307, and outputs it to the display device 105. In the subsequent step S309, the context transmission unit 22 determines whether the processing of the client terminal 2 has ended. If it is determined that the processing has not ended, the context transmission unit 22 returns to step S301. If it is determined that the processing has ended, the context transmission unit 22 ends the processing shown in FIG. 11. For example, the context transmission unit 22 determines that the processing has ended when the power switch of the client terminal 2 is pressed during operation, and determines that the processing has not ended if the power switch is not pressed.
[0043] FIG. 12 is a flowchart showing the processing of the server 1. First, in step S321, the relevance determination unit 13 of the server 1 verbalizes the scene situation and the user situation. For example, the relevance determination unit 13 generates a sentence such as "The user is in a train. There are two seats, two lights, and two doors in this train. The user is sitting on a seat and facing forward." Since the server 1 can obtain the data of the virtual space from the database 3 and can obtain the data regarding the position or movement of the user 5 in the virtual space from the client terminal 2, the execution of step S321 is possible. In the subsequent step S322, the relevance determination unit 13 receives from the client terminal 2 the user action transmitted in step S306 of FIG. 11, that is, the character string of the utterance of the user 5 and the name of the specified component model.
[0044] In the next step S323, the relevance determination unit 13 evaluates the proximity in the language vector between the situation and the action of the user 5 and each action script included in the action script collection 805. Specifically, the relevance determination unit 13 first acquires the action script collection 805 from the database 3. Next, the relevance determination unit 13 combines the scene situation and the user situation verbalized in step S321 with the character string of the user action received in step S322 and the name of the specified part model, and vectorizes them using the language model 11. This vector is called "vector U" in the following description. Next, the relevance determination unit 13 vectorizes the task description 8057 of each action script in the action script collection 805 using the language model 11. In the following description, these vectors are called "vector S1", "vector S2", "vector S3", ... The number after "S" indicates the number of the action script for convenience. Then, the relevance determination unit 13 calculates the distance D1 which is the distance between the vector U and the vector S1, and the distance D2 which is the distance between the vector U and the vector S2, ...
[0045] In the next step S324, the relevance determination unit 13 evaluates the keywords described in the tag list 80107 of the part data model 801 of the target object of each action script. Specifically, the relevance determination unit 13 extracts all part IDs described in the part set 8055 of each action script, and identifies all tag lists 80107 corresponding to each relevant part ID from the part data model 801. A set of tags described in all the identified tag lists 80107 becomes the tag corresponding to the action script. Then, the relevance determination unit 13 counts the number of each action script included in the character string corresponding to the above-mentioned vector U. That is, the evaluation of the tag list 80107 in this step means counting the number of tags included in the tag list 80107 related to each action script that are included in the character string corresponding to the above-mentioned vector U.
[0046] In the subsequent step S325, the relevance determination unit 13 determines an action script to be input to the language model 11. The determination of this action script is determined by multiplying the distance calculated in step S323 by a coefficient based on the evaluation result of step S324. The relevance determination unit 13 corrects the distance of the action script by a previously specified amount according to the count number for each action script in step S324. For example, when one tag corresponding to the above character string is included, it is multiplied by 0.9, and when two tags are included, it is multiplied by 0.85. The relevance determination unit 13 sets an action script whose value multiplied by the coefficient is smaller than a predetermined threshold, or a predetermined number of action scripts starting from the one with the smaller value multiplied by the coefficient as the action script to be input to the language model 11.
[0047] In step S326, the relevance determination unit 13 inputs a question sentence using the action script determined in step S325 as an option to the language model 11 to obtain the likelihood of execution by the user for each action script. This question sentence is composed of, for example, an explanation of the situation by the character string corresponding to the vector U described above, an explanation that each action script determined in step S325 is a candidate for execution, and an instruction to output the probability that the user 5 wishes to execute each action script as a numerical value from 0 to 10. In this case, each action script may describe the action name 8052 or the operation explanation 8057.
[0048] In the subsequent steps S327 to S332, the relevance determination unit 13 repeatedly executes each of the action scripts input to the language model 11 in order as the target script. For example, when 10 action scripts are determined in step S325, the relevance determination unit 13 repeats the processes of steps S328 to S331 10 times.
[0049] In step S328, the relevance determination unit 13 determines whether the execution likelihood of the processing script exceeds a preset threshold value. If the relevance determination unit 13 determines that the execution likelihood of the processing script exceeds the preset threshold value, it proceeds to step S329; if it determines that the execution likelihood of the processing script does not exceed the preset threshold value, it proceeds to step S330. In step S329, the relevance determination unit 13 instructs the client terminal 2 to execute the processing script and proceeds to step S331. As will be described later, when the client terminal 2 is instructed to execute the action script and does not have the component data necessary for the execution of the action script, it downloads data from the server 1.
[0050] In step S330, which is executed when a negative determination is made in step S328, the relevance determination unit 13 updates the importance W of each component necessary for the execution of the target script, that is, each component included in the component set 8055 of the target script, and proceeds to step S331. The update of the importance W will be described later. In step S331, the relevance determination unit 13 updates the importance of each component affected by the execution of the target script and proceeds to step S332. Also, the components targeted in step S331 are the components affected by the execution of the target script. For example, when the target script is the lighting of a light, the components illuminated by the lit light are the components targeted in step S331. Note that the relevance determination unit 13 may consider the magnitude of the likelihood of the target script in the update of the importance in steps S330 and S331.
[0051] When the relevance determination unit 13 finishes the processing of steps S328 to S331 for all the action scripts determined in step S325, it proceeds to step S333. In step S333, the relevance determination unit 13 causes the client terminal 2 to download the components whose importance has exceeded the threshold value due to the updates in steps S330 and S331, and ends the processing shown in FIG. 12.
[0052] The relevance determination unit 13 calculates the importance W for each component. The importance W is the sum of the user importance Wu indicating the relationship with the user, the event importance We indicating the relationship of the event, and the correlation importance Wp indicating the relationship with other components. The larger the value of the importance W, the more important it is.
[0053] The user importance Wu is set for components existing within a predetermined range from the perspective of user 5, and components pointed to by the mouse operated by user 5 or mentioned verbally. The user importance Wu may be a predetermined constant value, or may be set to a lower value as it deviates from the center of the perspective of user 5. The event importance We is set for components necessary for the execution of the action script. The event importance We may be a predetermined constant value, or a value corresponding to the likelihood of the execution of the action script. Specifically, the higher the likelihood, the higher the event importance We may be set.
[0054] The correlation importance Wp is an index of the correlation with all other components in which at least one of the user importance Wu and the event importance We is set. The correlation importance Wp(A) of component A is calculated as follows using three distance functions indicating the relationship with other component X, namely the joint distance function D_connect(A,X), the attribute distance function D_feature(A,X), and the construction method distance function D_process(A,X).
[0055] Wp(A) = Σ_X e^ -fn( D_connect(A,X), D_feature(A,X), D_process(A,X) ) ···(Equation 1)
[0056] However, X is a component other than component A for which at least one of the user importance Wu and the event importance We is set. The greater the number of other components for which the user importance Wu or the event importance We is set, the greater the value of the correlation importance Wp(A). The function fn is, for example, a function that outputs the sum of three arguments. Below, the definitions of the joining distance function, the attribute distance function, and the construction method distance function will be described. Note that below, the joining distance is also referred to as the "joining relationship", the attribute distance is also referred to as the "attribute relationship", and the construction method distance is also referred to as the "construction method relationship".
[0057] The joining distance function D_connect(A,B) is "d_connect" when component A is fixed to component B. Also, when object C is fixed to component B, the joining distance function D_connect(A,C) is "d_connect*2". The connection between components can be read from the items of the fixed parent component 80114 and the fixed child component 80116 of the component data model 801. When component A and component B are not in a fixed relationship with each other, the joining distance function D_connect(A,B) is infinite (a sufficiently large value).
[0058] The attribute distance function D_feature(A,B) sets the attribute distance to be small when component A and component B have the same model number, etc. For example, when component A and component B are included in the same type of device 80110, the attribute distance is set to "d_feature", and for objects included in the related similarity 80111, the attribute distance is set to "d_feature*2". Also, when component A and component B are made of the same material, the attribute distance is set to "d_feature*3". Furthermore, when component A and component B are not in a reference relationship with each other, the attribute distance is set to be infinite (a sufficiently large value).
[0059] The process distance function D_process(A,B) is set to have a smaller process distance as the similarity between part A and part B in the operation is higher. The process distance function utilizes three variables, namely d_sameworker, d_sameperiod, and d_sametool. For example, when part A and part B are parts of the same worker in the operation, a value smaller than 1 is set for d_sameworker; when they are parts of the same day's operation, a value smaller than 1 is set for d_sameperiod; and when they are parts of the operation using the same equipment, a value smaller than 1 is set for d_sametool. The process distance function D_process(A,B) is the product of the constant d0, d_sameworker, d_sameperiod, and d_sametool.
[0060] Note that each parameter of the distance calculation may be changed according to the utterance of user 5 and the profile of user 5 who made the utterance. For example, when user 5 is actively involved in the conversation and is responsible for the business of part attachment management using the vocabulary included in the utterance content of user 5 and the information of the operation speaker, d_connect may be set to a small value. Also, if the utterance of user 5 is related to the operation content, d_sameworker may be set to a small value. Further, if the utterance includes a date, d_sameperiod may be set to a small value.
[0061] Figure 13 is a flowchart showing the processing of the part data acquisition unit 23 of the client terminal 2. The part data acquisition unit 23 executes the processing shown in Figure 13 every time the importance of a part is notified from the server 1. The processing shown in Figure 13 has the first step S341 as the start of the loop and the last step S347 as the end of the loop, and the processing of steps S342 to S346 is repeated with the part whose importance has been notified as the "target part" in order. Therefore, the processing of steps S342 to S346 will be described below.
[0062] The component data acquisition unit 23 determines whether the component data of the target component, that is, the record of the component data model 801, is stored inside the client terminal 2 in step S342. If the component data acquisition unit 23 determines that the component data of the target component is stored inside the client terminal 2, it proceeds to step S347. If it determines that the component data of the target component is not stored inside the client terminal 2, it proceeds to step S343.
[0063] In step S343, the component data acquisition unit 23 determines whether the importance level W of the target component exceeds a threshold value. If the component data acquisition unit 23 determines that the importance level W of the target component exceeds the threshold value, it proceeds to step S344. If it determines that the importance level W of the target component does not exceed the threshold value, it proceeds to step S345. In step S344, the component data acquisition unit 23 starts downloading the component data of the target component, that is, starts acquiring it from the server 1 and proceeds to step S347. If the component data of the target component is already being downloaded, in step S344, it proceeds to step S347 without any special processing.
[0064] In step S345, the component data acquisition unit 23 determines whether the component data of the target component is being downloaded. If the component data acquisition unit 23 determines that the component data of the target component is being downloaded, it proceeds to step S346. If it determines that the component data of the target component is not being downloaded, it proceeds to step S347. In step S346, the component data acquisition unit 23 aborts the download of the component data of the target component and proceeds to step S347.
[0065] FIG. 14 is a flowchart showing the processing of the script execution unit 24 of the client terminal 2. The script execution unit 24 executes the processing shown in FIG. 14 every time it is instructed to execute an action script from the server 1. The script execution unit 24 receives the action script to be executed from the server 1, the record of the component operation set 806 called by this action script, and the start time of the action script.
[0066] In step S351, the script execution unit 24 analyzes the action script to be executed and enumerates all the components included in the action script. Specifically, the script execution unit 24 extracts all the component IDs described in the component set 8055 of the action script. In the subsequent step S352, the script execution unit 24 determines whether all the component data used by the action script is stored inside the client terminal 2. If the script execution unit 24 determines that all the component data is stored inside the client terminal 2, it proceeds to step S354; if it determines that there is component data not stored inside the client terminal 2, it proceeds to step S353.
[0067] In step S353, the script execution unit 24 starts downloading the missing component data and proceeds to step S354. In step S354, the script execution unit 24 executes the action script and ends the process shown in FIG. 14. At this time, the script execution unit 24 sets the time of "0" that serves as the reference for the operation time 8066 in the component operation set 806 to the start time of the action script received from the server 1. Then, the script execution unit 24 appropriately rewrites the position information 80104 and rotation information 80105 of each component of the component data model 801 in the client terminal 2 according to the description in the component operation set 806.
[0068] FIG. 15 is a flowchart showing the process of the drawing unit 25 of the client terminal 2. The drawing unit 25 executes the process shown in FIG. 15 at regular time intervals, for example, with a period of 100 ms. First, in step S361, the drawing unit 25 receives the movement of the user 5 in the virtual space. For example, the drawing unit 25 receives instructions for movement and line-of-sight movement within the virtual space using the mouse or keyboard, which is the input device 104 of the user 5. In the subsequent step S362, the drawing unit 25 transmits the movement and line-of-sight movement of the user 5 received in step S361 within the virtual space to the server 1.
[0069] In the subsequent step S363, the drawing unit 25 reads the data of each component in the component data model 801 stored in the client terminal 2. Since the position information 80104 of the component data is appropriately rewritten by the process of step S354 in FIG. 14, the latest position and orientation updated for each component are reflected in the drawing. In the subsequent step S364, the drawing unit 25 performs rendering using the component data read in step S365, draws the virtual space around the user 5 on the display device 105, and ends the process shown in FIG. 15. However, in step S364, components for which the download has not been completed are not drawn, and when the drawing is completed, those components are drawn. The drawing unit 25 refers to the importance level W of the component notified from the server 1 and targets for drawing components whose importance level W is greater than a predetermined threshold value.
[0070] According to the first embodiment described above, the following operational effects can be obtained. (1) The server 1, which can also be called an arithmetic unit, controls the display in the virtual space. The server 1 acquires the user context, which is data indicating the intention of the user 5 and is transmitted from the context transmission unit 22 of the client terminal 2 (S322 in FIG. 12). The server 1 acquires an action script indicating the operation of the component in the virtual space (S323 in FIG. 12). The server 1 inputs the state of the virtual space, the user context, and the options of the plurality of action scripts into the language model 11 to obtain the likelihoods of the options of the plurality of action scripts (S326 in FIG. 12). The server 1 changes the display in the virtual space based on the likelihood by transmitting the calculated importance level W to the client terminal 2 (S333 in FIG. 12). Therefore, by setting the importance level W for each component based on the actions of the user 5, the display in the virtual space can be controlled.
[0071] (2) The relevance determination unit 13 of the server 1 sets the event importance level We, which is the importance level of the component required for the execution of the action script, based on the likelihood of the action script. The relevance determination unit 13 transmits the importance level W of each component to the client terminal 2 to cause the client terminal 2 to determine the presence or absence of the display of the component based on the importance level W.
[0072] (3) The server 1 includes a network IF 109 that communicates with a client terminal 2 having a rendering unit 25 that generates an image of a virtual space. The relevance determination unit 13 sets an event importance We, which is the importance W of the parts necessary for the execution of the action script, based on the likelihood of the action script. The relevance determination unit 13 causes the client terminal 2 to download part data based on the magnitude of the importance W by transmitting the importance W of each part to the client terminal 2.
[0073] (4) The relevance determination unit 13 further sets the importance W using at least one of a joining distance function D_connect(A,X), an attribute distance function D_feature(A,X), and a construction method distance function D_process(A,X) for the parts with the event importance We set.
[0074] (5) The relevance determination unit 13 sets a user importance Wu for the parts based on the viewpoint in the virtual space of the user 5, sets a correlation importance Wp based on at least one of a joining distance function D_connect(A,X), an attribute distance function D_feature(A,X), and a construction method distance function D_process(A,X) for at least the parts with the user importance Wu set, and sets the importance W based on at least the user importance Wu and the correlation importance Wp.
[0075] (6) The server 1 includes at least one processor. The processor of the server 1 acquires user context, which is data indicating the intention of the user 5 (S322 in FIG. 12), acquires an action script indicating the operation of the parts in the virtual space (S323 in FIG. 12), inputs the state of the virtual space, the user context, and the options of a plurality of action scripts into a language model 11 to acquire the likelihoods of the options of the plurality of action scripts (S326 in FIG. 12), sets the importance for each part based on the likelihoods, and changes the display in the virtual space.
[0076] (Modification Example 1) In the above-described embodiment, the importance W was the sum of the user importance Wu, the event importance We, and the correlation importance Wp. However, the importance W may be the sum of any two of the three, namely the user importance Wu, the event importance We, and the correlation importance Wp, or any one of the user importance Wu, the event importance We, and the correlation importance Wp may be regarded as the importance W. Also, the correlation importance Wp was expressed using the three functions: the joining distance function D_connect(A,X), the attribute distance function D_feature(A,X), and the construction method distance function D_process(A,X). However, it is not necessary to use all three, and at least one of them may be used.
[0077] (Modification Example 2) In the above-described embodiment, the virtual space display system included the server 1 and the client terminal 2. However, the server 1 and the client terminal 2 may be integrally configured. Also, the database 3 may be integrally configured with the client terminal 2. In this case, since it is not necessary to download the component data in the client terminal 2, it is not necessary to determine whether or not to download.
[0078] However, even when the database 3 is integrally configured with the client terminal 2, since it takes time to read data from a low-speed external storage device to a high-speed main memory device, this data reading may be treated in the same way as data downloading via the Internet. That is, when the database 3 is stored in a low-speed external storage device and read into the high-speed main memory device of the client terminal 2, instead of determining whether or not to download as in the above-described embodiment, it may be determined whether or not to read data from the external storage device to the main memory device. In this case, there is an advantage that the display ability can be allocated to more necessary targets in an environment where the capacity of the high-speed main memory device is limited.
[0079] (Modification Example 3) In the function fn described in Equation 1 for calculating the correlation importance Wp, the ratios of the joining distance function D_connect(A,X), the attribute distance function D_feature(A,X), and the construction method distance function D_process(A,X) may be varied according to the user context received from the client terminal 2. For example, when the speech of user 5 is about the material of a component, the proportion of the joining distance function is increased, and when the speech of user 5 is about the construction method, the ratio of the construction method distance function is increased.
[0080] Whether the user context has a strong relevance to any of the joining distance function, the attribute distance function, and the construction method distance function can be determined using the language model 11. For example, the relevance determination unit 13 inputs the user context, the definition of each distance function, and an instruction text "Please evaluate the strength of the user's interest for each distance function as a decimal number from 0 to 1 based on the user's speech" into the language model 11. Then, the relevance determination unit 13 can reflect the output of the language model 11 in the function fn.
[0081] In each of the above-described embodiments and modifications, the configuration of the functional blocks is merely an example. Some of the functional configurations shown as separate functional blocks may be integrated, or the configuration represented by one functional block diagram may be divided into two or more functions. Also, a configuration may be adopted in which a part of the functions of each functional block is provided by other functional blocks.
[0082] In each of the above-described embodiments and modifications, the server 1 and the client terminal 2 are provided with input / output interfaces (not shown), and when necessary, a program may be read from another device via the input / output interfaces and a medium available to the server 1 and the client terminal 2. Here, the medium refers to, for example, a removable storage medium attached to the input / output interface, or a communication medium, that is, a wired, wireless, optical, etc. network, or a carrier wave or digital signal propagating through the network. Also, part or all of the functions realized by the program may be realized by a hardware circuit or an FPGA.
[0083] Each of the above-described embodiments and modifications may be combined. Although various embodiments and modifications have been described above, the present invention is not limited to these contents. Other aspects conceivable within the scope of the technical idea of the present invention are also included in the scope of the present invention.
Explanation of Signs
[0084] 1: Server 2: Client terminal 3: Database 5: User 6: Virtual object 11: Language model 12: Data acquisition unit 13: Relevance determination unit 21: Speech recognition unit 22: Context transmission unit 23: Component data acquisition unit 24: Script execution unit 25: Drawing unit 102: CPU 109: Network IF 805: Action script collection W: Importance We: Event importance Wp: Correlation importance Wu: User importance
Claims
1. A control method for a computing device to control a display in a virtual space, comprising: a context acquisition process for acquiring a user context, which is data indicating the intention of a user; a script acquisition process for acquiring an action script indicating the operation of components in the virtual space; a likelihood calculation process for inputting the state of the virtual space, the user context, and options of a plurality of the action scripts into a language model to obtain the likelihoods of the options of the plurality of action scripts; and a display change process for changing the display in the virtual space based on the likelihoods.
2. The control method according to claim 1, wherein in the display change process, an event importance, which is the importance of the components required for the execution of the action script, is set based on the likelihood of the action script, and the presence or absence of the display of the components is determined based on the magnitude of the importance.
3. The control method according to claim 1, wherein the computing device further comprises a communication unit for communicating with a client terminal including a rendering unit that generates a video of the virtual space, and in the display change process, an event importance, which is the importance of the components required for the execution of the action script, is set based on the likelihood of the action script, and the client terminal is caused to download the data of the components based on the magnitude of the importance.
4. The control method according to claim 2, wherein in the display change process, a correlation importance is set based on at least one of a joining relationship, an attribute relationship, and a construction method relationship with the components for which the event importance is set, and the importance is set based on at least the event importance and the correlation importance.
5. The control method according to claim 4, wherein in the display change process, based on the user context, a ratio of the joining relationship, the attribute relationship, and the construction method relationship for setting the correlation importance is determined.
6. The control method according to claim 3, wherein in the display change process, a correlation importance is set based on at least one of a joining relationship, an attribute relationship, and a construction method relationship with the components for which the event importance is set, and the importance is set based on at least the event importance and the correlation importance.
7. The control method according to claim 6, wherein In the display change process, a control method for determining ratios of the joining relationship, the attribute relationship, and the construction method relationship for setting the correlation importance based on the user context.
8. The control method according to claim 1, wherein in the display change process, a user importance is set for the component based on the viewpoint of the user in the virtual space, a correlation importance is set based on at least one of the joining relationship, the attribute relationship, and the construction method relationship with the component for which the user importance is set, and the presence or absence of display of the component is determined based on the user importance and the correlation importance.
9. The control method according to claim 1, wherein at least one of the action scripts includes a process of displaying data related to the component in the virtual space.
10. An arithmetic device including at least one processor, wherein the processor acquires a user context which is data indicating the intention of the user, acquires an action script indicating the operation of a component in the virtual space, inputs the state of the virtual space, the user context, and options of a plurality of the action scripts into a language model to obtain likelihoods of the options of the plurality of action scripts, and changes the display in the virtual space based on the likelihoods.
Citation Information
Patent Citations
Model display control method and model display control program and model display control system
JP2019174982A